Convolutional blind source separation method and system based on impulse response compression and storage medium
By constructing a global impulse response network and nonnegative matrix factorization, combined with iterative source control, the convolution blind source separation method is optimized, solving the problems of robustness and computational complexity in signal separation under high reverberation environments, and achieving efficient signal separation results.
Patent Information
- Application Number
- CN202511664377.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-13
- Publication Date
- 2025-12-30
AI Technical Summary
Existing convolutional blind source separation methods suffer from performance degradation in high reverberation environments. Key parameters rely on empirical settings, resulting in insufficient robustness and difficulty in handling blind signal processing scenarios without prior knowledge.
A global impulse response network is constructed, and the impulse response is adjusted by desired and undesired windows. The objective function is optimized using the gradient descent algorithm, and the separation of sound source signals is optimized by combining nonnegative matrix factorization and iterative source control.
It effectively suppresses the effects of high reverberation, improves the robustness of the method, reduces computational complexity, is suitable for blind signal processing scenarios, and improves signal separation accuracy.
Smart Images

Figure CN121237114A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of blind source separation, and particularly relates to a convolution blind source separation method and system based on impulse response compression and a storage medium. BACKGROUND
[0002] The blind source separation (BSS) method model can be divided into a linear instantaneous mixing model and a linear convolution mixing model, and the convolution mixing is more complex than the instantaneous mixing. The convolution mixing in the time domain can be converted into instantaneous mixing at each frequency point in the frequency domain through short-time Fourier transform (STFT), and then the instantaneous BSS algorithm can be used for separation at each frequency point. However, the influence of the amplitude ambiguity and the ordering ambiguity of the separated signals at each frequency point on the method performance still needs to be considered. In addition, the traditional BSS method uses short-time Fourier transform under the linear narrowband approximation assumption, and requires that the impulse response length is much shorter than the STFT window length. This is not practical in the case of high reverberation, i.e. long impulse response length, and therefore the performance of the traditional BSS method deteriorates or even fails under the condition of high reverberation.
[0003] Through the search of existing literature, it is found that H. Kagami et al. published in IEEE International Conference on Acoustics, Speech and Signal Processing (2018, Canada, pp. 31-35) “Joint separation and dereverberation of reverberant mixtures with determined multichannel nonnegative matrix factorization” proposes a method of combining the weighted prediction error (WPE) dereverberation method with the independent low rank matrix analysis (ILRMA) optimization method, named WPE-ILRMA, to suppress the influence of reverberation environment on the method performance, but the traditional WPE technology needs to introduce a long prediction filter to dereverberate the mixed signal, resulting in high computational complexity. Wang Taiyan et al. published in Acta Acustica Sinica (2024, vol. 49, no. 1, pp. 163-170) “Low complexity joint optimization method of blind source separation and dereverberation” proposes a low complexity WPE-ILRMA, but this method still needs to initialize the dereverberation signal and the separation signal, and when the filter order is only 20, the single iteration time is 3.29s. Even compared with the traditional WPE-ILRMA single iteration time, it has been greatly reduced, but it still cannot realize signal separation in high reverberation environment. Jie Yu et al. published in Electronics Letters (2023, vol. 51, no. 11) “Blind source separation combining impulse response reshaping and expectation maximization” proposes an impulse response reshaping model using gradient descent method to optimize the pre-filter, which shortens the impulse response length and reduces the influence of high reverberation environment, but the objective function constructed based on the impulse response reshaping model needs to set the norm value, and the norm value is set by experience, and improper setting will easily lead to the decline of impulse response reshaping effect, and in the blind signal processing scene without any prior information, it is not practical to adjust the norm value of the objective function according to the impulse response reshaping effect.
[0004] Existing literature shows that although the existing convolutional mixing BSS method in reverberation environment has made certain progress, there are still engineering problems to be solved. The key parameters of the existing method model that can effectively deal with high reverberation environment are set by experience, and the robustness is insufficient, which is difficult to deal with the blind signal processing scene without any prior knowledge. SUMMARY
[0005] The purpose of this invention is to solve the engineering problem that existing convolutional blind source separation methods cannot cope with high reverberation environments, and to provide a convolutional blind source separation method, system and storage medium based on impulse response compression.
[0006] The objective of this invention is achieved through the following technical solution:
[0007] A convolutional blind source separation method based on impulse response compression includes:
[0008] Step 1: Construct a convolutional mixed signal model and a global impulse response network. The global impulse response network is implemented by installing a loudspeaker in front of the receiving microphone to reduce the impact of the high reverberation environment on the performance of the method.
[0009] Step 2: Construct an objective function based on a global impulse response network. The objective function adjusts the global impulse response by introducing desired and undesirable windows to minimize the ratio of undesirable impulse response to desired impulse response.
[0010] Step 3: Update the impulse response generated between the sound source and the loudspeaker using the gradient descent algorithm to optimize the objective function;
[0011] Step 4: Determine if the gradient descent iterative model has reached its maximum number of iterations; if not, return to Step 3 and iterate again; otherwise, output the optimized global impulse response.
[0012] Step 5: Perform a short-time Fourier transform on the received signal to convert the signal to the time-frequency domain, and construct a source signal model based on non-negative matrix decomposition to simulate the sound source signal. Solve for the unmixing matrix and source signal parameters by minimizing the negative log-likelihood function.
[0013] Step 6: Update the source basis matrix and source activation matrix, introduce iterative source control to update the sound source signal, and normalize the updated source basis matrix and sound source signal;
[0014] Step 7: Determine if the nonnegative matrix decomposition iterative model has reached its maximum number of iterations; if not, return to Step 6 and iterate again; otherwise, output the separated signals at each frequency point.
[0015] Step 8: Use the back projection formula to resolve the ambiguity in the order of the separated signals at each frequency point, and perform a short-time Fourier inverse transform on the frequency domain separated signals to obtain the recovered time-domain sound source signal.
[0016] Furthermore, the convolutional mixed signal model constructed in step one is as follows:
[0017]
[0018] in, , The number of sampling points. The maximum number of sampling points, Let be the impulse response length, and let the additive Gaussian noise in the channel be represented as . ; For time delay; For time delay The aliasing system matrix below; The sound source signal; For delay The sound source signal after;
[0019] Install before receiving the microphone The loudspeaker, generated by the entire global impulse response network, is the first... The sound source signal propagates to the first... The length of each microphone is Spatial impulse response for:
[0020]
[0021] Among them, the The sound source signal propagates to the first... The length generated during the process of each loudspeaker is The spatial impulse response is expressed as , No. The loudspeaker broadcast to the... The length generated during the process of each microphone is The spatial impulse response is expressed as , Indicates by generated 3D convolution matrix, , constitute indivual 3D global impulse response matrix Then, the convolutional mixed signal model based on the global impulse response network is expressed as:
[0022]
[0023] Furthermore, the desired window in step two And Unexpected Window ,
[0024]
[0025]
[0026]
[0027] in, , , , , Represents the signal sampling frequency. , and The window function adjustment factor is used to obtain the desired impulse response from the global impulse response. Unexpected impulse response ,in Dot product;
[0028] The objective function is constructed based on a global impulse network as follows:
[0029]
[0030] in, Represents the natural logarithm, and
[0031]
[0032]
[0033] in, This represents taking the maximum value, because , Therefore, the impulse response is not desired. The generated by The function value of the variable is ; u is a symbol representing the undesirable value. For the undesired impulse response The generated by is the function value of the variable; d is the symbol representing the expectation.
[0034] Furthermore, in step three, the partial derivative of the objective function is calculated:
[0035]
[0036] in,
[0037]
[0038]
[0039]
[0040] in, Represents matrix diagonalization. Represents the symbolic function. represent Maximum value position index represent Located at the position index elements, represent Located at the position index elements, represent Located at the position index elements, , Represents the convolution matrix Located at the position index The row vector.
[0041] The update rule using the gradient descent algorithm is as follows:
[0042]
[0043] in Represents the number of iterations. The learning rate is represented by the maximum number of iterations. .
[0044] Furthermore, the convolutional mixed signal based on the global impulse response network in step five is represented by a short-time Fourier transform as follows:
[0045]
[0046] in, Represents the short-time Fourier transform at frequency points The following is a mixture matrix, , and These represent the received signal in the time-frequency domain, the sound source signal in the time-frequency domain, and the noise in the time-frequency domain, respectively. and These represent frequency points and time frames, respectively. and These represent the maximum frequency point and the maximum time frame, respectively. Time-frequency domain sound source signal of 3D variance matrix Satisfies nonnegative matrix decomposition. ,Right now In this model, it is assumed that the source signals have zero mean and are independent and identically distributed. In OK Column elements for Expected value;
[0047] The nonnegative matrix factorization model is constructed as follows:
[0048]
[0049] in, and Representing the first Time-frequency domain sound source signal variance matrix Decomposition Wikipedia matrix and 3D activation matrix, The cardinality of the nonnegative matrix factorization model; defining frequency points. Below The dimensional unmixing matrix is The recovered frequency domain source signal is:
[0050]
[0051] Solving by minimizing the negative log-likelihood function. Non-negative matrices , The negative log-likelihood function is expressed as follows:
[0052]
[0053] in For the unmixing matrix The column vectors, and Representing the basis matrices respectively of OK Column elements and activation matrix of OK Column elements, This represents taking the determinant of the matrix; ; ; ; .
[0054] Furthermore, step six initializes the nonnegative matrix. and The elements in the matrix are all random numbers that follow a uniform distribution between [0,1]. This represents the number of iterations; the maximum number of iterations is set to... , and Representing the first basis matrix of OK Column elements and activation matrix of OK Column elements; at frequency points On The 3D separated sound source signal matrix is represented as follows , represent of OK Column elements; , Representing the Variance matrix of OK Column elements;
[0055] First, minimizing the negative log-likelihood function is performed by iteratively updating the following spatial model parameters:
[0056]
[0057]
[0058]
[0059]
[0060] Then, an iterative source control algorithm is introduced to update the sound source signal:
[0061]
[0062] in, Represents the updated frequency. On 3D sound source signal matrix, Represents the frequency before the update The first road The sound source signal, and represent dimensional iterative source control vector, for vector medium elements Update ,when and When they are not equal:
[0063]
[0064]
[0065]
[0066] when and When they are equal:
[0067]
[0068] in, Represents taking the conjugate;
[0069] Finally, the updated source basis matrix and the sound source signal are normalized; Representative frequency point Sound source signal matrix of OK For column elements, the normalization operation is as follows:
[0070]
[0071]
[0072]
[0073] Furthermore, step eight involves selecting a frequency point. The first road 3D convolution mixed signal As a reference signal This represents the frequency points output after the iteration of the nonnegative matrix factorization iterative model. On Separate the sound source signal. This represents the frequency point recovered after applying the back projection formula. On Separate the sound source signal;
[0074] The back projection is as follows:
[0075]
[0076] Frequency domain separated signal Perform a short-time inverse Fourier transform to obtain the final time-domain sound source separation signal. .
[0077] A computer device / apparatus / system includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of a convolutional blind source separation method based on impulse response compression.
[0078] A computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of a convolutional blind source separation method based on impulse response compression.
[0079] The beneficial effects of this invention are as follows:
[0080] This invention designs a novel impulse response compression model to suppress the impact of high reverberation on the performance of the BSS method. Compared with the traditional impulse response reshaping model, the objective function constructed based on the impulse response compression model does not require empirical setting of the P-norm value. The method performance is not affected by the setting of key parameters, exhibiting stronger robustness and making it more suitable for blind signal processing scenarios. Furthermore, the introduction of an iterative source control vector to directly update the source separation signal avoids the large number of matrix inversion operations in the traditional ILRMA algorithm, thus preventing algorithm instability and high hardware resource consumption caused by excessive matrix inversion, and facilitating the engineering implementation of the algorithm. Attached Figure Description
[0081] Figure 1 : Schematic diagram of a convolutional blind source separation method based on impulse response compression;
[0082] Figure 2 : Schematic diagram of reverberation space;
[0083] Figure 3 Signal-to-interference ratio curve as a function of reverberation time
[0084] Figure 4 : Signal distortion ratio curve as a function of reverberation time;
[0085] Figure 5 : Signal artifact ratio curve as a function of reverberation time;
[0086] Figure 6 : Signal-to-interference ratio curve as a function of signal-to-noise ratio;
[0087] Figure 7 : Signal distortion ratio curve as a function of signal-to-noise ratio;
[0088] Figure 8 : Signal artifact ratio curve as a function of signal-to-noise ratio. Detailed Implementation
[0089] The present invention will now be further described with reference to the accompanying drawings. In this description, it should be noted that the terms "first," "second," and "third" used in the embodiments of the present invention are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first," "second," or "third" may explicitly or implicitly include one or more of that feature.
[0090] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0091] Specific Implementation Plan 1: Combining Figure 1As shown, the present invention provides a method for separating convolutional blind sources, comprising the following steps:
[0092] Step 1: Construct a convolutional mixed signal model and a global impulse response network, specifically including:
[0093] The convolutional mixed signal model is constructed as follows:
[0094]
[0095] in, , The number of sampling points. The maximum number of sampling points, , The number of receiving microphones, , This method studies the number of sound sources. The problem of convolutional blind source separation in the case of [missing information]. The received mixed signal can be represented as [missing information]. The sound source signal can be represented as , No. The sound source signal propagates to the first... The spatial impulse response generated by each microphone during the process is represented as follows: , Given the impulse response length, additive Gaussian noise in the channel can be expressed as: , For time delay; For time delay The aliasing system matrix below; For delay The sound source signal after that. It can constitute indivual 3D space impulse response matrix The convolutional mixed signal model can also be expressed as:
[0096]
[0097] To mitigate the impact of high reverberation environments on method performance, a global impulse response network is constructed. This network is installed before the receiving microphone. The loudspeaker, the first The sound source signal propagates to the first... The length generated during the process of each loudspeaker is The spatial impulse response is expressed as , No. The loudspeaker broadcast to the... The length generated during the process of each microphone is The spatial impulse response is expressed as The length generated by the entire global impulse response network is then... The spatial impulse response can be expressed as:
[0098]
[0099] in, Represents convolution multiplication, then Therefore, the above formula can also be expressed as:
[0100]
[0101] in, Indicates by generated 3D convolution matrix. It can constitute indivual 3D global impulse response matrix Then, the convolutional mixed signal model based on the global impulse response network can be expressed as:
[0102]
[0103] Step two involves constructing the objective function based on the global impulse response network, specifically including:
[0104] To adjust the generated global impulse response, two window functions are introduced, namely the desired window function and the window function. And Unexpected Window The definition is as follows:
[0105]
[0106]
[0107]
[0108] in, , , , , Represents the signal sampling frequency. , and The window function adjustment factor is used to obtain the desired impulse response from the global impulse response. Unexpected impulse response ,in The dot product is represented. The objective function, constructed based on a global impulse network, is as follows:
[0109]
[0110] in, Represents the natural logarithm, and
[0111]
[0112]
[0113] in, This represents taking the maximum value; because , Therefore, the impulse response is not desired. The generated by The function value of the variable is ; u is a symbol representing the undesirable value. For the undesired impulse response The generated by is the function value of the variable; d is the symbol representing the expectation.
[0114] Step 3: Update the impulse response generated between the sound source and the loudspeaker using the gradient descent algorithm. Specifically, this includes...
[0115] The gradient descent method is used to solve the objective function based on a global impulse network to optimize the impulse response generated between the sound source and the loudspeaker. The partial derivative of the objective function is then calculated as follows:
[0116]
[0117] in
[0118]
[0119]
[0120]
[0121] in, Represents matrix diagonalization. Represents the symbolic function. represent Maximum value position index represent Located at the position index elements, represent Located at the position index elements, represent Located at the position index elements, , Represents the convolution matrix Located at the position index The row vectors. The update rules based on the above analysis are as follows:
[0122]
[0123] in, Represents the number of iterations. The learning rate is represented by the maximum number of iterations. .
[0124] Step 4: Determine whether the gradient descent iterative model has reached its maximum number of iterations. If the condition is not met, return to step three and iterate again; otherwise, output the global impulse response. .
[0125] Step 5: Perform a short-time Fourier transform on the received signal, and construct a source signal model based on non-negative matrix decomposition to simulate the sound source signal. Construct a negative log-likelihood function, specifically including...
[0126] The convolutional mixture signal based on a global impulse response network can be expressed as follows after short-time Fourier transform:
[0127]
[0128] in, Represents the short-time Fourier transform at frequency points The following is a mixture matrix, , and These represent the received signal in the time-frequency domain, the sound source signal in the time-frequency domain, and the noise in the time-frequency domain, respectively. and These represent frequency points and time frames, respectively. and These represent the maximum frequency point and the maximum time frame, respectively. Time-frequency domain sound source signal of 3D variance matrix Satisfies nonnegative matrix decomposition. ,Right now In this model, it is assumed that the source signals have zero mean and are independent and identically distributed. In OK Column elements for The expected value. The nonnegative matrix factorization model is constructed as follows:
[0129]
[0130] in, and Representing the first Time-frequency domain sound source signal variance matrix Decomposition Wikipedia matrix and 3D activation matrix, The cardinality of the nonnegative matrix factorization model; defining frequency points. Below The dimensional unmixing matrix is The recovered frequency domain source signal is:
[0131]
[0132] The solution can be found by minimizing the negative log-likelihood function. Non-negative matrices , The negative log-likelihood function is expressed as follows:
[0133]
[0134] in, For the unmixing matrix The column vectors, and Representing the basis matrices respectively of OK Column elements and activation matrix of OK Column elements, This represents taking the determinant of the matrix; ; ; ; .
[0135] Step six involves updating the source basis matrix and source activation matrix, introducing iterative source control to update the sound source signal, and normalizing the updated source basis matrix and sound source signal. Specifically, this includes...
[0136] Initialization nonnegative matrix and The elements in the matrix are all random numbers that follow a uniform distribution between [0,1]. This represents the number of iterations; the maximum number of iterations is set to... , and Representing the first basis matrix of OK Column elements and activation matrix of OK Column elements; at frequency points On The 3D separated sound source signal matrix is represented as follows , represent of OK Column elements; , Representing the Variance matrix of OK Column elements.
[0137] First, minimizing the negative log-likelihood function can be performed by iterating over the following update rules for the spatial model parameters:
[0138]
[0139]
[0140]
[0141]
[0142] Then, an iterative source control algorithm is introduced to update the sound source signal:
[0143]
[0144] in Represents the updated frequency. On 3D sound source signal matrix, Represents the frequency before the update The first road The sound source signal, and represent dimensional iterative source control vector, for vector medium elements Update ,when and When they are not equal
[0145]
[0146]
[0147]
[0148] when and When equal
[0149]
[0150] in The term represents conjugate.
[0151] Finally, the updated source basis matrix and the sound source signal are normalized. Representative frequency point Sound source signal matrix of OK For column elements, the normalization operation is as follows:
[0152]
[0153]
[0154]
[0155] Step 7: Determine whether the nonnegative matrix factorization iterative model has reached its maximum number of iterations. If the desired result is not achieved, return to step six and iterate again; otherwise, output the separated signals for each frequency point. .
[0156] Step eight involves resolving the ambiguity in the order of the separated signals at each frequency point using the back projection formula, and then performing a short-time Fourier inverse transform to obtain the recovered sound source signal. Specifically, this includes...
[0157] Select frequency point The first road 3D convolution mixed signal As a reference signal This represents the frequency points output after the iteration of the nonnegative matrix factorization iterative model. On Separate the sound source signal. This represents the frequency point recovered after applying the back projection formula. On Separate the sound source signal. The back projection formula is as follows:
[0158]
[0159] Frequency domain separated signal The final time-domain sound source separation signal can be obtained by performing a short-time inverse Fourier transform. .
[0160] Specific Implementation Method Two: Combining Figure 1 As shown, the present invention provides a convolutional blind source separation system, which has a program module corresponding to the above steps, and executes the steps in the above convolutional blind source separation method when running.
[0161] The other combinations and connections in this implementation scheme are the same as in Specific Implementation Scheme 1.
[0162] Specific Implementation Method 3: The present invention provides a computer-readable storage medium storing a computer program configured to implement the steps of a convolutional blind source separation method when called by a processor.
[0163] The other combinations and connections in this implementation scheme are the same as in Specific Implementation Scheme 1.
[0164] Simulation Experiment
[0165] Using Junjie Yang et al.'s "Under-Determined Convolutive Blind Source" published in IEEE Transactions on Civics and Systems – I: Regular Papers (vol. 66, no. 8, pp. 3015-3027, 2019)
[0166] The environmental simulation method in "Separation Combining Density-Based Clustering and Sparse Reconstruction in Time-Frequency Domain" is also an internationally accepted method. It creates a reverberation space with dimensions 1m × 1m × 2m, placing two sound sources at positions [0.5 0 0.36] and [0.5 0.3 0.36], two loudspeakers at positions [0.6 0.05 0.26] and [0.6 0.25 0.26], and two microphones at positions [0.7 0.1 0.16] and [0.7 0.2 0.16]. A schematic diagram of the reverberation space is shown below. Figure 2 As shown, the animal sound signals “hengjiao” and “paoxiao” from dataset C in this literature were selected as the sound source signals.
[0167] Referring to the model parameters of "Blind Signal Separation Combining Impulse Response Reshaping and Expectation Maximization" published by Jie Yuan et al. in the *Journal of Electronics* (2023, vol. 51, no. 11), the parameters of this method are set as follows: sampling frequency , ,in Represents reverberation time. , , , , For STFT, the length of the Hamming window is set to 512, the jump size is set to 20, and the number of Fast Fourier Transform points is set to 512.
[0168] For ease of description, the proposed method is named Impulse Response Compression - Independent Low Rank Matrix Analysis Based on Iterative Source Steering (IRC-ILRMAISS), and compared with Impulse Response Remodeling - Expectation Maximization (IRR-EM) and the original ILRMA algorithm. To ensure fairness, the spatial impulse response length in IRR-EM and ILRMA is consistent with that in IRC-ILRMAISS. Other parameter settings of IRR-EM are referenced from the paper "Blind Signal Separation Combining Impulse Response Remodeling and Expectation Maximization" published in the *Journal of Electronics* (2023, vol. 51, no. 11). For other parameter settings of ILRMA, refer to "Determined Blind SourceSeparation Unifying Independent Vector Analysis and Nonnegative MatrixFactorization" published by D. Kitamura et al. in IEEE / ACM Transactions on Audio, Speech, and Language Processing (2016, vol. 24, no. 9, pp. 1626-1641).
[0169] Using internationally recognized evaluation criteria—Signal-to-distortion ratio (SDR), Signal-to-interference ratio (SIR), and Source-to-artifacts ratio (SAR)—the higher the value, the better the blind source separation performance. Therefore, the magnitude of the SDR, SIR, and SAR values is used to measure the quality of blind source separation.
[0170] Figures 3-5 The figures show the comparison curves of SIR, SDR, and SAR as a function of reverberation time for three convolutional BSS methods. The simulation results at different reverberation times are the average of 50 experiments. It can be seen that the proposed method has higher blind source separation accuracy, especially in high reverberation conditions. The traditional ILRMA performs poorly, while the performance of the IRR-EM algorithm is limited by the setting of key parameters of the impulse response reshaping model.
[0171] Figures 6-8 The comparison curves of SIR, SDR, and SAR of three convolutional BSS methods as a function of signal-to-noise ratio are shown. The simulation results at different signal-to-noise ratios are the average of 50 experiments. It can be seen that the proposed method has better robustness compared to other methods.
[0172] In particular, in some preferred embodiments of the present invention, a computer device is also provided, including a memory and a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the convolutional blind source separation method based on impulse response compression described in any of the above embodiments.
[0173] In some other preferred embodiments of the present invention, a computer-readable storage medium is also provided, on which a computer program / instruction is stored, wherein when the computer program is executed by a processor, the steps of the convolutional blind source separation method based on impulse response compression described in any of the above embodiments are implemented.
[0174] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above embodiments of the convolutional blind source separation method based on impulse response compression, which will not be repeated here.
[0175] Computer-readable storage media encompass a variety of types, including persistent and non-persistent, portable and fixed. These media store information using different technologies, and the content can be machine instructions, data structures, program modules, or other types of data. Some typical examples of computer storage media include: phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), various types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory and other storage technologies, optical storage media such as CD-ROM and digital video disc (DVD), magnetic storage devices such as magnetic tape and disks, and other non-transferable media used to store information accessible to computing devices. It is important to note that the computer-readable media described herein do not include temporary storage media, such as modulated data signals and carrier waves.
[0176] Those skilled in the art will further recognize that the operation of the module can be achieved using existing technical protocols or programs, without relying on new computer programs themselves. The units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0177] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0178] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method of blind source separation based on impulse response compression convolution, characterized in that, The method comprises the following steps: Step 1: constructing a convolution mixed signal model and a global impulse response network, wherein the global impulse response network is realized by installing a microphone before a loudspeaker, and is used to weaken the influence of a high reverberation environment on the performance of the method; Step 2: constructing a target function based on the global impulse response network, wherein the target function is used to adjust the global impulse response by introducing a desired window and an undesired window, so as to minimize the ratio of the undesired impulse response to the desired impulse response; Step 3: updating the impulse response generated between the sound source and the loudspeaker by using a gradient descent algorithm, so as to optimize the target function; Step 4: judging whether the gradient descent iteration model reaches its maximum iteration number; if not, returning to step 3 for reiteration; otherwise, outputting the optimized global impulse response; Step 5: performing short-time Fourier transform on the received signal, converting the signal into a time-frequency domain, constructing a source signal model based on non-negative matrix factorization to simulate the sound source signal, and solving the demixing matrix and the source signal parameter by minimizing a negative log-likelihood function; Step 6: updating the source basis matrix and the source activation matrix, introducing an iterative source control, updating the sound source signal, and normalizing the updated source basis matrix and the sound source signal; Step 7: judging whether the non-negative matrix factorization iteration model reaches its maximum iteration number; if not, returning to step 6 for reiteration; otherwise, outputting the separated signal at each frequency point; Step 8: solving the sorting ambiguity of the separated signal at each frequency point by using a back projection formula, and performing inverse short-time Fourier transform on the frequency-domain separated signal to obtain the recovered time-domain sound source signal.
2. The method according to claim 1, wherein, The convolution mixed signal model constructed in step 1 is as follows: wherein , is the number of samples, is the maximum number of samples, is the length of the impulse response, the additive Gaussian noise in the channel is denoted as ; is the time delay; is the time delay is the aliasing system matrix under is the sound source signal; is the sound source signal after a delay of Pre-installation of a microphone in front of a loudspeaker The overall global impulse response network produces a spatial impulse response of length from the first sound source signal to the first microphone is: wherein the first sound source signal propagating to the first loudspeaker produces a spatial impulse response of length denoted by , the first loudspeaker propagating to the first microphone produces a spatial impulse response of length denoted by , denotes an -dimensional convolution matrix generated by , , constitutes an -dimensional global impulse response matrix , the convolutional mixture signal model based on the global impulse response network is represented as: 3. The convolutional blind source separation method based on impulse response compression according to claim 2, characterized in that, The desired window in step two and the undesired window , wherein , , , , represents the signal sampling frequency, , and represents the window function adjustment factor; the desired impulse response is obtained from the global impulse response , the undesired impulse response wherein represents the dot product; The target function constructed based on the global impulse network is as follows: wherein represents a natural logarithm, and in, This represents taking the maximum value, because , Therefore, the impulse response is not expected. The generated by The function value of the variable is ; u is a symbol representing the undesirable value. For the undesired impulse response The generated by is the function value of the variable; d is the symbol representing the expectation.
4. The method according to claim 3, wherein, The partial derivative of the target function in step 3 is as follows: wherein, wherein, denotes matrix diagonalization, denotes sign function, denotes maximum position index, denotes denotes the element at position index denotes denotes denotes the element at position index denotes denotes the element at position index denotes the element at position index denotes , denotes the row vector of convolution matrix at position index . The updating rule of the gradient descent algorithm is as follows: wherein represents the number of iterations, represents the learning rate, the maximum number of iterations is set to .
5. The method according to claim 4, wherein the impulse response compression is performed by the following equation: ###0001### where x is the input signal, y is the output signal, h is the impulse response, and a is a constant. 5 The convolution mixed signal based on the global impulse response network in step 5 is represented by short-time Fourier transform as follows: in, Represents the short-time Fourier transform at frequency points The following is a mixture matrix, , and These represent the received signal in the time-frequency domain, the sound source signal in the time-frequency domain, and the noise in the time-frequency domain, respectively. and These represent frequency points and time frames, respectively. and These represent the maximum frequency point and the maximum time frame, respectively. Time-frequency domain sound source signal of 3D variance matrix Satisfies nonnegative matrix decomposition. ,Right now In this model, it is assumed that the source signals have zero mean and are independently and identically distributed. In OK Column elements for Expected value; The non-negative matrix factorization model is constructed as follows: in, and Representing the first Time-frequency domain sound source signal variance matrix Decomposition Wikipedia matrix and 3D activation matrix, The cardinality of the nonnegative matrix factorization model; defining frequency points. Below The dimensional unmixing matrix is The recovered frequency domain source signal is: Solving by minimizing the negative log-likelihood function and non-negative matrices , The negative log-likelihood function is expressed as follows: wherein is the de-mixing matrix is the first column vector of the de-mixing matrix and represent the row column elements of the basis matrix and the row column elements of the activation matrix respectively, denotes the determinant of the taking matrix; ; ; ; .
6. The method according to claim 5, wherein the impulse response compression is performed by the following equation: ###0001### where x is the input signal, y is the output signal, h is the impulse response, and a is a constant. 5 Step six initializes the nonnegative matrix. and The elements in the matrix are all random numbers that follow a uniform distribution between [0,1]. This represents the number of iterations; the maximum number of iterations is set to... , and Representing the first basis matrix of OK Column elements and activation matrix of OK Column elements; at frequency points On The 3D separated sound source signal matrix is represented as follows , represent of OK Column elements; , Representing the Variance matrix of OK Column elements; Firstly, the minimization of the negative log-likelihood function is performed by iteratively updating the following space model parameters: Then, the iterative source control algorithm is introduced to update the sound source signal: wherein, represents the updated frequency point on which the dimensional sound source signal matrix, represents the first dimensional sound source signal on the frequency point before updating, and represents the dimensional iteration source control vector, and the element in the vector is updated, when is not equal to When With Equal: wherein represents a conjugate; Finally, the updated source basis matrix and the sound source signals are normalized; representative frequency point on the sound source signal matrix of row column elements, the normalization operation is as follows: 。 7. The method according to claim 6, wherein the impulse response compression is performed by a compression function defined as: ###0002### where x is the input signal, y is the output signal, and a is a constant. Step eight involves selecting a frequency point. The first road 3D convolution mixed signal As a reference signal This represents the frequency points output after the iteration of the nonnegative matrix factorization iterative model. On Separate the sound source signal. This represents the frequency point recovered after applying the back projection formula. On Separate the sound source signal; The back projection is as follows: to the frequency domain separated signal performing an inverse short-time Fourier transform to obtain a final time domain sound source separated signal .
8. A computer apparatus / device / system comprising a memory, a processor, and a computer program stored on the memory, characterized in that: The processor executes the computer program to realize the steps of the method in claim 7.
9. A computer readable storage medium having stored thereon computer programs / instructions, characterized in that: The computer program / instructions are executed by the processor to realize the steps of the method in claim 7.