A single-channel sound separation method
The sound separation network is constructed through the semi-non-negative matrix decomposition algorithm, which solves the problem of target sound signal separation in single-channel mixed sound signals, and realizes the effective separation effect in multiple aliasing situations.
Patent Information
- Application Number
- CN202210805069.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-08
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2042-07-08
AI Technical Summary
The prior art is difficult to effectively separate the target sound from a single-channel mixed sound signal, especially in the case of time domain, frequency domain and time frequency domain aliasing, and existing methods generally require multi-channel signal processing or are limited to specific signal characteristics.
A semi-non-negative matrix decomposition algorithm is used to construct a sound separation network, and a sound separation model is established by training a single-channel target and mixed sound signal, which is used to separate the target acoustic signal from a single-channel mixed sound signal.
Effective target acoustic signal separation in the case of time domain, frequency domain and time frequency domain aliasing, has good ability to extract common characteristics of the same signal, and improves the effect of sound separation.
Smart Images

Figure CN115331693B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the technical field of single-channel mixed sound signal separation, and in particular to a single-channel sound separation method. Background Art
[0002] Acoustic technology has been widely researched and applied in areas such as product quality inspection, equipment status monitoring, medical diagnosis, and the localization of unusual sounds. Because the sound field of its application environment can be complex, the collected sound signal may be a mixture of the target sound and ambient noise. Therefore, it is generally necessary to first separate the target sound from the mixed sound signal for subsequent signal processing and analysis. In addition, due to limitations such as size, equipment cost, and installation issues, it may be necessary (or best practice) to install only one sound sensor. This requires separating the target sound signal from the single-channel mixed sound signal collected by a single microphone. Single-channel mixed sound signal separation is a common method for solving the above tasks.
[0003] The Chinese invention patent "Method and device for separating underdetermined sound signals based on Hilbert transform" applied for by the Third Design Institute of Machinery Industry proposes the use of Hilbert transform for separation of underdetermined sound signals. It does not specify that this is for single-channel mixed sound signals, and its method is not applicable to frequency domain aliasing. The Chinese invention patent "Method for separating multi-vehicle sound signals in wireless sensor networks based on particle filtering" applied for by the Institute of Microsystems of the Jiaxing Center of the Chinese Academy of Sciences and the Chinese invention patent "Method for detecting fault sounds of power equipment based on joint approximate diagonalization blind source separation algorithm" applied for by Shandong University are both for sound separation of multi-channel mixed sound signals collected by microphone arrays. The Chinese invention patent "A method for separating sound signals based on semi-nonnegative matrix decomposition" applied for by the Guangdong Institute of Intelligent Manufacturing is only applicable to the separation of single-channel mixed sound signals with frequency domain overlap.
[0004] The paper "Blind Source Separation of Single-Channel Signals Based on Variational Mode Decomposition" by Wang Kang et al. from Anhui University of Science and Technology proposes a blind source separation method for single-channel signals based on variational mode decomposition. First, variational mode decomposition is used to increase the dimensionality of the single-channel observation signal and estimate the number of source signals. Blind source separation is then performed. This method involves dimensionality increase (i.e., mapping the single-channel signal into a multi-channel signal) and is not suitable for frequency domain aliasing. The paper "Low-Complexity Blind Separation Algorithm for Single-Channel Co-Frequency Mixed Signals Based on SIC" by Guo Yiming et al. from the PLA Information Engineering University uses oversampling to construct multi-channel conditions, then constructs a channel matrix and uses a continuous interference cancellation algorithm to achieve blind separation of single-channel co-frequency mixed signals. This method also involves mapping the single-channel signal into a multi-channel signal before performing signal separation. However, it is significantly affected by delay differences and the oversampling factor of the received signal, resulting in demodulation blind spots. The paper "Blind Source Separation Algorithm for Single-Channel Communication Signals" by Yang Hailan et al. from Jiangnan University proposes a blind source separation algorithm for single-channel communication signals based on the Hilbert-Huang transform and independent component analysis. This method is not suitable for frequency domain aliasing. Zhu Huijie et al. from the PLA University of Science and Technology published a paper titled "Single-Channel Blind Source Separation of Mechanical Signals Based on Shift-Invariant Sparse Coding." They proposed a single-channel blind source separation method based on shift-invariant sparse coding for mechanical signals with recurring features. The algorithm treats the source signal as a convolution of multiple bases and coefficients. Based on the signal's statistical distribution and inherent characteristics, it adaptively learns matching bases and sparse coefficients. This method is specifically targeted at mechanical signals with recurring features. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a single-channel sound separation method in view of the deficiencies in the prior art.
[0006] The present invention solves the above technical problems with the following technical solutions: A single-channel sound separation method comprises the following steps:
[0007] S1: importing a first single-channel target sound signal x and a first single-channel interference sound signal y, and mixing the first single-channel target sound signal x and the first single-channel interference sound signal y to obtain a first single-channel mixed sound signal s;
[0008] S2: constructing a sound separation network W based on a semi-nonnegative matrix factorization algorithm, and training the sound separation network W using the first single-channel target sound signal x and the first single-channel mixed sound signal s to obtain a sound separation model M;
[0009] S3: Input the second single-channel mixed sound signal s′ into the sound separation model M for separation to obtain the second target sound signal
[0010] Another technical solution of the present invention to solve the above technical problem is as follows: a single-channel sound separation device, comprising:
[0011] a signal mixing module, configured to import a first single-channel target acoustic signal x and a first single-channel interference acoustic signal y, and mix the first single-channel target acoustic signal x and the first single-channel interference acoustic signal y to obtain a first single-channel mixed acoustic signal s;
[0012] A model training module is configured to construct a sound separation network W based on a semi-nonnegative matrix factorization algorithm, and train the sound separation network W using the first single-channel target sound signal x and the first single-channel mixed sound signal s to obtain a sound separation model M;
[0013] The target sound signal acquisition module is used to input the second single-channel mixed sound signal s' into the sound separation model M for separation to obtain the second target sound signal
[0014] Another technical solution of the present invention to solve the above-mentioned technical problem is as follows: a single-channel sound separation device includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the single-channel sound separation method as described above is implemented.
[0015] Another technical solution of the present invention to solve the above technical problem is as follows: a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the single-channel sound separation method as described above is implemented.
[0016] The beneficial effects of the present invention are: a first single-channel mixed sound signal is obtained by mixing a first single-channel target sound signal and a single-channel interference sound signal, a sound separation network is obtained based on a semi-non-negative matrix decomposition algorithm, a sound separation model is obtained by training the sound separation network with the first single-channel target sound signal and the first single-channel mixed sound signal, the second single-channel mixed sound signal is input into the sound separation model for separation to obtain a target sound signal, and the sound separation model can be automatically learned from the target sound and the mixed sound. The model can be used to separate the target sound from the mixed sound of time domain aliasing, frequency domain aliasing, and time-frequency domain aliasing. In addition, semi-non-negative matrix decomposition has the advantage of better extracting common features of the same signal. Constructing a sound separation network based on semi-non-negative matrix decomposition can enable the network to have a better ability to extract the target sound signal components, thereby better reconstructing the target sound and achieving a better sound separation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A schematic flow chart of a single-channel sound separation method provided by an embodiment of the present invention;
[0018] Figure 2 A schematic diagram of a flow chart of a sound separation network W provided in an embodiment of the present invention;
[0019] Figure 3 This is a module block diagram of a single-channel sound separation device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0021] Figure 1 A schematic flow chart of a single-channel sound separation method provided in an embodiment of the present invention.
[0022] like Figure 1 As shown, a single-channel sound separation method includes the following steps:
[0023] S1: importing a first single-channel target sound signal x and a first single-channel interference sound signal y, and mixing the first single-channel target sound signal x and the first single-channel interference sound signal y to obtain a first single-channel mixed sound signal s;
[0024] S2: constructing a sound separation network W based on a semi-nonnegative matrix factorization algorithm, and training the sound separation network W using the first single-channel target sound signal x and the first single-channel mixed sound signal s to obtain a sound separation model M;
[0025] S3: Input the second single-channel mixed sound signal s′ into the sound separation model M for separation to obtain the second target sound signal
[0026] It should be understood that the signal (i.e., the target acoustic signal ) represents an estimation of the signal x (ie, the single-channel original acoustic signal x), and the two have a high similarity.
[0027] It should be understood that the first single-channel interference sound signal y is a single-channel interference sound signal y 1 、y 2 ,…,y n A mixture of any one or more or all of the signals;
[0028] The second single-channel mixed sound signal s' is a mixed signal of the first single-channel target sound signal x and the second single-channel interference sound signal y', and the second single-channel interference sound signal y' is a mixed signal of the single-channel interference sound signal y 1 、y 2 ,…,yn A mixture of any one or more or all of the signals;
[0029] The single-channel interference sound signal y 1 、y 2 ,…,y n is n (n≥1) types of interference sound signals;
[0030] The first single-channel target sound signal x, the first single-channel interference sound signal y, the single-channel interference sound signal y 1 、y 2 ,…,y n , the first single-channel mixed sound signal s, the second single-channel mixed sound signal s′ and the second target sound signal The size of is 1×n1.
[0031] In the above embodiment, a first single-channel mixed sound signal is obtained by mixing a first single-channel target sound signal and a single-channel interference sound signal, a sound separation network is obtained based on a semi-non-negative matrix decomposition algorithm, a sound separation model is obtained by training the sound separation network with the first single-channel target sound signal and the first single-channel mixed sound signal, and the second single-channel mixed sound signal is input into the sound separation model for separation to obtain a target sound signal. The sound separation model can be automatically learned from the target sound and the mixed sound, and the model can be used to separate the target sound from the mixed sound of time domain aliasing, frequency domain aliasing, and time-frequency domain aliasing. In addition, semi-non-negative matrix decomposition has the advantage of better extracting common features of the same signal. Constructing a sound separation network based on semi-non-negative matrix decomposition can enable the network to have a better ability to extract the target sound signal components, thereby better reconstructing the target sound and achieving a better sound separation effect.
[0032] Optionally, as an embodiment of the present invention, Figure 1 and 2 As shown, the sound separation network W includes an input layer, a Fourier transform layer, an energy spectrum separation layer, a one-dimensional convolution unit C1, a one-dimensional convolution unit C2 and a target sound reconstruction layer.
[0033] The process in step S2 includes:
[0034] S21: the input layer inputs the first single-channel mixed sound signal s;
[0035] S22: The Fourier transform layer performs a convolution operation on the first single-channel mixed sound signal s and the one-dimensional convolution unit C1 to obtain a real part R of the Fourier transform;
[0036] Performing a convolution operation on the first single-channel mixed sound signal s and the one-dimensional convolution unit C2 to obtain an imaginary part I of a Fourier transform;
[0037] Calculating the real part R of the Fourier transform and the imaginary part I of the Fourier transform to obtain a Fourier transform energy spectrum E and a Fourier transform phase spectrum P;
[0038] S23: The energy spectrum separation layer performs semi-nonnegative matrix decomposition and partial reconstruction on the Fourier transform energy spectrum E to obtain an energy spectrum E1;
[0039] S24: The target acoustic reconstruction layer calculates the energy spectrum E1 and the Fourier transform phase spectrum P to obtain a real part R1 and an imaginary part I1;
[0040] Combining the real part R1 and the imaginary part I1 to obtain a complex number X1;
[0041] Performing an inverse Fourier transform on the complex number X1 to obtain a complex number x1;
[0042] Extracting the real part of the complex number x1 to obtain the real part x2;
[0043] Extract the first n1 elements of the real part x2 as the third target acoustic signal
[0044] The real part R of the Fourier transform, the imaginary part I of the Fourier transform, the Fourier transform energy spectrum E, the Fourier transform phase spectrum P, the energy spectrum E1, the real part R1, the imaginary part I1, the complex number X1, the complex number x1 and the real part x2 are all 1×n2, and n2≥n1;
[0045] S25: The first single-channel target acoustic signal x and the third target acoustic signal x are calculated by the first formula. Calculate the loss function L of the sound separation network W to obtain the loss function L, where the first formula is:
[0046]
[0047] Where L is the loss function, x is the first single-channel target sound signal, is the third target sound signal, ∑x 2 To square and then sum the elements in the first single-channel target sound signal x, For The elements in are first squared and then summed, log 10 To calculate the base 10 logarithm;
[0048] S26: Update the parameters of the sound separation network W according to the loss function L to obtain a sound separation model M.
[0049] Specifically, if Figure 2As shown, the specific structure and application steps of the sound separation network W based on semi-nonnegative matrix decomposition (i.e., the sound separation network W) are as follows:
[0050] S21, input layer: input the first single-channel mixed sound signal s;
[0051] S22, Fourier transform layer: convolve s (i.e., the first single-channel mixed sound signal s) with the one-dimensional convolution unit C1 to obtain the real part R of the Fourier transform; convolve s (i.e., the first single-channel mixed sound signal s) with the one-dimensional convolution unit C2 to obtain the imaginary part I of the Fourier transform; calculate the Fourier transform energy spectrum E and the Fourier transform phase spectrum P based on the real part R (i.e., the real part R of the Fourier transform) and the imaginary part I (i.e., the imaginary part I of the Fourier transform);
[0052] S23, energy spectrum separation layer: performing semi-nonnegative matrix decomposition and partial reconstruction on the energy spectrum E (i.e., the Fourier transform energy spectrum E) to obtain an energy spectrum E1;
[0053] S24, target sound reconstruction layer: using the energy spectrum E1 and the phase spectrum P (i.e., the Fourier transform phase spectrum P), calculate the real part R1 and the imaginary part I1; combine the real part R1 and the imaginary part I1 into a complex number X1, perform an inverse Fourier transform on X1 (i.e., the complex number X1) to obtain a complex number x1, take the real part of x1 (i.e., the complex number x1) to obtain x2 (i.e., the complex number x2), and take the first n1 elements of x2 (i.e., the complex number x2) as the separated target sound signal (i.e., the third target acoustic signal ).
[0054] In this embodiment, the sizes of R (i.e., the real part R of the Fourier transform), I (i.e., the imaginary part I of the Fourier transform), E (i.e., the Fourier transform energy spectrum E), P (i.e., the Fourier transform phase spectrum P), E1 (i.e., the energy spectrum E1), R1 (i.e., the real part R1), I1 (i.e., the imaginary part I1), X1 (i.e., the complex number X1), x1 (i.e., the complex number x1), and x2 (i.e., the complex number x2) are all 1×n2, and n2≥n1.
[0055] In this embodiment, the loss function L used for training the sound separation network W based on semi-nonnegative matrix factorization (i.e., the sound separation network W) is:
[0056]
[0057] Among them, ∑x 2 It means that the elements in x are squared and then summed. Express The elements in are first squared and then summed, log10 Calculates the base 10 logarithm.
[0058] In the above embodiment, a sound separation network is constructed by a semi-non-negative matrix decomposition algorithm, and a sound separation model is obtained by training the sound separation network using a first single-channel target sound signal and a first single-channel mixed sound signal. This model has the advantage of better extracting common features of the same signal. Constructing a sound separation network based on semi-non-negative matrix decomposition allows the network to have a better ability to extract target sound signal components, thereby better reconstructing the target sound and achieving a better sound separation effect.
[0059] Optionally, as an embodiment of the present invention, the size of the convolution kernel of the one-dimensional convolution unit C1 is n2×n1, and the initial value of the convolution kernel of the one-dimensional convolution unit C1 is calculated by the second formula to obtain the initial value C1_R of the convolution kernel, and the second formula is:
[0060]
[0061] Among them, C1_R is the initial value of the convolution kernel of the one-dimensional convolution unit C1;
[0062] The size of the convolution kernel of the one-dimensional convolution unit C2 is n2×n1. The initial value of the convolution kernel of the one-dimensional convolution unit C2 is calculated by the third formula to obtain the initial value C2_R of the convolution kernel. The third formula is:
[0063]
[0064] Among them, C2_R is the initial value of the convolution kernel of the one-dimensional convolution unit C2;
[0065] The process of calculating the real part R of the Fourier transform and the imaginary part I of the Fourier transform to obtain the Fourier transform energy spectrum E and the Fourier transform phase spectrum P is specifically as follows:
[0066] The Fourier transform energy spectrum E is calculated by performing Fourier transform on the real part R of the Fourier transform and the imaginary part I of the Fourier transform through the fourth formula to obtain the Fourier transform energy spectrum E. The fourth formula is:
[0067]
[0068] Where E is the Fourier transform energy spectrum, R is the real part of the Fourier transform, and I is the imaginary part of the Fourier transform;
[0069] The Fourier transform phase spectrum P is calculated by performing Fourier transform on the real part R of the Fourier transform and the imaginary part I of the Fourier transform through the fifth formula to obtain the Fourier transform phase spectrum P. The fifth formula is:
[0070]
[0071] Where P is the Fourier transform phase spectrum, R is the real part of the Fourier transform, and I is the imaginary part of the Fourier transform.
[0072] Specifically, the size of the convolution kernel C1_R of the one-dimensional convolution unit C1 in step S22 is n2×n1. During the weight initialization process of the network W (i.e., the sound separation network W), the initial value of C1_R is the cosine triangular basis function set of the Fourier transform. The initial value of C1_R is:
[0073]
[0074] The size of the convolution kernel C2_R of the one-dimensional convolution unit C2 in step S22 is n2×n1. During the weight initialization process of the network W, the initial value of C2_R is the sine triangular basis function set of the Fourier transform. The initial value of C2_R is:
[0075]
[0076] The calculation formula of the Fourier transform energy spectrum E in step S22 is:
[0077]
[0078] The calculation formula of the Fourier transform phase spectrum P in step S22 is:
[0079]
[0080] In the above embodiment, the Fourier transform energy spectrum and the Fourier transform phase spectrum are obtained by calculating the real part of the Fourier transform and the imaginary part of the Fourier transform, and a sound separation model can be automatically learned from the target sound and the mixed sound. The model can be used to separate the target sound from the mixed sound of time domain aliasing, frequency domain aliasing, and time-frequency domain aliasing.
[0081] Optionally, as an embodiment of the present invention, the process of S23 includes:
[0082] S231: Initializing the sound separation network W to obtain weights F1 and F2, wherein the initial values of the weights F1 and F2 are both within the range of (-1, 1);
[0083] Initializing the coefficient matrix of the semi-nonnegative matrix decomposition algorithm to obtain a coefficient matrix G, wherein the initial value of the coefficient matrix G is within the interval (0, 1);
[0084] S232: Add the weight F1 and the weight F2 to obtain a basis matrix F of semi-non-negative matrix decomposition;
[0085] The coefficient matrix G is iteratively updated N times by using the sixth formula and the basis matrix F of the semi-nonnegative matrix decomposition and the Fourier transform energy spectrum E to obtain an updated coefficient matrix G. The sixth formula is:
[0086]
[0087] Among them, G is the coefficient matrix of semi-non-negative matrix decomposition, which is a non-negative matrix, E is the Fourier transform energy spectrum, F is the basis matrix of semi-non-negative matrix decomposition, F T is the transpose of the basis matrix F of the semi-nonnegative matrix factorization, (EF T ) + EF T The positive elements in (EF T ) - EF T Negative elements in (FF T ) + For FF T The positive elements in (FF T ) - For FF T Negative elements in ;
[0088] S233: Calculate the energy spectrum of the updated coefficient matrix G and the weight F1 using the seventh formula to obtain the energy spectrum E1. The seventh formula is:
[0089] E1=G*F1,
[0090] Among them, E1 is the energy spectrum, G is the updated coefficient matrix, and F1 is the weight.
[0091] It should be understood that the value of N in step S232 can be an integer between 60 and 120, specifically 100.
[0092] Specifically, S231, initialize the weights F1 and F2 of the sound separation network W, where the initial values of F1 (i.e., the weight F1) and F2 (i.e., the weight F2) are in the range of (-1, 1); initialize the coefficient matrix G of the semi-non-negative matrix decomposition, where the initial value of G is in the range of (0, 1);
[0093] S232, add F1 (i.e., the weight F1) and F2 (i.e., the weight F2) to obtain the basis matrix F of the semi-nonnegative matrix decomposition, using the formula Performing N iterative updates on the coefficient matrix G;
[0094] In the above formula, G is the coefficient matrix of semi-non-negative matrix decomposition, which is a non-negative matrix, E is the Fourier transform energy spectrum, F is the basis matrix of semi-non-negative matrix decomposition, F T is the transpose of F; (EFT ) + EF T The positive elements in (EF T ) - EF T Negative elements in (FF T ) + For FF T The positive elements in (FF T ) - For FF T Negative elements in .
[0095] S233 , multiply the coefficient matrix G (ie, the updated coefficient matrix G) by F1 (ie, the weight F1) to obtain the energy spectrum E1, ie, E1=G*F1.
[0096] In the above embodiment, by obtaining the energy spectrum through semi-nonnegative matrix decomposition and partial reconstruction of the Fourier transform energy spectrum, the network can have a better ability to extract the target sound signal components, thereby better reconstructing the target sound and achieving a better sound separation effect.
[0097] Optionally, as an embodiment of the present invention, in step S24, the process of calculating the energy spectrum E1 and the Fourier transform phase spectrum P to obtain a real part R1 and an imaginary part I1; and combining the real part R1 and the imaginary part I1 to obtain a complex number X1 includes:
[0098] The real part R1 is calculated for the energy spectrum E1 and the Fourier transform phase spectrum P using the eighth formula to obtain the real part R1. The eighth formula is:
[0099] R1=E1*cos(P),
[0100] Where R1 is the real part, E1 is the energy spectrum, and P is the Fourier transform phase spectrum;
[0101] The imaginary part I1 is calculated for the energy spectrum E1 and the Fourier transform phase spectrum P using the ninth formula to obtain the imaginary part I1. The ninth formula is:
[0102] I1=E1*sin(P),
[0103] Where I1 is the imaginary part, E1 is the energy spectrum, and P is the Fourier transform phase spectrum;
[0104] The real part R1 and the imaginary part I1 are combined by the tenth formula to obtain the complex number X1. The tenth formula is:
[0105] X1=R1+j*I1,
[0106] Where X1 is a complex number, R1 is the real part, I1 is the imaginary part, and j is the imaginary number sign.
[0107] It should be understood that the calculation formula of the real part R1 in step S24 is:
[0108] R1=E1*cos(P),
[0109] The calculation formula of the imaginary part I1 in step S24 is:
[0110] I1=E1*sin(P),
[0111] The calculation formula for combining the real part R1 and the imaginary part I1 into the complex number X1 in step S24 is:
[0112] X1=R1+j*I1,
[0113] Wherein, j is the imaginary number symbol.
[0114] In the above embodiment, the real part and the imaginary part are obtained by calculating the energy spectrum and the Fourier transform phase spectrum, and the complex number is obtained by combining the real part and the imaginary part, which has the advantage of better extracting the common characteristics of the same signal. The sound separation network is constructed based on semi-non-negative matrix decomposition, which enables the network to have a better ability to extract the target sound signal components, thereby better reconstructing the target sound and achieving a better sound separation effect.
[0115] Optionally, as an embodiment of the present invention, in the sound separation network W, the trained network weights include: the convolution kernel of the one-dimensional convolution unit C1, the convolution kernel of the one-dimensional convolution unit C2, the weight F1 and the weight F2.
[0116] It should be understood that in the sound separation network W, the network weights that need to be trained include: the convolution kernel C1_R of the one-dimensional convolution unit C1 (that is, the convolution kernel of the one-dimensional convolution unit C1), the convolution kernel C2_R of the one-dimensional convolution unit C2 (that is, the convolution kernel of the one-dimensional convolution unit C2), the weight F1 and the weight F2.
[0117] In the above embodiment, the network weights are trained to obtain a better network model and achieve better sound separation effect.
[0118] Optionally, as another embodiment of the present invention, the sound separation model M and the sound separation network W have the same network structure.
[0119] Optionally, as another embodiment of the present invention, the beneficial effects of the present invention are as follows:
[0120] The present invention utilizes a sound separation network based on semi-nonnegative matrix factorization to automatically learn a sound separation model from a target sound and mixed sounds. This model can be used to separate the target sound from mixed sounds with time domain aliasing, frequency domain aliasing, and time-frequency domain aliasing. Furthermore, semi-nonnegative matrix factorization is well-suited for extracting common features of homogeneous signals. Constructing a sound separation network based on semi-nonnegative matrix factorization allows the network to better extract target sound signal components, thereby better reconstructing the target sound and achieving superior sound separation results.
[0121] Alternatively, as another embodiment of the present invention, the present invention is used to separate a target sound from a single-channel mixed sound. This method employs an AI-based sound separation approach, using the target sound and a mixed sound of the target sound and interference sound to train a sound separation network based on semi-nonnegative matrix factorization, thereby obtaining a sound separation model.
[0122] Optionally, as another embodiment of the present invention, the present invention discloses a single-channel sound separation method, comprising: mixing a single-channel target sound signal and a single-channel interfering sound signal to obtain a single-channel mixed sound signal; using the target sound signal and the mixed sound signal to train a sound separation network based on semi-non-negative matrix decomposition to obtain a sound separation model; and inputting the single-channel mixed sound signal into the sound separation model to separate the target sound signal. The present invention utilizes a sound separation network based on semi-non-negative matrix decomposition to automatically learn a sound separation model from the target sound and the mixed sound. The model can be used to separate the target sound from mixed sounds with time domain aliasing, frequency domain aliasing, or time-frequency domain aliasing.
[0123] Optionally, as another embodiment of the present invention, the effect of the present invention can be further illustrated by the following experiment:
[0124] 1) Experimental data
[0125] The experimental data contains three sounds: the sound of an air compressor (denoted as sound A for convenience), the sound of a running conveyor belt (denoted as sound B for convenience), and the human voice (denoted as sound C for convenience). Sound A has 100 training samples and 10 test samples; sound B has 100 training samples and 10 test samples; and sound C has 10 samples. Each sample is 0.8 seconds long, and the sampling rate is 48 kHz. Sound A is assumed to be the target sound, and sounds B and C are interference sounds. Most of the main frequency bands of sounds A, B, and C overlap.
[0126] Each training sample of sound A is mixed with each training sample of sound B in pairs, and 10,000 mixed sound samples (denoted as Tr_AB sound for convenience) to be used for training can be obtained.
[0127] Each test sample of sound A is mixed with each test sample of sound B in pairs, and 100 mixed sound samples (for the convenience of description, denoted as Te_AB sound) to be used for testing can be obtained.
[0128] Each test sample of the A sound is mixed with each sample of the C sound in pairs, and 100 mixed sound samples to be used for testing can be obtained (for convenience of expression, recorded as Te_AC sound).
[0129] 2) Experimental conditions
[0130] The experimental program of the present invention was written using Python 3.6.5 software, and the code related to the sound separation network was written based on TensorFlow 1.15.3. The value of N in S232 was 100. The sound separation network was trained using the Tr_AB sound pair to obtain a sound separation model. The sound separation model was tested using the Te_AB sound pair and the Te_AC sound pair, and the separation effect was evaluated.
[0131] 3) Experimental results
[0132] The signal distortion ratio (SDR) is used as an evaluation index of the separation effect of the present invention, and the unit of SDR is db.
[0133] Before signal separation, the SDR (denoted as SDR) of the original target sound signal (A sound) and the mixed sound signal (Te_AB sound or Te_AC sound) is calculated. 1 ) is:
[0134]
[0135] After signal separation, the SDR (denoted as SDR) of the original target sound signal (A sound) and the separated target sound signal (the target sound separated from the Te_AB sound or Te_AC sound) is calculated. 2 ) is:
[0136]
[0137] In the above two formulas, x represents the original target sound signal, that is, A sound, s′ is the mixed sound signal of signal x and interference sound signal, that is, Te_AB sound and Te_AC sound, It represents the target sound signal separated from the mixed sound using the sound separation model. 2 With SDR 1 The greater the difference, the better the separation effect; on the contrary, if SDR 2 <SDR 1 , which means that after separation, the signal distortion is higher, that is, the separation model has a negative effect. The experimental results are shown in Table 1.
[0138] Table 1 shows the average SDR of the test mixed sound 1 , average SDR 2 Comparison with average SDR improvement
[0139]
[0140]
[0141] Average SDR in Table 1 1 Indicates the average SDR of 100 mixed sound samples, average SDR 2 It represents the average SDR of 100 separated target sounds. The average SDR improvement represents the improvement of SDR before and after separation, that is, SDR 2 -SDR 1 From the average SDR improvement values in Table 1, it can be seen that the present invention has a better sound separation effect.
[0142] Figure 3 This is a module block diagram of a single-channel sound separation device provided by an embodiment of the present invention.
[0143] Alternatively, as another embodiment of the present invention, Figure 3 As shown, a single-channel sound separation device includes:
[0144] a signal mixing module, configured to import a single-channel original acoustic signal x and a single-channel interference acoustic signal y, and mix the first single-channel target acoustic signal x and the single-channel interference acoustic signal y to obtain a first single-channel mixed acoustic signal s;
[0145] a model training module, configured to obtain a sound separation network W based on a semi-nonnegative matrix factorization algorithm, and train the sound separation network W according to the first single-channel target sound signal x and the first single-channel mixed sound signal s to obtain a sound separation model M;
[0146] The target sound signal acquisition module is used to input the first single-channel mixed sound signal s' into the sound separation model M for separation to obtain the target sound signal
[0147] Optionally, as an embodiment of the present invention, the sound separation network W includes an input layer, a Fourier transform layer, an energy spectrum separation layer, a one-dimensional convolution unit C1, a one-dimensional convolution unit C2 and a target sound reconstruction layer.
[0148] The model training module is specifically used for:
[0149] The input layer inputs the first single-channel mixed sound signal s;
[0150] The Fourier transform layer performs a convolution operation on the first single-channel mixed sound signal s and the one-dimensional convolution unit C1 to obtain a real part R of the Fourier transform;
[0151] Performing a convolution operation on the first single-channel mixed sound signal s and the one-dimensional convolution unit C2 to obtain an imaginary part I of a Fourier transform;
[0152] Calculating the real part R of the Fourier transform and the imaginary part I of the Fourier transform to obtain a Fourier transform energy spectrum E and a Fourier transform phase spectrum P;
[0153] The energy spectrum separation layer performs semi-nonnegative matrix decomposition and partial reconstruction on the Fourier transform energy spectrum E to obtain an energy spectrum E1;
[0154] The target acoustic reconstruction layer calculates the energy spectrum E1 and the Fourier transform phase spectrum P to obtain a real part R1 and an imaginary part I1;
[0155] Combining the real part R1 and the imaginary part I1 to obtain a complex number X1;
[0156] Performing an inverse Fourier transform on the complex number X1 to obtain a complex number x1;
[0157] Extracting the real part of the complex number x1 to obtain the real part x2;
[0158] Extract the first n1 elements of the real part x2 as the third target acoustic signal
[0159] The real part R of the Fourier transform, the imaginary part I of the Fourier transform, the Fourier transform energy spectrum E, the Fourier transform phase spectrum P, the energy spectrum E1, the real part R1, the imaginary part I1, the complex number X1, the complex number x1 and the real part x2 are all 1×n2, and n2≥n1;
[0160] The first single-channel target acoustic signal x and the third target acoustic signal x are calculated by the first formula Calculate the loss function L of the sound separation network W to obtain the loss function L, where the first formula is:
[0161]
[0162] Where L is the loss function, x is the first single-channel target sound signal, is the third target sound signal, ∑x 2 To square and then sum the elements in the first single-channel target sound signal x, For The elements in are first squared and then summed, log 10To calculate the base 10 logarithm;
[0163] The parameters of the sound separation network W are updated according to the loss function L to obtain a sound separation model M.
[0164] Alternatively, another embodiment of the present invention provides a single-channel sound separation device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the single-channel sound separation method described above is implemented. The device may be a computer or other device.
[0165] Optionally, another embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the single-channel sound separation method as described above is implemented.
[0166] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0167] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0168] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or ignoring or not implementing certain features.
[0169] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the objectives of the embodiments of the present invention.
[0170] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0171] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0172] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A single-channel sound separation method, characterized in that: The steps include: S1: importing a first single-channel target sound signal x and a first single-channel interference sound signal y, and mixing the first single-channel target sound signal x and the first single-channel interference sound signal y to obtain a first single-channel mixed sound signal s; S2: constructing a sound separation network W based on a semi-nonnegative matrix factorization algorithm, and training the sound separation network W using the first single-channel target sound signal x and the first single-channel mixed sound signal s to obtain a sound separation model M; S3: Input the second single-channel mixed sound signal s′ into the sound separation model M for separation to obtain the second target sound signal The sound separation network W includes an input layer, a Fourier transform layer, an energy spectrum separation layer, a one-dimensional convolution unit C1, a one-dimensional convolution unit C2 and a target sound reconstruction layer. The process in step S2 includes: S21: the input layer inputs the first single-channel mixed sound signal s; S22: The Fourier transform layer performs a convolution operation on the first single-channel mixed sound signal s and the one-dimensional convolution unit C1 to obtain a real part R of the Fourier transform; Performing a convolution operation on the first single-channel mixed sound signal s and the one-dimensional convolution unit C2 to obtain an imaginary part I of a Fourier transform; Calculating the real part R of the Fourier transform and the imaginary part I of the Fourier transform to obtain a Fourier transform energy spectrum E and a Fourier transform phase spectrum P; S23: The energy spectrum separation layer performs semi-nonnegative matrix decomposition and partial reconstruction on the Fourier transform energy spectrum E to obtain an energy spectrum E1; S24: The target acoustic reconstruction layer calculates the energy spectrum E1 and the Fourier transform phase spectrum P to obtain a real part R1 and an imaginary part I1; Combining the real part R1 and the imaginary part I1 to obtain a complex number X1; Performing an inverse Fourier transform on the complex number X1 to obtain a complex number x1; Extracting the real part of the complex number x1 to obtain the real part x2; Extract the first n1 elements of the real part x2 as the third target acoustic signal The real part R of the Fourier transform, the imaginary part I of the Fourier transform, the Fourier transform energy spectrum E, the Fourier transform phase spectrum P, the energy spectrum E1, the real part R1, the imaginary part I1, the complex number X1, the complex number x1 and the real part x2 are all 1×n2, and n2≥n1; S25: The first single-channel target acoustic signal x and the third target acoustic signal x are calculated by the first formula. Calculate the loss function L of the sound separation network W to obtain the loss function L, where the first formula is: Where L is the loss function, x is the first single-channel target sound signal, is the third target sound signal, ∑x 2 To square and then sum the elements in the first single-channel target sound signal x, For The elements in are first squared and then summed, log 10 To calculate the base 10 logarithm; S26: Update the parameters of the sound separation network W according to the loss function L to obtain a sound separation model M.
2. The single-channel sound separation method according to claim 1, characterized in that The size of the convolution kernel of the one-dimensional convolution unit C1 is n2×n1. The initial value of the convolution kernel of the one-dimensional convolution unit C1 is calculated by the second formula to obtain the initial value C1_R of the convolution kernel. The second formula is: Where C1_R is the initial value of the convolution kernel of the one-dimensional convolution unit C1; The size of the convolution kernel of the one-dimensional convolution unit C2 is n2×n1. The initial value of the convolution kernel of the one-dimensional convolution unit C2 is calculated by the third formula to obtain the initial value C2_R of the convolution kernel. The third formula is: Among them, C2_R is the initial value of the convolution kernel of the one-dimensional convolution unit C2; The process of calculating the real part R of the Fourier transform and the imaginary part I of the Fourier transform to obtain the Fourier transform energy spectrum E and the Fourier transform phase spectrum P is specifically as follows: The Fourier transform energy spectrum E is calculated by performing Fourier transform on the real part R of the Fourier transform and the imaginary part I of the Fourier transform through the fourth formula to obtain the Fourier transform energy spectrum E. The fourth formula is: Where E is the Fourier transform energy spectrum, R is the real part of the Fourier transform, and I is the imaginary part of the Fourier transform; The Fourier transform phase spectrum P is calculated by performing Fourier transform on the real part R of the Fourier transform and the imaginary part I of the Fourier transform through the fifth formula to obtain the Fourier transform phase spectrum P. The fifth formula is: Where P is the Fourier transform phase spectrum, R is the real part of the Fourier transform, and I is the imaginary part of the Fourier transform.
3. The single-channel sound separation method according to claim 2, characterized in that The process of S23 includes: S231: Initializing the sound separation network W to obtain weights F1 and F2, wherein the initial values of the weights F1 and F2 are both within the range of (-1, 1); Initializing the coefficient matrix of the semi-nonnegative matrix decomposition algorithm to obtain a coefficient matrix G, wherein the initial value of the coefficient matrix G is within the interval (0, 1); S232: Add the weight F1 and the weight F2 to obtain a basis matrix F of semi-non-negative matrix decomposition; The coefficient matrix G is iteratively updated N times by using the sixth formula and the basis matrix F of the semi-nonnegative matrix decomposition and the Fourier transform energy spectrum E to obtain an updated coefficient matrix G. The sixth formula is: Among them, G is the coefficient matrix of semi-non-negative matrix decomposition, which is a non-negative matrix, E is the Fourier transform energy spectrum, F is the basis matrix of semi-non-negative matrix decomposition, F T is the transpose of the basis matrix F of the semi-nonnegative matrix factorization, (EF T ) + EF T The positive elements in (EF T ) - EF T Negative elements in (FF T ) + For FF T The positive elements in (FF T ) - For FF T Negative elements in ; S233: Calculate the energy spectrum of the updated coefficient matrix G and the weight F1 using the seventh formula to obtain the energy spectrum E1. The seventh formula is: E1=G*F1, Among them, E1 is the energy spectrum, G is the updated coefficient matrix, and F1 is the weight.
4. The single-channel sound separation method according to claim 1, characterized in that In step S24, the process of calculating the energy spectrum E1 and the Fourier transform phase spectrum P to obtain a real part R1 and an imaginary part I1; and combining the real part R1 and the imaginary part I1 to obtain a complex number X1 includes: The real part R1 is calculated for the energy spectrum E1 and the Fourier transform phase spectrum P using the eighth formula to obtain the real part R1. The eighth formula is: R1=E1*cos(P), Where R1 is the real part, E1 is the energy spectrum, and P is the Fourier transform phase spectrum; The imaginary part I1 is calculated for the energy spectrum E1 and the Fourier transform phase spectrum P using the ninth formula to obtain the imaginary part I1. The ninth formula is: I1=E1*sin(P), Where I1 is the imaginary part, E1 is the energy spectrum, and P is the Fourier transform phase spectrum; The real part R1 and the imaginary part I1 are combined by the tenth formula to obtain the complex number X1. The tenth formula is: X1=R1+j*I1, Where X1 is a complex number, R1 is the real part, I1 is the imaginary part, and j is the imaginary number sign.
5. The single-channel sound separation method according to claim 3, characterized in that: In the sound separation network W, the trained network weights include: the convolution kernel of the one-dimensional convolution unit C1, the convolution kernel of the one-dimensional convolution unit C2, the weight F1 and the weight F2.
Citation Information
Patent Citations
Speaker-independent single-channel voice separation method
CN111583954A
Semi-non-negative matrix factorization-based sound signal separation method
CN111837119A