A single-channel speech enhancement method for fiber microphones integrating OMLSA and TQWT
By combining OMLSA and TQWT to develop a single-channel speech enhancement method for fiber optic microphones, the noise suppression problem of fiber optic microphones in low signal-to-noise ratio environments is solved, and efficient enhancement of speech signals and quality fidelity are achieved.
Patent Information
- Application Number
- CN202411871991.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-12-18
AI Technical Summary
The existing single-channel speech enhancement algorithm has poor noise suppression effect in fiber optic microphone applications, especially in low signal-to-noise ratio conditions where there is a lot of residual noise. In addition, improper TQWT parameter settings affect the enhancement effect and make it difficult to effectively suppress low-frequency noise.
The method of fusing OMLSA and TQWT is adopted. After preliminary enhancement by the OMLSA algorithm, the segmented signal-to-noise ratio is used to determine whether further TQWT processing is needed. The TQWT parameters are optimized with the basis pursuit sparsification technology to achieve efficient enhancement of the speech signal.
Under the premise of ensuring voice quality, it effectively reduces residual noise, improves the signal-to-noise ratio, saves computing resources, and achieves efficient enhancement of the fiber optic microphone signal.
Smart Images

Figure CN119785808B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microphone speech signal processing, and in particular to a single-channel speech enhancement method for an optical fiber microphone integrating OMLSA and TQWT. Background Art
[0002] Fiber optic microphones are acoustic sensors that use optical fiber as a sensing medium. They detect phase changes in the fiber caused by sound waves and then recover the sound wave information. They are passive, resistant to electromagnetic interference, and have a wide dynamic range. Their high sensitivity enables high-fidelity detection of sound waves, and they have been widely used in fields such as sound recognition and sound source localization. However, due to interference such as ambient noise and system background noise, the clarity and intelligibility of speech signals are inevitably affected. Therefore, performing speech enhancement processing on the speech signals collected by fiber optic microphones is of great significance for further improving speech quality.
[0003] Depending on whether the sound signal is collected by a single microphone or multiple microphones, it can be divided into single-channel and multi-channel speech enhancement algorithms. Compared to multiple microphones, single microphones have attracted much attention due to their small size, low cost, and low power consumption. Single-channel speech enhancement algorithms primarily leverage the inherent characteristics of speech for speech enhancement. By extracting the differences between speech and noise signal characteristics in the time and transform domains, they suppress noise components and achieve speech enhancement. Speech enhancement algorithms can be divided into traditional speech enhancement algorithms and deep learning-based speech enhancement algorithms. Deep learning-based speech enhancement requires a large amount of data as a training set to train the model, which is time-consuming, cannot be deployed immediately, and its performance depends to a certain extent on the distribution of the training set. Traditional single-channel speech enhancement includes various methods, among which the minimum mean square error class can further reduce residual noise without compromising speech quality. For example, the optimally modified log-spectral amplitude (OMLSA) method, based on the probability of speech presence, achieves noise suppression by optimizing the speech amplitude spectrum. In addition, the Tunable-Q Wavelet Transform (TQWT) decomposes the signal at multiple scales and removes interference noise by sparsifying the decomposed wavelet coefficients, thereby achieving speech enhancement.
[0004] Defects and shortcomings of existing technology:
[0005] (1) Although some existing single-channel speech enhancement algorithms have excellent effects on noise suppression, their implementation requires a clean speech signal as a reference, such as the Wiener filter and the adaptive filter, which is not desirable in practical applications. Therefore, it is necessary to implement a high-performance algorithm that can achieve denoising based only on the noisy speech itself.
[0006] (2) Due to the diversity of fiber optic microphone application scenarios, speech signals may have low signal-to-noise ratio and non-stationary noise. The noise estimation error will significantly affect the speech enhancement performance of OMLSA, resulting in a large amount of residual noise in the signal after OMLSA processing, which seriously affects the quality of the speech signal. The algorithm needs to be improved to be suitable for low signal-to-noise ratio situations.
[0007] (3) The performance of TQWT depends on the selection of parameters, such as the Q factor, oversampling rate, number of decomposition levels, and the threshold for wavelet coefficient sparsification. Improper parameter settings may result in the decomposed subband signal failing to accurately capture the main features of speech, thus affecting the enhancement effect. In speech signal processing, low-frequency components are usually important speech information, while the wavelet coefficients decomposed by TQWT are mainly composed of high-frequency components, which may result in low-frequency noise not being completely suppressed. Summary of the Invention
[0008] To solve the above technical problems, the present invention provides a single-channel speech enhancement method for a fiber optic microphone that integrates OMLSA and TQWT. Compared with traditional electronic microphones, the present invention uses a highly sensitive fiber optic microphone to collect signals, providing high-quality original data for subsequent speech enhancement processing. Moreover, the present invention only enhances the signal collected by the fiber optic microphone itself without providing additional clean speech signals, thereby achieving good speech enhancement effects and further improving language quality. The method can be widely used in practice.
[0009] The technical solution adopted by the present invention is:
[0010] A single-channel speech enhancement method for a fiber optic microphone integrating OMLSA and TQWT comprises the following steps:
[0011] Step 1: Assume that the speech signal detected by the fiber optic microphone is y(n), where n is the sampling point index, 0≤n≤N-1, and N is the total number of sampling points, which must be an even number.
[0012] Step 2: Select the Hanning window to frame the speech signal y(n) to obtain the time domain signal y(n,l), where l is the frame index, 0≤l≤L-1, and L is the total number of frames;
[0013] Step 3: Use discrete Fourier transform to convert the time domain signal y(n,l) into the frequency domain signal Y(k,l), where k is the frequency index, k = 0, 1, 2, ..., N-1;
[0014] Step 4: Enhance the speech based on the OMLSA algorithm and calculate the enhanced spectrum amplitude Calculated by the following formula:
[0015]
[0016] Among them: G min is the gain function when speech does not exist; p(k,l) is the posterior probability of the existence of speech in the lth frame at frequency k; G H1 (k, l) is the gain function when the lth frame of speech exists at frequency k; (G min |Y(k,l)|) (1-p(k,l)) Indicates that when the l-th frame of speech does not exist at frequency k, a gain G is applied to |Y(k,l)| min To suppress noise, (G H1 (k,l)·|Y(k,l)|) p(k,l) Indicates that when the l-th frame of speech exists at frequency k, a gain G is applied to |Y(k,l)| H1 (k, l) emphasizes the speech component, and the symbol |·| indicates the absolute value.
[0017] Step 5: Perform inverse Fourier transform on the enhanced signal spectrum amplitude of all frames, and restore the speech signal x1(n) enhanced by the OMLSA algorithm based on the Hanning window function;
[0018] Step 6: Determine whether the segmented signal-to-noise ratio segSNR of the speech signal x1(n) satisfies the following conditions:
[0019] segSNR≥5dB;
[0020] If satisfied, skip steps 7, 8, and 9. At this time, x1(n) is the final enhanced speech signal x(n). Otherwise, proceed to step 7.
[0021] Step 7: Perform TQWT decomposition on x1(n) to obtain J+1 wavelet coefficients w{j}, where J is the optimal total decomposition level, j is the level index, 1≤j≤J+1;
[0022] Step 8: Use basis pursuit to sparse the wavelet coefficients w{j} to obtain the sparse wavelet coefficients w1{j};
[0023] Step 9: Perform inverse TQWT on the sparse wavelet coefficients w1{j} to obtain the final enhanced speech signal x(n).
[0024] In the step 2, Indicates rounding up, N f is the Hanning window frame length, N m Frame shift.
[0025] The step 4 comprises the following steps:
[0026] Step 4.1: Initialize the following parameters:
[0027] The initial value of the gain function G when speech exists H1 (k,0)=1, the initial value of the posterior signal-to-noise ratio of the speech γ(k,0)=1, the initial noise power spectrum λ d (k,0)=|Y(k,0)| 2 , initial noise power spectrum before offset compensation
[0028] Among them, Y(k,0) is the spectrum information of the initial frame, k represents the frequency index; the initial value of the first power spectrum time smoothing S(k,0) = S f (k,0), the initial value S of the first power spectrum time smoothing minimum value min (k,0)=S f (k,0), the initial value of the second power spectrum time smoothing The initial value of the second power spectrum time smoothing minimum S f (k,0) is the initial value of the speech frequency smoothed power spectrum; l is initialized to 1.
[0029] Step 4.2: Calculate the noise power spectrum of the lth frame before offset compensation at frequency k The calculation formula is:
[0030]
[0031] in: is the noise power spectrum of the l-1 frame before offset compensation at frequency k, Y(k,l-1) is the frequency domain signal of the l-1 frame at frequency k, is the time-varying smoothing parameter of the l-1th frame at frequency k, and its calculation formula is:
[0032]
[0033] Where: α d Represents the smoothing parameter before bias compensation, 0<α d <1; p(k,l-1) is the posterior probability of the presence of speech in the l-1th frame at frequency k;
[0034] Step 4.3: Calculate the noise power spectrum λ of the lth frame at frequency k d The calculation formula for (k,l) is as follows:
[0035]
[0036] Where: β is the bias factor, and the calculation formula is:
[0037]
[0038] Among them: γ1 is the regulatory factor; represents the base of natural logarithm e raised to the power of -γ1;
[0039] Step 4.4: Calculate the posterior signal-to-noise ratio γ(k,l) of the lth frame of speech at frequency k. The calculation formula is:
[0040]
[0041] Step 4.5: Calculate the prior signal-to-noise ratio ξ(k,l) of the lth frame of speech at frequency k. The calculation formula is:
[0042]
[0043] Where: α is the weight factor that controls the trade-off between noise reduction and speech distortion; G H1 (k, l-1) is the gain function when the l-1th frame of speech exists at frequency k; γ(k, l-1) is the posterior signal-to-noise ratio of the l-1th frame of speech at frequency k; max{γ(k, l)-1,0} means the maximum value between γ(k, l)-1 and 0;
[0044] Step 4.6: Calculate the gain function G when the lth frame of speech exists at frequency k H1 (k,l), which is calculated as follows:
[0045]
[0046] in: represents the balance between the priori signal-to-noise ratio and the posterior signal-to-noise ratio of the l-th frame of speech at frequency k;
[0047] Step 4.7: Calculate the first time-smoothed power spectrum value S(k,l) of the l-th frame of speech at frequency k. The calculation formula is:
[0048] S(k,l)=α s ·S(k,l-1)+(1-α s )·S f (k,l);
[0049] Where: α s is the time smoothing parameter of the first power spectrum, S(k,l-1) is the time smoothing value of the first power spectrum of the l-1 frame speech at frequency k, S f (k, l) is the first power spectrum frequency smoothing value of the l-th frame speech at frequency k, and its calculation formula is:
[0050]
[0051] Where: b(i) is the normalized window function, the window length is 2w+1, w represents the half length of the window function, and i is the window function index, that is, Y(ki,l) is the frequency domain signal of the lth frame at frequency ki;
[0052] Step 4.8: Calculate the second power spectrum frequency smoothing of the l-th frame signal at frequency k and the second power spectrum time smoothing The calculation formulas are:
[0053]
[0054] in: is the second power spectrum time smoothing of the l-1 frame signal at frequency k, and the value of I(ki,l) is given by the following formula:
[0055]
[0056] Where: γ0 and ζ0 are two thresholds for determining the presence of speech, and “&&” represents the logical operation “and”;
[0057]
[0058] γ min (ki,l) represents the minimum signal-to-noise ratio of the l-th frame of speech at frequency ki, and ζ(ki,l) represents the short-term stationary index of the l-th frame of speech at frequency ki;
[0059] S(ki,l) is the first power spectrum time smoothing of the l-th frame speech at frequency ki, B min is the noise spectrum minimum estimation bias compensation factor, the minimum value S of the first power spectrum time smoothing of the l-th frame speech at frequency ki min (ki,l)=min{S min (ki,l-1),S(ki,l)},S min (ki,l-1) is the minimum value of the first time smoothing of the power spectrum of the l-1 frame of speech at frequency ki;
[0060] Step 4.9: Calculate the prior probability q(k,l) that the lth frame of speech does not exist at frequency k. The calculation formula is:
[0061]
[0062] in: S min (k, l-1) is the minimum value of the first power spectrum time smoothing of the l-1 frame speech at frequency k; γ1 represents the adjustment factor;
[0063] Step 4.10: Calculate p(k,l), which is calculated as follows:
[0064]
[0065] Step 4.11: Change the current G H1 Substitute (k, l) and p(k, l) into the calculation formula in step 4, as follows:
[0066]
[0067] Output frequency k at the enhanced spectrum amplitude of the lth frame
[0068] Step 4.12: Determine whether the following loop conditions are met:
[0069] l <L;
[0070] If satisfied, increase the value of l by 1 and repeat steps 4.2 to 4.11; otherwise, proceed to step 5.
[0071] In step 5, the enhanced signal spectrum amplitude of all frames is subjected to inverse Fourier transform, and the calculation formula is:
[0072]
[0073] in: is the time domain signal of the lth frame of speech, n1 is the sampling point index of the lth frame of time domain signal, 0≤n1≤N f -1, e is the base of natural logarithm, i1 is the imaginary unit;
[0074] And according to the Hanning window function, the speech signal x1(n) enhanced by the OMLSA algorithm is restored. The specific steps are:
[0075] right Apply the Hanning window to obtain the time domain signal a(n1,l) after windowing the l-th frame of speech:
[0076]
[0077] Where: hanning(n1) is the Hanning window formula, cos is the cosine trigonometric function; x1(n) is obtained by overlapping and adding a(n1,l) in frames. The calculation formula is:
[0078]
[0079] Among them: hanning_overlap(n1) Hanning window overlap frame weight, its calculation formula is:
[0080]
[0081] In step 6, calculating segSNR includes the following steps:
[0082] Step 6.1: Divide x1(n) into M segments, where the value of M is:
[0083]
[0084] in: Indicates rounding down, S m is the size of each segment, usually 256, 512 or 1024, etc., m is the index of the segment, 1≤m≤M.
[0085] Step 6.2: For the t-th sampling point noisy in the m-th segment of the residual noise in x1(n), 1_m The value of (t) is estimated:
[0086]
[0087] Where: x 1_m (t) is the signal of the t-th sampling point in the m-th segment of x1(n), and th is the energy threshold of the speech silence segment.
[0088] Step 6.3: Calculate segSNR:
[0089]
[0090] In step 7, decomposing x1(n)TQWT to obtain J+1 wavelet coefficients w{j} includes the following steps:
[0091] Step 7.1: Use the reconstruction error E to find the optimal value of the quality factor Q opt and the optimal value r of redundancy r opt , the specific steps are:
[0092] S7.1.1: Set the maximum value that Q can take to be Q max , the maximum value that r can take is r max , initialize Q = 1, r takes [3,r max ] any value in the interval;
[0093] S7.1.2: Calculate the number of decomposition levels J1 from Q and r using the following formula:
[0094]
[0095] The Q value at this time is stored in the sequence Q_cell. The current three parameters Q, r, and J1 are used to perform steps 7.2 to 7.8. Step 8 is skipped. At this time, w{j} is w1{j} in step 9, and step 9 is performed.
[0096] S7.1.3: Calculate the reconstruction error E = max(|x1(n) - x(n)|) and store it in the sequence E_cell.
[0097] S7.1.4: Determine whether the following loop conditions are met:
[0098] Q max ;
[0099] If satisfied, add 0.2 to the value of Q and repeat S7.1.2 to S7.1.3; otherwise, proceed to S7.1.5;
[0100] S7.1.5: The minimum reconstruction error min(E_cell) corresponds to the value in Q_cell, which is Q opt .
[0101] S7.1.6: Initialize r = 3, Q to [1, Q max ] any value in the interval;
[0102] S7.1.7: Calculate the decomposition level J1, store the current value of r in the sequence r_cell, and proceed to steps 7.2 to 7.8 using the current parameters, skipping step 8. At this point, w{j} is w1{j} in step 9, and proceed to step 9.
[0103] S7.1.8: Calculate the reconstruction error E = max(|x1(n) - x(n)|) and store it in the sequence E_cell_1.
[0104] S7.1.9: Determine whether the following loop conditions are met:
[0105] r <r max ;
[0106] If satisfied, add 0.2 to the value of r and repeat S7.1.7 to S7.1.8; otherwise, proceed to S7.1.10;
[0107] S7.1.10: The minimum reconstruction error min(E_cell_1) corresponds to the value in r_cell, which is r opt .
[0108] S7.1.11: Use Q opt and r opt Find the optimal decomposition level J.
[0109] Step 7.2: Perform unitary discrete Fourier transform uDFT(x1(n)) on x1(n) to obtain X1(k), which is given by:
[0110]
[0111] Where: X_1(k) is the discrete Fourier transform of x1(n), and j=1 is initialized.
[0112] Step 7.3: Calculate the length of the low-pass subband of layer j and high-pass subband length
[0113]
[0114] Among them: "round" means rounding operation,
[0115] Step 7.4: Calculate the passband length P of the j-th low-pass layer j , high-pass passband length S j , the length of the transition zone T j :
[0116]
[0117] Step 7.5: Obtain the j-th layer low-pass sub-band sequence and high-pass sub-band sequence. The specific steps are as follows:
[0118] S7.5.1: Obtain the j-th layer low-pass subband sequence k_low is the low-pass subband frequency index,
[0119]
[0120] Where: represents the initial value of the j-th layer low-pass subband sequence, X1(0) represents the initial value of X1(k) obtained by performing unitary discrete Fourier transform on x1(n);
[0121]
[0122] Where: Represents the j-th layer low-pass subband sequence from frequency 1 to frequency P j The value of X1(1:P j ) indicates that X1(k) goes from frequency 1 to frequency P j The value of
[0123]
[0124] Where: Represents the j-th layer low-pass subband sequence from frequency P j +1 to frequency P j +T j The value of
[0125] X1(P j +(1:T j )) indicates that X1(k) is from frequency P j +1 to frequency P j +Tj The value of θ(1:T j ) means taking values from 1 to T in sequence j Perform the following theta operation:
[0126]
[0127] Where: D represents the parameter value for θ operation;
[0128]
[0129] Where: Indicates the j-th layer low-pass subband sequence at frequency The value of
[0130]
[0131] Where: Indicates that the j-th layer low-pass subband sequence is from frequency To frequency The value of X1(NP j -(1:T j )) indicates that X1(k) is from frequency NP j -1 to frequency NP j -T j The value of
[0132]
[0133] Where: A1:B1 means the sequence is indexed from A1 to B1, and “.*” means multiplying the corresponding elements of the sequence;
[0134] S7.5.2: Obtain the j-th layer high-pass subband sequence V1 j (k_high), k_high is the high-pass subband frequency index,
[0135] V1 j (0)=0;
[0136] Where: V1 j (0) represents the initial value of the j-th layer high-pass subband sequence;
[0137] V1 j (1:T j )=X1(P j +(1:T j )).*θ(T j -(1:T j )+1);
[0138] Where: V1 j (1:Tj ) represents the j-th layer high-pass subband sequence from frequency 1 to frequency T j The value of X1(P j +(1:T j )) indicates that X1(k) is from frequency P j +1 to frequency P j +T j The value of θ(T j -(1:T j )+1) means taking the value T in sequence j To 1, perform theta operation;
[0139] V1 j (T j +(1:S j ))=X1(P j +T j +(1:S j ));
[0140] Where: V1 j (T j +(1:S j )) represents the j-th layer high-pass subband sequence from frequency T j +1 to frequency T j +S j The value of
[0141] X1(P j +T j +(1:S j )) indicates that X1(k) is from frequency P j +T j +1 to frequency P j +T j +S j The value of
[0142]
[0143] Where: Indicates that the j-th layer high-pass subband sequence is at frequency X1(N / 2) represents the value of X1(k) at k=N / 2;
[0144]
[0145] Where: Indicates that the j-th layer high-pass subband sequence is from frequency To frequency The value of X1(NP j -T j -(1:S j )) indicates that X1(k) is from frequency NP j-T j -1 to frequency NP j -T j -S j The value of
[0146]
[0147] Where: Indicates that the j-th layer high-pass subband sequence is from frequency To frequency The value of
[0148] X1(NP j -(1:T j )) indicates that X1(k) is from frequency NP j -1 to frequency NP j -T j The value of
[0149] Step 7.6: V1 j (k_high) performs inverse unitary discrete Fourier transform to obtain the j-th wavelet coefficient w{j} of TQWT decomposition:
[0150] w{j}=uDFTinv(V1 j (k_high));
[0151] Where: uDFTinv represents the unitary inverse discrete Fourier transform, and its relationship with the inverse discrete Fourier transform is:
[0152]
[0153] Where: w_inv{j} is V1 j Inverse discrete Fourier transform of (k_high);
[0154] Step 7.7: Determine whether the following loop conditions are met:
[0155] j <J;
[0156] If so, add 1 to the value of j and repeat steps 7.3 to 7.6; otherwise, proceed to step 7.8.
[0157] Step 7.8: Calculate the J+1th wavelet coefficient of the TQWT decomposition Represents the J-th layer low-pass subband sequence; so far, all the J+1 wavelet coefficients decomposed have been obtained.
[0158] In step 8, basis pursuit is used to perform sparsification on w{j}, which includes the following steps:
[0159] Step 8.1: Initialize j=1;
[0160] Step 8.2: Generate an l j ×l j The identity matrix A j , l j is the length of the jth wavelet coefficient; vector b j =A j T *w{j} represents b j Projection onto the identity matrix, where A j T A j The transpose of ; matrix B j =[A j T *A j ,-A j T *A j ;-A j T *A j ,A j T *A j ] is used to construct quadratic terms in optimization problems;
[0161] Step 8.3: Calculate the regularization parameter τ of the jth layer j , and its calculation formula is:
[0162]
[0163] in: is a regulation factor reflecting the change of sparsity with the number of layers j, is the coefficient for adjusting the amplitude of w{j}, ln() is the natural logarithm, and median() is the median operation;
[0164] Step 8.4: Construct a j The column vector c j , as the driving vector for sparse optimization, is generated by the following formula:
[0165]
[0166] Where: C j is of length 2·l j And a column vector of all 1s;
[0167] Step 8.5: Solve the quadratic optimization problem:
[0168]
[0169] Among them: z is the intermediate quantity when solving, and its length is 2·l jColumn vector of z T represents the transpose of z; c j T Indicates c j The transpose of It represents the optimal solution of z when minimizing the objective function, and the optimal solution is z j .
[0170] Step 8.6: By z j The sparse wavelet coefficient w1{j} is obtained, and its calculation formula is:
[0171] w1{j}=z j (1:l j )-z j (l j +1:2·l j );
[0172] Where: z j (1:l j ) represents the column vector z j Rows 1 to 1 j The value of the row, z j (l j +1:2·l j ) represents the column vector z j No. 1 j +1 row to 2.1 j The value of the row;
[0173] Step 8.7: Determine whether the following loop conditions are met:
[0174] j≤J;
[0175] If satisfied, the value of j is increased by 1 and steps 8.2 to 8.6 are repeated; otherwise, proceed to step 9.
[0176] In step 9, performing inverse TQWT on w1{j} includes the following steps:
[0177] Step 9.1: Find the unitary discrete Fourier transform of the J+1th sparse wavelet coefficient w1{J+1} as Y J+1 (k_2), where k_2 is its frequency index, Represents Y J+1 The length of (k_2); initialize j=J.
[0178] Step 9.2: Calculate the passband length P of the j-th low-pass layer after sparsification j _1. High-pass passband length S j _1. Length of transition zone T j _1 are:
[0179]
[0180]
[0181] Where: N j _1 is the length of the unitary discrete Fourier transform of the jth sparse wavelet coefficient w1{j}, N j+1 _1 is the length of the unitary discrete Fourier transform of the j+1th sparse wavelet coefficient w1{j+1}.
[0182] Step 9.3: Calculate the j-th layer reconstructed signal Y j (k1), 0≤k1≤M j -1, k1 represents the frequency index of the j-th layer reconstructed signal; comprising the following steps:
[0183] S9.3.1: Perform a unitary discrete Fourier transform on the j-th sparse wavelet coefficient w1{j} to W j (k_1), k_1 is its frequency index, 0≤k_1≤N j _1-1;
[0184] S9.3.2: Signal Y reconstructed by layer j+1 j+1 (k1) Find the j-th layer sparse low-pass subband sequence
[0185]
[0186] Where: represents the initial value of the j-th layer sparse low-pass subband sequence, Y j+1 (0) represents the j+1th layer reconstructed signal Y j+1 The initial value of (k1);
[0187]
[0188] Where: Represents the j-th layer of sparse low-pass subband sequence from frequency 1 to frequency P j The value of _1, Y j+1 (1:P j _1) indicates Y j+1 (k1) from frequency 1 to frequency P j The value of _1;
[0189]
[0190] Where: Represents the j-th layer sparse low-pass subband sequence from frequency P j _1+1 to frequency P j _1+T j _1 value; Y j+1 (Pj _1+(1:T j _1)) represents Y j+1 (k1) from frequency P j _1+1 to frequency P j _1+T j _1 value;θ(1:T j _1) means taking values from 1 to T in sequence j _1 performs theta operation;
[0191]
[0192] Where: Represents the j-th layer sparse low-pass subband sequence from frequency P j _1+T j _1+1 to frequency P j _1+T j _1+S j The value of _1;
[0193]
[0194] Where: Represents the j-th layer sparse low-pass subband sequence at frequency M j The value of
[0195]
[0196] Where: Denotes the j-th layer sparse low-pass subband sequence from frequency M j -P j _1-T j _1-1 to frequency M j -P j _1-T j _1-S j The value of _1;
[0197]
[0198] Where: Represents the j-th layer sparse low-pass subband sequence from frequency M j -P j _1-1 to frequency M j -P j _1-T j _1 value; Y j+1 (N j+1 _1-P j _1-(1:T j _1)) represents Y j+1 (k1) from frequency N j+1 _1-Pj _1-1 to frequency N j+1 _1-P j _1-T j The value of _1,
[0199]
[0200] Where: Represents the j-th layer sparse low-pass subband sequence from frequency M j -1 to frequency M j -P j _1 value; Y j+1 (N j+1 _1-(1:P j _1)) represents Y j+1 (k1) from frequency N j+1 _1-1 to frequency N j+1 _1-P j The value of _1;
[0201] S9.3.3: Find the j-th layer sparsified high-pass subband sequence Y1 j (k1):
[0202] Y1 j (0)=0;
[0203] Where: Y1 j (0) represents the initial value of the j-th layer sparsified high-pass subband sequence;
[0204] Y1 j (1:P j _1)=0;
[0205] Where: Y1 j (1:P j _1) represents the j-th layer sparse high-pass subband sequence from frequency 1 to frequency P j The value of _1;
[0206] Y1 j (P j _1+(1:T j _1))=W j (1:T j _1).*θ(T j _1-(1:T j _1)+1);
[0207] Where: Y1 j (P j _1+(1:T j _1)) represents the j-th layer sparse high-pass subband sequence from frequency P j _1+1 to frequency P j_1+T j _1 value; W j (1:T j _1) represents the unitary discrete Fourier transform of w1{j} to W j (k_1) from frequency 1 to frequency T j _1 value;θ(T j _1-(1:T j _1)+1) means taking the value T in sequence j _1 to 1 performs theta operation;
[0208] Y1 j (P j _1+T j _1+(1:S j _1))=W j (T j _1+(1:S j _1));
[0209] Where: Y1 j (P j _1+T j _1+(1:S j _1)) represents the j-th layer sparse high-pass subband sequence from frequency P j _1+T j _1+1 to frequency P j _1+T j _1+S j _1 value; W j (T j _1+(1:S j _1)) indicates W j (k_1) from frequency T j _1+1 to frequency T j _1+S j The value of _1;
[0210] Y1 j (M j / 2)=W j (N j _1 / 2);
[0211] Where: Y1 j (M j / 2) represents the j-th layer sparse high-pass subband sequence at frequency M j / 2 value; W j (N j _1 / 2) represents W j (k_1) when k_1=N j The value at _1 / 2;
[0212] Y1j (M j -P j _1-T j _1-(1:S j _1))=W j (N j _1-T j _1-(1:S j _1));
[0213] Where: Y1 j (M j -P j _1-T j _1-(1:S j _1)) represents the j-th layer sparse high-pass subband sequence from frequency M j -P j _1-T j _1-1 to frequency M j -P j _1-T j _1-S j _1 value; W j (N j _1-T j _1-(1:S j _1)) indicates W j (k_1) from frequency N j _1-T j _1-1 to frequency N j _1-T j _1-S j The value of _1;
[0214] Y1 j (M j -P j _1-(1:T j _1))=W j (N j _1-(1:T j _1)).*θ(T j _1-(1:T j _1)+1);
[0215] Where: Y1 j (M j -P j _1-(1:T j _1)) represents the j-th layer sparse high-pass subband sequence from frequency M j -P j _1-1 to frequency M j -P j _1-T j _1 value; Wj (N j _1-(1:T j _1)) indicates W j (k_1) from frequency N j _1-1 to frequency N j _1-T j The value of _1;
[0216] Y1 j (M j -(1:P j _1))=0;
[0217] Where: Y1 j (M j -(1:P j _1)) represents the j-th layer sparse high-pass subband sequence from frequency M j -1 to frequency M j -P j The value of _1;
[0218] S9.3.4: At this time, the j-th layer reconstructs the signal
[0219] Step 9.4: Determine whether the following loop conditions are met:
[0220] j>1
[0221] If satisfied, the value of j is reduced by 1 and steps 9.2 to 9.3 are repeated; otherwise, proceed to step 9.5;
[0222] Step 9.5: After the step 9.4 cycle is completed, the final reconstructed signal Y 1 (k) Perform inverse unitary discrete Fourier transform to obtain the final enhanced language signal x(n), x(n) = uDFTinv(Y 1 (k)).
[0223] uDFTinv stands for unitary inverse discrete Fourier transform, which is different from the ordinary inverse discrete Fourier transform.
[0224] The present invention provides a single-channel speech enhancement method for optical fiber microphones integrating OMLSA and TQWT, and the technical effects are as follows:
[0225] 1) This invention fully leverages the characteristics of OMLSA and TQWT, combining them to achieve complementary results. While OMLSA has strong denoising capabilities in the presence of speech, it still leaves significant residual noise when the signal-to-noise ratio is low. TQWT, through decomposition, generates wavelet coefficients, which further suppress OMLSA's residual noise when the signal-to-noise ratio is low.
[0226] 2) The present invention uses the reconstruction error to provide a reference for the Q factor and oversampling rate required for TQWT denoising. By minimizing the reconstruction error, the optimal parameters are selected to achieve the best speech enhancement performance of TQWT.
[0227] 3) In the present invention, when the segmented signal-to-noise ratio of the signal after OMLSA processing is high, it indicates that the residual noise of the speech signal is very small. If TQWT processing is still performed, it will cause part of the useful signal to be lost. Therefore, the present invention uses the segmented signal-to-noise ratio as the judgment condition for the signal after OMLSA processing to determine whether further TQWT processing is needed, thereby saving computing resources while ensuring optimal speech quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0228] Figure 1 Flowchart of the present invention.
[0229] Figure 2 This is how the reconstruction error changes with the Q value in Example 1.
[0230] Figure 3 This is how the reconstruction error changes with the r value in Example 1.
[0231] Figure 4 This is a comparison chart of the results of performing OMLSA on the signal in Example 1 and the speech enhancement results of the present invention. DETAILED DESCRIPTION
[0232] Example 1:
[0233] The present invention performs speech enhancement on the signal collected by the fiber optic microphone. The sampling frequency is 10kHz and the collected signal has 32812 sampling points. The parameters are set as follows:
[0234] Hanning window frame length and frame shift settings during framing:
[0235] Hanning window frame length N f Take 512, frame shift N m Take 128.
[0236] Gain function G when speech is absent min Value: G min =0.1257.
[0237] Smoothing parameter α before bias compensation d The value of: α d =0.85
[0238] The value of the adjustment factor γ1 is: γ1=3.
[0239] The value of the weight factor α that controls the trade-off between noise reduction and speech distortion is: α=0.92.
[0240] The first power spectrum time smoothing parameter α s The value of: α s =0.9.
[0241] The value of the half-length w of the normalized window function is: w=1.
[0242] The two thresholds γ0 and ζ0 for language existence judgment are set as follows:
[0243] Take γ0=4.6, ζ0=1.67.
[0244] Noise spectrum minimum estimation bias compensation factor B min Value: B min =1.66.
[0245] Segmented signal-to-noise ratio: S m =256.
[0246] The value of the speech silence segment energy threshold th is: th=0.02.
[0247] Maximum quality factor Q max and the maximum redundancy r max set up:
[0248] Take Q max =10, r max =6.
[0249] The signal is enhanced using the present invention and compared with the original noisy signal and the result of only OMLSA enhancement:
[0250] The calculated signal-to-noise ratio of the signal after only OMLSA enhancement is 4.25dB, which is less than 5dB. Therefore, TQWT is required for further denoising. Figure 2 and Figure 3 It can be seen that the best Q value for this speech is 1.8, the best r value is 3, and the corresponding decomposition layer number J is 11. The enhancement effect of the algorithm on the speech is observed by the time domain waveform of the signal. Figure 4 It can be seen that after OMLSA processing, a lot of noise in the speech signal is removed, but there is still some residual noise. Compared with the processing results of the present invention, the present invention can further reduce the residual noise on the basis of OMLSA, especially the blank part of the speech can see an obvious denoising effect, and the waveform of the part of the speech that exists is very similar to the original speech, which shows that the present invention can minimize the interference noise while ensuring the fidelity of the speech, thereby achieving speech enhancement.
[0251] Example 2:
[0252] Eight speech sentences were selected from the TIMIT speech database, and four noise types, F16, Pink, Tank, and White, were selected from the Noisex-92 noise database. These were mixed with the speech signals at signal-to-noise ratios of -5dB, 0dB, 5dB, and 10dB, respectively, to obtain speech signals with different signal-to-noise ratios. The noisy speech signals were enhanced using OMLSA and the present invention, respectively. The speech enhancement algorithm was evaluated using a comprehensive evaluation index, which includes three parameters: signal distortion Csig, background noise Cbak, and overall quality Covl. A higher score indicates better speech quality. The parameter settings for speech enhancement are the same as those in Example 1. The scores of the comprehensive evaluation indexes under different noise types are shown in Table 1:
[0253] Table 1 Comparison of comprehensive evaluation indicators under different noise types
[0254]
[0255]
[0256] As shown in Table 1, with the exception of a few indicators in areas with shading, the speech quality of the speech enhancement method of the present invention is significantly improved compared to that of OMLSA alone. Furthermore, in all cases, the Cbak indicator of the present invention is comparable to that of OMLSA alone, demonstrating the excellent performance of the present method in suppressing background noise. Furthermore, when the speech signal evaluation index is high, the index data after OMLSA processing and the present invention are the same, as shown by the bolded portion of the index data, indicating that the segmented signal-to-noise ratio of the speech after OMLSA processing is greater than 5dB, indicating that the speech already meets the application requirements and does not require further enhancement using TQWT.
[0257] The present invention first frames the collected speech signal, performs OMLSA speech enhancement on each frame, and uses the segmented signal-to-noise ratio as a decision criterion to determine whether TQWT is needed to remove residual noise. The reconstruction error is used to select the optimal values of the Q factor and redundancy r parameters required for TQWT decomposition, ensuring that the most suitable set of wavelet coefficients is obtained for the speech signal decomposition. A threshold is adaptively selected based on each wavelet coefficient, and the wavelet coefficients are sparsely distributed using basis pursuit. The final enhanced speech signal is reconstructed through the sparse wavelet coefficients. Compared to using only OMLSA for denoising, the present method can further effectively suppress residual noise on this basis, improving the speech enhancement effect.
Claims
1. A fiber microphone single-channel speech enhancement method integrating OMLSA and TQWT, characterized in that The following steps are involved: Step 1: Assume that the voice signal detected by the fiber optic microphone is y ( n ),in, n is the sampling point index, , N is the total number of sampling points; Step 2: Select the Hanning window for the speech signal y ( n ) frame to obtain the time domain signal y ( n, l ),in, l is the frame index, , L is the total number of frames; Step 3: Use discrete Fourier transform to transform the time domain signal y ( n, l ) is converted into a frequency domain signal Y ( k , l ),in, k is the frequency index, k = 0, 1, 2, …, N -1; Step 4: Enhance the speech based on the OMLSA algorithm and calculate the enhanced spectrum amplitude Calculated by the following formula: ; in: G min is the gain function when speech does not exist; p ( k, l ) is the frequency k Next l The posterior probability of the presence of speech in the frame; G H1 ( k, l ) is the frequency k Next l Gain function when frame speech exists; Indicates frequency k Next l When frame speech does not exist, the | Y ( k , l )|Apply gain G min To suppress noise, Indicates frequency k Next l When frame speech exists, the | Y ( k , l )|Apply gain G H1 ( k, l ) emphasizes the phonetic component, and the symbol |•| indicates taking the absolute value; Step 5: Perform inverse Fourier transform on the enhanced signal spectrum amplitude of all frames, and restore the speech signal enhanced by the OMLSA algorithm based on the Hanning window function. x 1( n ); Step 6: Determine the speech signal x 1( n ) meets the following conditions: segSNR ≥ 5 dB; If satisfied, skip steps 7, 8, and 9. x 1( n ) is the final enhanced speech signal x ( n ), otherwise, proceed to step 7; Step 7: Right x 1( n ) is decomposed by TQWT to obtain J + 1 wavelet coefficient w { j },in, J is the optimal total number of decomposition layers, j is the layer index, 1 ≤ j ≤ J + 1; Step 8: Use basis pursuit to align wavelet coefficients w { j } to perform sparseness and obtain the sparse wavelet coefficients w 1{ j }; Step 9: Sparse wavelet coefficients w 1{ j } Perform inverse TQWT to obtain the final enhanced speech signal x ( n ); In step 5, the enhanced signal spectrum amplitude of all frames is subjected to inverse Fourier transform, and the calculation formula is: ; in: For the l The time domain signal of frame speech, n 1st l The sampling point index of the frame time domain signal, , e is the base of natural logarithm, i 1 is the imaginary unit; And according to the Hanning window function, the speech signal enhanced by the OMLSA algorithm is restored. x 1( n ), the specific steps are: right Apply the Hanning window to get the l Time domain signal after frame speech windowing : ; in: is the Hanning window formula, , cos is the cosine trigonometric function; Overlap-add by frame x 1( n ), which is calculated as follows: ; in: The Hanning window overlapping frame weight is calculated as follows: ; In step 7, x 1( n ) TQWT decomposition yields J + 1 wavelet coefficient w { j }, including the following steps: Step 7.1: Utilize reconstruction error E Find the quality factor Q The best value Q opt and redundancy r The best value r opt ; Step 7.2: x 1( n ) performs unitary discrete Fourier transform ,get X 1( k ), the formula is: ; in: X_ 1( k )for x 1( n ), initialize j = 1; Step 7.3: Calculate the j Low-pass subband length of the layer and high-pass subband length : ; ; Among them: "round" means rounding operation, , ; Step 7.4: Calculate the j Passband length of layer low-pass , high-pass passband length , the length of the transition zone : ; ; ; Step 7.5: Get the j Layer low-pass sub-band sequence and high-pass sub-band sequence; Step 7.6: Perform unitary discrete Fourier inverse transform to obtain the first j Wavelet coefficients w { j }: ; in: uDFTinv represents the unitary inverse discrete Fourier transform, and its relationship with the inverse discrete Fourier transform is: ; in: w _ inv { j }for The inverse discrete Fourier transform of Step 7.7: Determine whether the following loop conditions are met: j < J ; If satisfied, then j Add 1 to the value and repeat steps 7.3 to 7.6, otherwise proceed to step 7.8; Step 7.8: Calculate the TQWT decomposition J + 1 wavelet coefficient , Indicates the J Layer low-pass subband sequence; decomposed so far J + 1 wavelet coefficients have all been calculated; In step 9, w 1{ j Performing inverse TQWT includes the following steps: Step 9.1: Find the J +1 sparse wavelet coefficient w 1{ J +1} is the unitary discrete Fourier transform ,in, k_ 2 is its frequency index, 0 ≤ k_ 2 ≤ , express Length; Initialization j = J; Step 9.2: Calculate the sparsified j Passband length of layer low-pass , high-pass passband length , the length of the transition zone They are: ; ; ; in: For the j Sparse wavelet coefficients w 1{ j }The length of the unitary discrete Fourier transform, For the j + 1 sparse wavelet coefficient w 1{ j+ 1} The length of the unitary discrete Fourier transform; Step 9.3: Calculate the j Layer reconstruction signal , 0 ≤ k 1≤ , k 1 means the j The frequency index of the layer reconstruction signal includes the following steps: S9.3.1: For j Sparse wavelet coefficients w 1{ j } Perform unitary discrete Fourier transform to , k_ 1 is its frequency index, 0≤ k_ 1≤ ; S9.3.2: By j + 1 layer of reconstructed signal Seeking the first j Layer-sparse low-pass subband sequence ; ; Where: Indicates the j The initial value of the low-pass subband sequence for layer sparsification, Indicates the j + 1 layer of reconstructed signal The initial value of ; Where: Indicates the j The low-pass subband sequence of the layer sparsification is from frequency 1 to frequency The value of express From frequency 1 to frequency The value of ; Where: Indicates the j The low-pass subband sequence of the layer sparsification is from the frequency To frequency The value of express From the frequency + 1 to frequency + The value of Indicates that the values 1 to implement θ Operation; ; Where: Indicates the j The low-pass subband sequence of the layer sparsification is from the frequency To frequency The value of ; Where: Indicates the j The low-pass subband sequence of the layer sparsification is The value of ; Where: Indicates the j The low-pass subband sequence of the layer sparsification is from the frequency To frequency The value of ; Where: Indicates the j The low-pass subband sequence of the layer sparsification is from the frequency To frequency The value of express From the frequency To frequency The value of ; Where: Indicates the j The low-pass subband sequence of the layer sparsification is from the frequency To frequency The value of express From the frequency To frequency The value of S9.3.3: Find the first j Layer-sparse high-pass subband sequence : ; Where: Indicates the j Initial value of the high-pass subband sequence for layer sparsification; ; Where: Indicates the j The high-pass subband sequence of the layer sparsification is from frequency 1 to frequency The value of ; Where: Indicates the j The high-pass subband sequence of the layer sparsification is from the frequency To frequency The value of Express w 1{ j } Perform unitary discrete Fourier transform to From frequency 1 to frequency The value of Indicates that the values are taken in sequence To 1 execution θ Operation; ; Where: Indicates the j The high-pass subband sequence of the layer sparsification is from the frequency To frequency The value of express From the frequency To frequency The value of ; Where: Indicates the j The high-pass subband sequence of the layer sparsification is The value at express exist The value at ; Where: Indicates the j The high-pass subband sequence of the layer sparsification is from the frequency To frequency The value of express From the frequency To frequency The value of ; Where: Indicates the j The high-pass subband sequence of the layer sparsification is from the frequency To frequency The value of express From the frequency To frequency The value of ; Where: Indicates the j The high-pass subband sequence of the layer sparsification is from the frequency To frequency The value of S9.3.4: At this time j Layer reconstruction signal ; Step 9.4: Determine whether the following loop conditions are met: j >1; If satisfied, then j The value of is reduced by 1 and steps 9.2 to 9.3 are repeated, otherwise proceed to step 9.5; Step 9.5: Determine the final reconstructed signal after the cycle of step 9.4 is completed Perform unitary discrete Fourier inverse transform to obtain the final enhanced language signal x ( n ), x ( n ) = uDFTinv ( Y 1 ( k )).
2. The method for single-channel speech enhancement of a fiber microphone by integrating OMLSA and TQWT according to claim 1 is characterized in that: In the step 2, , Indicates rounding up. N f is the Hanning window frame length, N m Frame shift.
3. The method for single-channel speech enhancement of a fiber microphone by integrating OMLSA and TQWT according to claim 1 is characterized in that: The step 4 comprises the following steps: Step 4.1: Initialize the following parameters: Initial value of the gain function when speech is present G H1 ( k, 0) = 1, the initial value of the posterior signal-to-noise ratio of the speech , initial noise power spectrum = | Y ( k , 0)| 2 , initial noise power spectrum before offset compensation = | Y ( k , 0)| 2 ; in, Y ( k , 0) is the spectrum information of the initial frame, k Indicates the frequency index; the initial value of the first power spectrum time smoothing S ( k , 0) = S f ( k, 0), the initial value of the first power spectrum time smoothing minimum value S min ( k, 0) = S f ( k, 0), the initial value of the second power spectrum time smoothing = S f ( k , 0), the initial value of the second power spectrum time smoothing minimum value = S f ( k , 0), S f ( k, 0) is the initial value of the speech frequency smoothed power spectrum; l Initialized to 1; Step 4.2: Calculate the frequency k Next l Noise power spectrum before frame offset compensation , and its calculation formula is: ; in: Frequency k Next l - 1 frame noise power spectrum before offset compensation, Y ( k , l - 1) is the frequency k Next l - 1 frame of frequency domain signal, Frequency k Next l - The time-varying smoothing parameter of 1 frame is calculated as follows: ; in: Indicates the smoothing parameter before bias compensation, 0 < < 1; p ( k,l - 1) is the frequency k Next l - The posterior probability of the presence of a frame of speech; Step 4.3: Calculate the frequency k Next l Noise power spectrum of the frame The calculation formula is as follows: ; in: is the bias factor, and the calculation formula is: ; in: is the regulating factor; The base e of the natural logarithm power; Step 4.4: Calculate the frequency k Next l Posterior SNR of speech frame , and its calculation formula is: ; Step 4.5: Calculate the frequency k Next l Prior signal-to-noise ratio of frame speech , and its calculation formula is: ; in: A weight factor to control the trade-off between noise reduction and speech distortion; G H1 ( k, l - 1) is the frequency k Next l -Gain function when 1 frame of speech exists; Frequency k Next l - Posterior signal-to-noise ratio of 1 frame of speech; Indicates taking The maximum value between 0 and 0; Step 4.6: Calculate the frequency k Next l Gain function when frame speech exists G H1 ( k, l ), which is calculated as follows: ; in: , indicating the frequency k Next l The balance between the priori signal-to-noise ratio and the posterior signal-to-noise ratio of the frame speech; Step 4.7: Calculate the frequency k Next l The first time smoothing value of the power spectrum of the frame speech S ( k , l ), which is calculated as follows: ; in: is the first power spectrum time smoothing parameter, S ( k , l -1) is the frequency k Next l -1 frame speech first power spectrum time smoothing value, S f ( k, l ) is the frequency k Next l The first power spectrum frequency smoothing value of the frame speech is calculated as follows: ; in: b ( i ) is the normalized window function, the window length is 2 w +1, w represents the half length of the window function, i is the window function index, that is ; Frequency k - i Next l Frequency domain signal of the frame; Step 4.8: Calculate the frequency k Next l Second power spectrum frequency smoothing of frame signal and the second power spectrum time smoothing , the calculation formulas are: ; in: Frequency k Next l- The second power spectrum time smoothing of 1 frame signal, The value of is given by the following formula: ; in: and These are the two thresholds for determining the presence of speech, and "&&" represents the logical operation "and"; , ; Indicates frequency k - i Next l The minimum signal-to-noise ratio of frame speech, Indicates frequency k - i Next l Short-term stability index of frame speech; S ( k - i, l ) is the frequency k - i Next l The first power spectrum time smoothing of the frame speech, B min Estimate the bias compensation factor for the noise spectrum minimum, frequency k - i Next l The minimum value of the first time smoothing of the power spectrum of the frame speech S min ( k - i, l ) = min{ S min ( k - i, l -1), S ( k - i, l )}, S min ( k - i, l -1) is the frequency k - i Next l - The minimum value of the first time smoothing of the power spectrum of a frame of speech; Step 4.9: Calculate the frequency k Next l Prior probability of the absence of speech in the frame q ( k, l ), which is calculated as follows: ; in: , , , S min ( k, l -1) is the frequency k Next l - The minimum value of the first time smoothing of the power spectrum of a frame of speech; represents the regulating factor; Step 4.10: Calculation p ( k, l ), which is calculated as follows: ; Step 4.11: Change the current G H1 ( k,l )and p ( k,l ) into the calculation formula in step 4, as follows: Output frequency k Next l Spectral amplitude after frame enhancement ; Step 4.12: Determine whether the following loop conditions are met: ; If satisfied, then l Add 1 to the value and repeat steps 4.2 to 4.11, otherwise proceed to step 5.
4. The method for single-channel speech enhancement of a fiber microphone by integrating OMLSA and TQWT according to claim 1 is characterized in that: In step 6, calculating segSNR includes the following steps: Step 6.1: x 1( n ) is divided into M part, M The value of is: ; in: Indicates rounding down. S m is the size of each segment, m is the index of the segment, 1 ≤ m ≤ M; Step 6.2: x 1( n ) in the residual noise m Section 1 t Sampling points noisy 1_m ( t ) is estimated: ; in: x 1_m ( t )for x 1( n ) No. m Section 1 t Sampling point signals, th is the energy threshold of speech silence segment; Step 6.3: Calculate segSNR: 。 5. The method for single-channel speech enhancement of a fiber microphone by integrating OMLSA and TQWT according to claim 1 is characterized in that: The step 7.1 includes the following steps: S7.1.1: Settings Q The maximum value that can be taken is Q max , r The maximum value that can be taken is r max ,initialization Q = 1, r Take [3, r max ] any value in the interval; S7.1.2: By Q and r Calculate the number of decomposition levels J 1. The calculation formula is: ; At this time Q Store values into a sequence Q _ cell In the current Q 、 r 、 J 1. Perform steps 7.2 to 7.8 for the three parameters and skip step 8. w { j } is the value in step 9 w 1{ j }, proceed to step 9; S7.1.3: Compute reconstruction error E = max(| x 1( n ) - x ( n )|), and store it in sequence E _ cell ; S7.1.4: Determine whether the following loop conditions are met: Q < Q max ; If satisfied, then Q Add 0.2 to the value and repeat S7.1.2 to S7.1.3; otherwise, proceed to S7.1.
5. S7.1.5: Minimum reconstruction error min( E _ cell )correspond Q _ cell The value in is Q opt ; S7.1.6: Initialization r = 3, Q Take [1, Q max ] any value in the interval; S7.1.7: Calculate the number of decomposition levels J 1. At this time r Store values into a sequence r _ cell In the current step, use the current parameters to proceed to steps 7.2 to 7.8, skipping step 8. w { j } is the value in step 9 w 1{ j }, proceed to step 9; S7.1.8: Compute reconstruction error E = max(| x 1( n ) - x ( n )|), and store it in sequence E _ cell_ 1; S7.1.9: Determine whether the following loop conditions are met: r < r max ; If satisfied, then r Add 0.2 to the value and repeat S7.1.7 to S7.1.8; otherwise, proceed to S7.1.
10. S7.1.10: Minimum reconstruction error min( E _ cell_ 1) Correspondence r _ cell The value in is r opt ; S7.1.11: Use Q opt and r opt Find the optimal number of decomposition levels J value.
6. The method for single-channel speech enhancement of a fiber microphone by integrating OMLSA and TQWT according to claim 1, characterized in that: The step 7.5 includes the following steps: S7.5.1: Obtain the j Layer low-pass subband sequence , k_low is the low-pass subband frequency index, 0≤ k_low ≤ : ; Where: Indicates the j The initial value of the layer low-pass subband sequence, Express x 1( n ) is subjected to unitary discrete Fourier transform to X 1( k )’s initial value; ; Where: Indicates the j Layer low-pass subband sequence from frequency 1 to frequency The value of express X 1( k ) from frequency 1 to frequency The value of ; Where: Indicates the j Layer low-pass subband sequence from frequency + 1 to frequency + The value of express X 1( k ) from the frequency + 1 to frequency + The value of Indicates that the values 1 to Execute as follows θ Operation: ; Where: D Indicates progress θ The parameter values of the operation; ; Where: Indicates the j Layer low-pass subband sequence at frequency The value of ; Where: Indicates the j Layer low-pass subband sequence from frequency To frequency The value of express X 1( k ) from the frequency To frequency The value of ; Where: A 1: B 1 means the sequence starts from A 1 index to B 1. ".*" means multiplying corresponding elements of the sequence; S7.5.2: Obtain the j Layer high-pass subband sequence , k_high is the high-pass subband frequency index, 0≤ k_high ≤ : ; Where: Indicates the j The initial value of the layer high-pass subband sequence; ; Where: Indicates the j Layer high-pass subband sequence from frequency 1 to frequency The value of express X 1( k ) from the frequency + 1 to frequency + The value of Indicates that the values are taken in sequence To 1 execution θ Operation; ; Where: Indicates the j Layer high-pass subband sequence from frequency To frequency The value of express X 1( k ) from the frequency To frequency The value of ; Where: Indicates the j Layer high-pass subband sequence at frequency The value of express X 1( k )exist k = N The value at / 2; ; Where: Indicates the j Layer high-pass subband sequence from frequency To frequency The value of express X 1( k ) from the frequency To frequency The value of ; Where: Indicates the j Layer high-pass subband sequence from frequency To frequency The value of express X 1( k ) from the frequency To frequency value.
7. The method for single-channel speech enhancement of a fiber microphone by integrating OMLSA and TQWT according to claim 1, characterized in that: In step 8, basis tracking is used to find w{ j }Sparseness is performed, including the following steps: Step 8.1: Initialization j =1; Step 8.2: Generate a l j × l j The identity matrix A j , l j For the j The length of the wavelet coefficients; vector b j = A j T * w { j }express b j Projection onto the identity matrix, where A j T for A j The transpose of B j = [ A j T * A j , - A j T * A j ; - A j T * A j , A j T * A j ] is used to construct quadratic terms in optimization problems; Step 8.3: Calculate the j Layer regularization parameter τ j , and its calculation formula is: ; in: To reflect the sparsity with the number of layers j The regulatory factors of change, For w { j } is the coefficient for amplitude adjustment, ln() is for taking the natural logarithm, and median() is for taking the median operation; Step 8.4: Construct a Column vector of c j , as the driving vector for sparse optimization, is generated by the following formula: ; in: C j Is the length of And a column vector of all 1s; Step 8.5: Solve the quadratic optimization problem ; in: z is the intermediate quantity when solving, and the length is Column vector of ; express z The transpose of express c j The transpose of When minimizing the objective function z The optimal solution of z j ; Step 8.6: By z j Obtain the sparse wavelet coefficients w 1{ j }, and its calculation formula is: ; Where: Represents a column vector z j Line 1 to l j The value of the row, Represents a column vector z j No. l j + 1 row to The value of the row; Step 8.7: Determine whether the following loop conditions are met: j ≤ J ; If satisfied, then j Add 1 to the value and repeat steps 8.2 to 8.6, otherwise proceed to step 9.
Citation Information
Patent Citations
Beam forming voice enhancing algorithm with invariant LCMV frequency based on log spectrum estimation
CN108922554A
Voice noise reduction method and device, computer equipment and storage medium
CN114242103A