Adaptive line spectrum enhancement method based on deep learning
By applying a fully connected neural network based on deep learning in the field of signal processing, replacing the iterative process of the traditional adaptive line spectrum enhancement method, and combining the noise suppression gate preprocessing method, the problem of poor processing effect of traditional methods at low signal-to-noise ratio is solved, and higher signal-to-noise ratio and processing gain is achieved.
Patent Information
- Application Number
- CN202510303567.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-14
Smart Images

Figure CN120199263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal processing, and in particular to an adaptive line spectrum enhancement method. Background Art
[0002] Conventional ALE utilizes the characteristic that the narrowband signal in the input signal and the narrowband signal in the delayed input signal remain correlated while the broadband noise is uncorrelated, and adaptively enhances the narrowband signal and suppresses the broadband noise.
[0003] x(nk)=[x(nk),x(nk-1),…x(nk-(L-1))] T
[0004] y(n)=w T (n)x(nk)
[0005] e(n)=x(n)-y(n)
[0006] w(n+1)=w(n)+μe(n)x(nk)
[0007] In each iteration of adaptive line spectrum enhancement, the delayed input signal x(nk) is multiplied by the weight coefficient w(n) to obtain the output signal at that moment, which is then subtracted from the reference signal at that moment (i.e., the undelayed input signal) to obtain the error signal, and the weight coefficient is then updated using the LMS algorithm. It can be seen that in each iteration, there is a delayed input signal x(nk) corresponding to an input signal x(n), and each delayed input signal corresponds to an output signal y(n), i.e., y(n) = f w (x(nk)), and then use the LMS algorithm to make the difference between x(n) and y(n) smaller and smaller (eliminating noise as much as possible and retaining the signal). Based on this principle, a neural network can be used to replace this process. The input signal and the delayed input signal are used to train the neural network. The neural network can also make the difference between the input signal x(n) and the output signal y(n) smaller and smaller, achieving the purpose of adaptively enhancing narrowband signals and suppressing noise. This is the basic idea of using neural networks for line spectrum enhancement. At the same time, neural networks have powerful nonlinear fitting capabilities, which can better fit the relationship between the input signal and the delayed input signal under low signal-to-noise ratio, and improve the processing effect of conventional ALE under low signal-to-noise ratio.
[0008] The processing effect of the conventional ALE method deteriorates significantly when the signal-to-noise ratio decreases, so it is necessary to further improve the processing gain of the conventional ALE method. Specifically, when the signal is interfered by strong noise, the traditional ALE method cannot accurately identify and enhance the target signal, resulting in a decrease in signal quality, which affects subsequent signal processing and application. Summary of the invention
[0009] To solve the above technical problems, the present invention proposes an adaptive line spectrum enhancement method based on deep learning, belonging to the field of signal processing. By means of deep learning, the present invention replaces the iterative process of traditional adaptive line spectrum enhancement, realizes the iteration of weights under low signal-to-noise ratio, and finally improves the signal-to-noise ratio of line spectrum signals. At the same time, preprocessing means are added before the fully connected neural network to enhance the line spectrum enhancement effect. Even under low signal-to-noise ratio, the present invention can accurately identify and enhance the target signal, improving the processing effect of the traditional adaptive line spectrum enhancement (ALE) method under low signal-to-noise ratio.
[0010] An adaptive line spectrum enhancement method based on deep learning, the adaptive line spectrum enhancement method specifically includes the following steps:
[0011] Step 1, generating target acoustic signal sample data;
[0012] Step 2, normalizing the target acoustic signal sample data;
[0013] Step 3, constructing a data set; the data set includes a training data set and a test data set;
[0014] Step 4, constructing a fully connected neural network model;
[0015] Step 5, training the fully connected neural network model;
[0016] Step 6, using the fully connected neural network model to verify the effect in actual signal processing.
[0017] Further, in step 1, the target acoustic signal sample data is generated by means of simulation or sonar. The target acoustic signal collected by means of simulation or sonar can ensure the diversity and authenticity of the target acoustic signal;
[0018] The data tool used in the simulation method is MATALB, and the specific simulation settings of MATALB are as follows:
[0019] First, the frequency, duration and power of the target acoustic signal are preset; then, a white noise background is preset to simulate the real environment; then, the bandwidth and power of the noise signal are set;
[0020] The target acoustic signal is denoted as x(n), and the target acoustic signal x(n) includes a line spectrum signal S(n) and a noise signal N noise (n);
[0021]
[0022] where A is the amplitude of the line spectrum signal; f0 is the frequency of the line spectrum signal; T is the total duration of the line spectrum signal; f s is the sampling frequency; n represents the nth sampling; is the random phase of the line spectrum signal, and the random phase satisfies the distribution U(0, 2π).
[0023] Furthermore, in step 2, the normalization processing method of the target acoustic signal sample data is as follows:
[0024] To ensure that the sample data has a consistent convergence speed during training, the min-max normalization method is used to normalize the target acoustic signal sample data. After normalization, the target acoustic signal is as follows:
[0025]
[0026] where min x and max x represent the minimum value and the maximum value of the target acoustic signal sequence respectively.
[0027] Furthermore, in step 3, the dataset is constructed as follows:
[0028] The normalized target acoustic signal sample data is used as the dataset, and the dataset includes a training set and a test set; take the first T s moments of the normalized target acoustic signal, that is, the signal with the number of sampling points of T s *f s as the training set data x T (n), and the signal of the remaining sampling points is the test set data x E (n).
[0029] Furthermore, in step 4, the fully connected neural network model includes an input layer, a hidden layer, and an output layer; the fully connected neural network model is used to replace the adaptive weight iteration process of ALE;
[0030] Furthermore, the input layer is the delay vector x T (n - k) of the training set data:
[0031] x T (n - k) = [x T (n - k), x T (n - k - 1), …, x T (n - k - (N - 1)] T ,
[0032] where k represents the number of pre-delay points, the pre-delay is set to 1s, that is, the number of digital sampling points of 1*f s [·] T is the transpose of the matrix, and N is the order of the input layer;
[0033] Using the training set data x T (n) and the delay vector x T(n - k) is used as the input of the fully - connected neural network model; among which, the data corresponding to the sampling points in the training set is the output signal of the fully - connected neural network model.
[0034] Furthermore, in step 5, the training process of the fully - connected neural network model includes:
[0035] Step 5.1, conduct preliminary training on the fully - connected neural network model:
[0036] Input the training set data x T (n) directly into the fully - connected neural network model, then the output signal y(n) of the fully - connected neural network model is:
[0037] y(n) = f θ (x T (n))
[0038] where f θ (·) represents the function mapping relationship of the fully - connected neural network model;
[0039] Obtain the parameters of the trained fully - connected neural network model
[0040]
[0041] where N represents the order of the input layer; i is the sampling sequence number; x(i) is the input acoustic signal at the i - th sampling; y(i) is the output acoustic signal at the i - th sampling; L is the loss function; θ is the parameter of the fully - connected neural network model;
[0042] The loss function L is:
[0043] L(x, y)=||x - y|| 2
[0044] where ‖‖·‖‖ represents the squared - error operator; x and y are the input acoustic signal and the output acoustic signal of sampling respectively;
[0045] Step 5.2, use the Adam algorithm to optimize the parameter θ of the fully - connected neural network model obtained in step 5.1 * for optimization:
[0046]
[0047] In the Adam algorithm, the learning rate α is taken as 0.001 by using the gradient - descent method to solve, β′ and β″ are the first - order moment and the second - order moment of the initial estimated gradient respectively, and the first - order moment and the second - order moment are taken as 0.9 and 0.999 respectively; the Adam algorithm can adaptively adjust the learning rate, does not require manual adjustment, and can effectively solve the problem of gradient noise, and at the same time converges quickly and achieves good results;
[0048] The optimization process of the Adam algorithm for the parameters θ of the fully connected neural network model * is as follows:
[0049] During the optimization process, the subscripts of letters all represent the number of current iterations;
[0050] Step 5.2.1, Given initial parameters: Given the learning rate α; Given the parameters β′, β″; The parameters β′, β″ are used to estimate the first moment and the second moment of the gradient; Define the optimization objective function f(θ) and the initial value θ0 of the optimization parameters;
[0051] Step 5.2.2, Set the initial values t = 0, m0 = 0, v0 = 0, where t is the number of iterations, m is the estimate of the first moment of the gradient, and v is the estimate of the second moment;
[0052] Step 5.2.3, Update the time t + 1 → t, and calculate the gradient g of the objective function t ; The gradient g of the objective function t is expressed as: where the update of the estimate m of the first moment of the gradient t is as follows: m t = β′m t-1 +(1 - β′)g t , calculate the bias-corrected estimate of the first moment of the gradient Update the estimate of the second moment of the gradient, and calculate the bias-corrected estimate of the second moment of the gradient
[0053] v t = β″v t-1 +(1 - β″)g t +g t
[0054]
[0055] Update the parameters:
[0056]
[0057] where ε = 10 -8 is an infinitesimal quantity to ensure that the denominator is not zero;
[0058] Step 5.2.4, Repeat steps 5.2.1 to 5.2.4 for iteration. When the Adam algorithm fully traverses the database more than 40 times, the optimal θ is finally obtained * ;
[0059] Step 5.3, In the training of the fully connected neural network model, add activation functions between the input layer, hidden layer, and output layer of the fully connected neural network model to ensure the non-linear transformation of the fully connected neural network model:
[0060] The rectified linear activation function is used as the activation function of the fully connected neural network model. The rectified linear activation function ReLU(z) is as follows:
[0061]
[0062] where z is the calculation result of the fully connected neural network model in the forward propagation process;
[0063] When the input of the rectified linear activation function is less than 0, the output of the fully connected neural network model is 0; the ReLU function obtains a sparse representation of data, enabling the fully connected neural network model to accurately mine the relevant line spectral frequencies in the acoustic signal; the ReLU function is more conducive to the rapid and accurate descent of the gradient.
[0064] Furthermore, in step 6, the effect verification process of the fully connected neural network model in the actual signal processing task is as follows:
[0065] Step 6.1, test the trained fully connected neural network model:
[0066] Input the test set data into the fully connected neural network model for adaptive line spectrum enhancement; the method of using the fully connected neural network model for adaptive line spectrum enhancement is DLE (Deep learning Line Enhancement: DLE), and the enhancement result of DLE is the output result x DLE (n) is:
[0067]
[0068] Step 6.2, preprocess the fully connected neural network model after testing to improve the signal-to-noise ratio:
[0069] Use the noise suppression gate (NSG) method as the preprocessing means before inputting into the fully connected neural network model. The noise suppression gate method is as follows:
[0070] The autocorrelation function T E (n) of the test set data signal x x is:
[0071] T x (n) = x E (-n) * x E (n)
[0072] The noise suppression gate is I(n) is:
[0073]
[0074] where αT is the zeroing time range, n is the nth sampling, and N s is the total number of sampling points;
[0075] The amplitude equalization window E(n) is:
[0076]
[0077] where A is the amplitude of the test set data signal x(n);
[0078] The preprocessed test set data x NSG (n) is:
[0079] x NSG (n) = T x (n)·I(n)·E(n)
[0080] Step 6.3, apply the preprocessed fully connected neural network model to actual signal processing.
[0081] The beneficial effects of the present invention are as follows:
[0082] (1) Use a fully connected neural network model to replace the iterative process of ALE, and can better handle complex acoustic environment changes through the powerful nonlinear fitting ability of the neural network;
[0083] (2) For different types of noise and signal characteristics, targeted optimization is carried out through a large amount of data training, rather than being limited to ALE that needs to readjust parameters in a new environment;
[0084] (3) Better realize the iteration of weights under low signal-to-noise ratio. The processing gain of the traditional adaptive line spectrum enhancement method is only 0.31 dB, while the processing gain of the adaptive line spectrum enhancement method based on deep learning is 9.38 dB, greatly improving the signal processing gain;
[0085] (4) Through a large amount of data training, iterate the environment-related features, which is different from ALE relying on manually designed algorithms and feature extraction methods;
[0086] (5) For the weights under a specific low signal-to-noise ratio, the processing gain has a relatively large improvement under the same input conditions. Brief Description of the Drawings
[0087] Figure 1 is the flowchart of the adaptive line spectrum enhancement method based on deep learning;
[0088] Figure 2 is the structure diagram of the fully connected neural network model;
[0089] Figure 3 is the algorithm block diagram of the adaptive line spectrum enhancement method based on deep learning;
[0090] Figure 4 Time-frequency diagram for the analysis of the simulated signal with SNR = 7 dB;
[0091] Figure 4 In, Figure (a) is the time-frequency diagram of the signal FFT (Fast Fourier Transform) method with SNR = 7 dB, Figure (b) is the time-frequency diagram of the ALE (Adaptive Line Enhancement) method, Figure (c) is the time-frequency diagram of the method DLE (Deep Learning-based Line Enhancement) proposed by the present invention, and Figure (d) is the time-frequency diagram of the method NSG-DLE (Deep Learning-based Line Enhancement through Noise Suppression Gate) proposed by the present invention;
[0092] Figure 5 Time-frequency diagram for the analysis of the simulated signal with SNR = 3 dB;
[0093] Figure 5 In, Figure (a) is the time-frequency diagram of the signal FFT (Fast Fourier Transform) method with SNR = 3 dB, Figure (b) is the time-frequency diagram of the ALE (Adaptive Line Enhancement) method, Figure (c) is the time-frequency diagram of the method DLE (Deep Learning-based Line Enhancement) proposed by the present invention, and Figure (d) is the time-frequency diagram of the method NSG-DLE (Deep Learning-based Line Enhancement through Noise Suppression Gate) proposed by the present invention;
[0094] Figure 6 Gain calculation results for input signals with different signal-to-noise ratios;
[0095] Figure 6 In, (a) is the output signal-to-noise ratio of the fully connected neural network model; (b) is the gain of the fully connected neural network model;
[0096] Figure 7 Time-frequency diagram of the experimental acquisition data processing results;
[0097] Figure 7 In, Figure (a) is the time-frequency diagram of the measured signal FFT (Fast Fourier Transform) method, Figure (b) is the time-frequency diagram of the ALE (Adaptive Line Enhancement) method, Figure (c) is the time-frequency diagram of the method DLE (Deep Learning-based Line Enhancement) proposed by the present invention, and Figure (d) is the time-frequency diagram of the method NSG-DLE (Deep Learning-based Line Enhancement through Noise Suppression Gate) proposed by the present invention. Detailed implementation manners
[0098] The mathematical software includes all software similar to MATALB that can implement the function of generating simulated signals.
[0099] To make up for the deficiencies of simulation data, the underwater acoustic signal detection field prefers to use measured data sets; by setting up experimental scenarios, simulating the emission of line spectrum signals, and using acoustic sensors for data acquisition; the simulation environment is carried out in a laboratory environment or a real water area to ensure that the data can cover a variety of underwater sound targets, including fish, whales, submarines, and ships; different noise backgrounds include ocean ambient noise and biological noise.
[0100] An adaptive line spectrum enhancement method based on deep learning belongs to the field of signal processing. In the present invention, the iterative process of traditional adaptive line spectrum enhancement is replaced by a deep learning method, realizing the iteration of weights under low signal-to-noise ratio, and finally improving the signal-to-noise ratio of line spectrum signals; at the same time, a preprocessing means is added before the fully connected neural network to enhance the line spectrum enhancement effect; even in the case of low signal-to-noise ratio, the present invention can accurately identify and enhance the target signal, improving the processing effect of the traditional adaptive line spectrum enhancement (ALE) method under low signal-to-noise ratio.
[0101] As Figure 1 shown, an adaptive line spectrum enhancement method based on deep learning, the adaptive line spectrum enhancement method specifically includes the following steps:
[0102] Step 1, generating target acoustic signal sample data;
[0103] Step 2, normalizing the target acoustic signal sample data;
[0104] Step 3, constructing a data set; the data set includes a training data set and a test data set;
[0105] Step 4, constructing a fully connected neural network model;
[0106] Step 5, training the fully connected neural network model;
[0107] Step 6, using the fully connected neural network model to verify the effect in actual signal processing.
[0108] In step 1, the target acoustic signal sample data is generated by means of simulation or sonar, and the target acoustic signal collected by the simulation or sonar method can ensure the diversity and authenticity of the target acoustic signal;
[0109] The data tool used in the simulation method is MATALB, and the specific simulation settings of MATALB are as follows:
[0110] First, preset the frequency, duration, and power of the target acoustic signal; then, preset a white noise background to simulate the real environment; then, set the bandwidth and power of the noise signal;
[0111] The signal frequency is 300 Hz, the sampling frequency is 2 kHz, the broadband noise range is [100, 500], and the signal duration is 5 s. Design signal-to-noise ratios of 7 dB and 3 dB respectively, and use experimental sampling data;
[0112] The target acoustic signal is denoted as x(n), and the target acoustic signal x(n) includes a line spectrum signal S(n) and a noise signal N noise (n);
[0113]
[0114] Among them, A is the amplitude of the line spectrum signal; f0 is the frequency of the line spectrum signal; T is the total duration of the line spectrum signal; f s is the sampling frequency; n represents the nth sampling; is the random phase of the line spectrum signal, and the random phase satisfies the distribution U(0, 2π).
[0115] In step 2, the normalization processing method of the target acoustic signal sample data is as follows:
[0116] To ensure that the sample data has a consistent convergence speed, the minimum-maximum normalization method is used to normalize the target acoustic signal sample data. The normalized target acoustic signal is as follows:
[0117]
[0118] Among them, min x and max x respectively represent the minimum and maximum values of the target acoustic signal sequence.
[0119] In step 3, the dataset is constructed as follows:
[0120] Take the normalized target acoustic signal sample data as the dataset, and the dataset includes a training set and a test set; Take the signal of the first T s moments of the normalized target acoustic signal, that is, the signal of the number of sampling points of T s *f s as the training set data x T (n), and the signals of the remaining sampling points are the test set data x E (n).
[0121] As Figure 2 shown, in step 4, the fully connected neural network model includes an input layer, a hidden layer, and an output layer; In this example, the number of neurons in each layer is taken as 2000, 800, and 1 respectively;
[0122] Use the fully connected neural network model to replace the adaptive weight iteration process of ALE; The target acoustic signal x T(n) Serve as the input to the fully connected neural network model;
[0123] The input vector of the input layer:
[0124] x T (n - k) = [x T (n - k), x T (n - k - 1), …, x T (n - k - (N - 1)] T ,
[0125] where x T (n - k) represents the delay of the training set data, where k represents the number of pre - delay points, and the pre - delay is set to 1s, that is, 1 * f s the number of digital sampling points, [·] T is the transpose of the matrix, and N is the order of the input layer;
[0126] Take the trained set data x T (n) and the delay vector x T (n - k) of the training set data as the input to the fully connected neural network model; where the data corresponding to the sampling points of the training set data is the output signal of the fully connected neural network model.
[0127] As Figure 3 shown, in step 5, the training process of the fully connected neural network model includes:
[0128] Step 5.1, conduct preliminary training on the fully connected neural network model:
[0129] Directly input the training set data x T (n) into the fully connected neural network model, then the output signal y(n) of the fully connected neural network model is:
[0130] y(n) = f θ (x T (n))
[0131] where f θ (·) represents the function mapping relationship of the fully connected neural network model;
[0132] Obtain the parameters of the trained fully connected neural network model
[0133]
[0134] where N represents the order of the input layer; i is the sampling sequence number; x(i) is the input sound signal for the i - th sampling; y(i) is the output sound signal for the i - th sampling; L is the loss function; θ is the parameter of the fully connected neural network model;
[0135] The loss function L is as follows:
[0136] L(x,y) = ||x - y|| 2
[0137] where ||·|| represents the squared error operator; x and y are the sampled input acoustic signal and the sampled output acoustic signal respectively;
[0138] Step 5.2, use the Adam algorithm to optimize the parameters θ of the fully connected neural network model obtained in Step 5.1 * as follows:
[0139]
[0140] In the Adam algorithm, the learning rate α is taken as 0.001 by using the gradient descent method to solve. β′ and β″ are the first moment and the second moment of the initial estimated gradient respectively, and the first moment and the second moment are taken as 0.9 and 0.999 respectively; the Adam algorithm can adaptively adjust the learning rate without manual adjustment, and can effectively solve the problem of gradient noise, and at the same time converge quickly to achieve good results;
[0141] The optimization process of the Adam algorithm for the parameters θ of the fully connected neural network model * is as follows:
[0142] In the optimization process, the subscripts of letters all represent the current iteration number;
[0143] First, given the initial parameters: given the learning rate α; given the parameters β′, β″; the parameters β′, β″ are used to estimate the first moment and the second moment of the gradient; given the optimization objective function f(θ) and the initial value θ0 of the optimization parameters;
[0144] Second, set the initial values t = 0, m0 = 0, v0 = 0, where t is the iteration number, m is the first moment estimate of the gradient, and v is the second moment estimate;
[0145] Then, update the time t + 1 → t, and calculate the gradient g of the objective function t ; the gradient g of the objective function t is expressed as: where the first moment estimate m of the updated gradient t is as follows: m t = β′m t-1 + (1 - β′)g t , calculate the first moment estimate of the gradient after bias correction Update the second moment estimate of the gradient, and calculate the second moment estimate of the gradient after bias correction
[0146] v t = β″v t-1 + (1 - β″)gt +g t
[0147]
[0148] Update parameter:
[0149]
[0150] where ε = 10 -8 is an infinitesimal to ensure that the denominator is not zero;
[0151] Finally, repeat the above iterative process to finally obtain the optimal θ *
[0152] Step 5.3, in the training of the fully connected neural network model, add activation functions between the input layer, hidden layer, and output layer of the fully connected neural network model to ensure the non-linear transformation of the fully connected neural network model:
[0153] Adopt the rectified linear activation function as the activation function of the fully connected neural network model. The rectified linear activation function ReLU(z) is:
[0154]
[0155] where z is the calculation result in the forward propagation process of the fully connected neural network model;
[0156] When the input of the rectified linear activation function is less than 0, the output of the fully connected neural network model is 0; the ReLU function obtains a sparse representation of a data, enabling the fully connected neural network model to accurately mine the relevant line spectral frequencies in the acoustic signal; the ReLU function is more conducive to the rapid and accurate descent of the gradient.
[0157] In step 6, the effect verification process of the fully connected neural network model in the actual signal processing task is as follows:
[0158] Step 6.1, test the trained fully connected neural network model:
[0159] Input the test set data into the fully connected neural network model for adaptive line spectrum enhancement; the method of using the fully connected neural network model for adaptive line spectrum enhancement is DLE (Deep learning Line Enhancement: DLE), and the enhancement result of DLE is the output result x DLE (n) is:
[0160]
[0161] Step 6.2, preprocess the tested fully connected neural network model to improve the signal-to-noise ratio:
[0162] Use the Noise Suppression Gate (NSG) method as the preprocessing means before inputting the fully connected neural network model. The noise suppression gate method is as follows:
[0163] The autocorrelation function T E (n) of the test set data signal x x is:
[0164] T x (n) = x E (-n) * x E (n)
[0165] The noise suppression gate I(n) is:
[0166]
[0167] where α T is the zeroing time range, n is the nth sampling, and N s is the total number of sampling points;
[0168] The amplitude equalization window E(n) is:
[0169]
[0170] where A is the amplitude of the test set data signal x(n);
[0171] The preprocessed test set data x NSG (n) is:
[0172] x NSG (n) = T x (n) · I(n) · E(n)
[0173] Step 6.3, apply the preprocessed fully connected neural network model to actual signal processing.
[0174] Input the original data and the preprocessed data passing through the noise suppression gate into the above-trained model respectively, and the following effects can be obtained:
[0175] Calculate the processing gain of each method using FFT (Fast Fourier Transform), ALE (Adaptive Line Enhancement), the proposed DLE (Deep Learning Line Enhancement) in the present invention, and NSG-DLE (Line Enhancement through Noise Suppression Gate first and then Deep Learning) respectively. When the input signal-to-noise ratio is 7 dB, the output results are as Figure 4As shown in (a)-(c), the output signal-to-noise ratios (SNRs) and gain calculation results of the conventional adaptive line spectrum enhancement method and the deep learning-based adaptive line spectrum enhancement method are shown in Table 1, where the gain is defined as the output SNR of different methods minus the output SNR of the FFT result.
[0176] Table 1 Gain calculation results of various line spectrum enhancement methods when SNR = 7 dB
[0177]
[0178] Similarly, when the input SNR is changed to 3 dB, as Figure 5 shown in (a)-(c), the output SNR and gain calculation results are also shown in Table 2:
[0179] Table 2 Gain calculation results of various line spectrum enhancement methods when SNR = 3 dB
[0180]
[0181] Figure 6 The curves of the output SNR and gain of various methods changing with the input SNR are given. The gain calculation is based on the FFT method. The input SNR increases from -15 dB to 15 dB at intervals of 2 dB, and the number of Monte Carlo experiments is 100 times. The curves of the output SNR and gain changes are as Figure 6 shown.
[0182] The deep learning-based adaptive line spectrum enhancement method proposed in the present invention is verified using experimental data.
[0183] The frequency of the transmitted signal is 170 Hz, the sampling frequency of the system is 8192 Hz, the signal duration is 30 s, the number of taps of the adaptive line spectrum enhancer is 1000, the step size is 2×10 -6 , the number of pre-delayed digital sampling points is 20, the number of neurons in the three layers is 1000, 300, and 1 respectively, the input vector dimension of the neural network model is 1000, the length of each frame of the time-frequency diagram is 2 s, and the overlap rate between frames is 50%. As Figure 7 shown in (a)-(c), the following table gives the output SNR and gain obtained by various line spectrum enhancement methods for processing experimental data
[0184] Table 3 Gain calculation results of various line spectrum enhancement methods for processing experimental data
[0185]
[0186]
[0187] The processing gain of ALE is 0.31 dB, and its processing effect is poor at low signal-to-noise ratios. However, the processing gains of the DLE method and NSG-DLE proposed in the present invention are 9.38 dB and 20.85 dB, respectively, which have relatively high processing gains, verifying the correctness of the simulation analysis and demonstrating good high-gain processing effects.
Claims
1. An adaptive line spectrum enhancement method based on deep learning, characterized in that: The adaptive line spectrum enhancement method specifically comprises the following steps: Step 1, generating target sound signal sample data; Step 2: normalizing the target sound signal sample data; Step 3, construct a data set; the data set includes a training data set and a test data set; Step 4, construct a fully connected neural network model; Step 5, training a fully connected neural network model; Step 6: Use the fully connected neural network model to verify the effect in actual signal processing.
2. The method for adaptive line spectrum enhancement based on deep learning according to claim 1, characterized in that: In step 1, target acoustic signal sample data is generated by simulation or sonar; the data tool used in the simulation method is MATALB, and the specific simulation settings of MATALB are as follows: First, the frequency, duration and power of the target sound signal are pre-set; then, a white noise background is preset to simulate the real environment; then, the bandwidth and power of the noise signal are set; The target sound signal is recorded as x(n), and the target sound signal x(n) includes a line spectrum signal S(n) and a noise signal N noise (n); Where A is the amplitude of the line spectrum signal; f0 is the frequency of the line spectrum signal; T is the total duration of the line spectrum signal; f s is the sampling frequency; n represents the nth sampling; is the random phase of the line spectrum signal, and the random phase satisfies the distribution U(0,2π).
3. The method for adaptive line spectrum enhancement based on deep learning according to claim 1, characterized in that: In step 2, the normalization processing method of the target acoustic signal sample data is as follows: The minimum-maximum normalization method is used to normalize the target sound signal sample data. The target sound signal after normalization is as follows: Among them, min x and max x Respectively represent the minimum and maximum values of the target acoustic signal sequence.
4. The method for adaptive line spectrum enhancement based on deep learning according to claim 1, characterized in that: In step 3, the data set is constructed as follows: The normalized target acoustic signal sample data is used as a data set, which includes a training set and a test set. The first T s Time, i.e. T s *f s The signal of the number of sampling points is used as the training set data x T (n), the signal of the remaining sampling points is the test set data x E (n).
5. The method for adaptive line spectrum enhancement based on deep learning according to claim 1, characterized in that: In step 4, the fully connected neural network model includes an input layer, a hidden layer and an output layer; the target sound signal x T (n) serving as input for a fully connected neural network model; The input layer is the delayed vector x of the training set data. T (nk): x T (n-k)=[x T (n-k),x T (n-k-1),…,x T (n-k-(N-1)] T , Where k represents the number of pre-delay points, and the pre-delay is set to 1s, that is, 1*f s The number of digital sampling points, [·] T is the transpose of the matrix, N is the order of the input layer; The training set data x T (n) and the delay vector x of the training set data T (nk) is used as the input of the fully connected neural network model; the data of the sampling points corresponding to the training set data is the output signal of the fully connected neural network model.
6. The method for adaptive line spectrum enhancement based on deep learning according to claim 1, characterized in that: In step 5, the training process of the fully connected neural network model includes: Step 5.1, perform preliminary training on the fully connected neural network model: The training set data x T (n) is directly input into the fully connected neural network model, then the output signal y(n) of the fully connected neural network model is: y(n)=f θ (x T (n)) Among them, f θ (·) represents the function mapping relationship of the fully connected neural network model; Get the trained fully connected neural network model parameters Where N is the order of the input layer; i is the sampling number; x(i) is the i-th sampled input sound signal; y(i) is the i-th sampled output sound signal; L is the loss function; θ is the fully connected neural network model parameter; The loss function L is: L(x,y)=||xy|| 2 Wherein, ||·|| represents the square error operator; x and y are the sampled input acoustic signal and the sampled output acoustic signal respectively; Step 5.2: Use the Adam algorithm to calculate the fully connected neural network model parameters θ obtained in step 5.
1. * To optimize: In the Adam algorithm, the gradient descent method is used to solve the learning rate α, which is 0.
001. β′ and β″ are the first-order moment and second-order moment of the initial estimated gradient, respectively. The first-order moment and the second-order moment are 0.9 and 0.999, respectively. The Adam algorithm can adaptively adjust the learning rate without manual adjustment, and can effectively solve the problem of gradient noise, while converging quickly and achieving good results. Adam algorithm adjusts the fully connected neural network model parameters θ * The optimization process is as follows: Step 5.2.1, given initial parameters: given learning rate α; given parameters β′, β″; parameters β′, β″ are used to estimate the first-order moment and second-order moment of the gradient; set the optimization objective function f(θ) and the initial value of the optimization parameter θ0; Step 5.2.2, set the initial values t = 0, m0 = 0, v0 = 0, where t is the number of iterations, m is the first-order moment estimate of the gradient, and v is the second-order moment estimate; Step 5.2.3, update time t+1→t and calculate the objective function gradient g t ; Objective function gradient g t It is expressed as: Among them, the first-order moment estimate m of the updated gradient t As follows: t =β′m t-1 +(1-β′)g t , compute the bias-corrected first-order moment estimate of the gradient Update the second-order moment estimate of the gradient and calculate the bias-corrected second-order moment estimate of the gradient v t =β″v t-1 +(1-β″)g t +g t Update parameters: Where ε = 10 -8 is an infinitesimal quantity to ensure that the denominator is not zero; Step 5.2.4, repeat steps 5.2.1 to 5.2.4 for iteration. When the Adam algorithm completely traverses the database more than 40 times, the optimal θ is finally obtained. * ; Step 5.3, in the training of the fully connected neural network model, an activation function is added between the input layer, hidden layer and output layer of the fully connected neural network model to ensure the nonlinear transformation of the fully connected neural network model: The rectified linear activation function is used as the activation function of the fully connected neural network model. The rectified linear activation function ReLU(z) is: Among them, z is the calculation result of the fully connected neural network model in the forward propagation process; When the input of the modified linear activation function is less than 0, the output of the fully connected neural network model is 0; the ReLU function obtains a sparse representation of the data, allowing the fully connected neural network model to accurately mine the relevant line spectrum frequencies in the acoustic signal; the ReLU function is more conducive to the rapid and accurate descent of the gradient.
7. The method for adaptive line spectrum enhancement based on deep learning according to claim 1, characterized in that: In step 6, the effect verification process of the fully connected neural network model in the actual signal processing task is as follows: Step 6.1, test the trained fully connected neural network model: The test set data is input into the fully connected neural network model for adaptive line spectrum enhancement; The enhanced result of adaptive line spectrum enhancement is the output result x of the fully connected neural network model. DLE (n) is: Step 6.2, preprocess the tested fully connected neural network model to improve the signal-to-noise ratio: The noise suppression gate method is used as a preprocessing method before inputting the fully connected neural network model. The noise suppression gate method is as follows: Test set data signal x E The autocorrelation function T of (n) x (n) is: T x (n)=x E (-n)*x E (n) The noise suppression gate I(n) is: Among them, α T is the zeroing time range, n is the nth sampling, N s is the total number of sampling points; The amplitude equalization window E(n) is: Where A is the amplitude of the test set data signal x(n); Preprocessed test set data x NSG (n) is: x NSG (n)=T x (n)·I(n)·E(n) Step 6.3, apply the preprocessed fully connected neural network model to actual signal processing.
8. An electronic device, characterized in that: include: one or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, and the program codes can be called by a processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Noise classifying method based on convolutional neural network
CN110164472A
Radiation noise line spectrum frequency domain adaptive enhancement method
CN113343914A
Reconvolution recurrent neural network single-channel speech enhancement method based on masking effect
CN114999510A
Cited By
Low signal-to-noise ratio underwater acoustic line spectrum enhancement system based on deep learning
CN121636928A
Low snr underwater acoustic spectrum enhancement system based on deep learning
CN121636928B