Unmanned aerial vehicle speech enhancement method and system based on harmonic characteristics and deep learning
By extracting the periodic harmonic characteristics of drone rotor self-noise and predicting noise changes in deep learning models and dynamically adjusting parameters, the problem that traditional technology cannot effectively deal with drone rotor self-noise, achieving efficient voice enhancement effect.
Patent Information
- Application Number
- CN202510476287.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The prior art is difficult to effectively handle the self-noise of the drone rotor, especially during flight. The self-noise is dynamic, and the traditional deep learning voice enhancement system cannot perceive and process the self-noise characteristics of the rotor unmanned aerial vehicle.
By extracting the periodic harmonic characteristics of the drone rotor self-noise, and combining with the deep learning model to predict the noise change trend, a deep learning model is constructed to dynamically analyze the periodic harmonic characteristics of the drone self-noise, and dynamically adjust the parameters to achieve deep learning speech enhancement of the self-noise characteristic perception of the rotor UAV.
The precise separation and processing of the drone rotor self-noise is realized, which significantly improves the voice signal-to-noise ratio, reduces voice distortion, and makes the enhanced voice clearer and more understandable.
Smart Images

Figure CN120148540A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech signal processing, and particularly relates to a method and system for enhancing the speech of an unmanned aerial vehicle (UAV) based on harmonic features and deep learning. Background Art
[0002] In recent years, the UAV equipped with a camera has provided a wide range of applications for users in visual applications. However, in some special scenarios, visual processing cannot meet the needs of users. For example, in rescue operations in bad weather, due to weather reasons, the camera carried by the UAV cannot clearly capture the rescue situation. Therefore, a microphone array can be installed on the UAV so that the UAV has both hearing in addition to vision to solve some situations that cannot be satisfied by visual processing.
[0003] Currently, most of the methods adopted have limitations. For example, DJI (Dajiang) and Parrot improve the rotor design (such as shape and material) and motor structure to reduce noise at the source, but the noise cancellation effect is limited. The software-based noise cancellation methods provided by audio software processing companies such as iZotope and Adobe have limited support for real-time processing. The noise processing based on traditional signal processing technologies by MIT, Stanford, etc. has limited effects on non-linear noise. At the same time, with the development of deep learning, there are many speech enhancement systems based on deep learning, which basically process static noise. However, since the self-noise of the UAV is dynamic during flight, there is no deep learning speech enhancement system for perceiving the self-noise characteristics of rotor UAVs.
[0004] The present invention processes the harmonic characteristics of the UAV's self-noise to guide deep learning, constructs a deep learning model to dynamically analyze the periodic harmonic characteristics of the UAV's self-noise, and dynamically adjusts parameters according to the analysis results to achieve deep learning speech enhancement based on the perception of the self-noise characteristics of rotor UAVs. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a method and system for enhancing the speech of an unmanned aerial vehicle based on harmonic features and deep learning, and solve the problems mentioned in the background art.
[0006] To achieve the above object, the present invention provides the following technical solution: A method for enhancing the speech of an unmanned aerial vehicle based on periodic harmonic features and deep learning prediction, comprising the following steps:
[0007] Step S1: Collect the self-noise signal of the rotor of the UAV, that is, the original sampling frequency.
[0008] Step S2: Process the original sampling frequency to obtain a new sampling frequency, filter the new sampling frequency to obtain a normalized signal, extract the normalized signal to obtain the normalized signal spectrum, use the peak detection algorithm to find the peaks in the normalized signal spectrum, take the peaks as the fundamental frequency, extract the harmonic components from the normalized signal spectrum, and arrange the fundamental frequency and harmonic components in sequence to form a feature vector;
[0009] Step S3: Process the feature vector to generate a noise dynamic change label, combine the feature vector and the noise dynamic change label into a data set, construct a deep learning model, input the data set into the deep learning model for processing, and output the prediction result of the noise dynamic change label;
[0010] Step S4: Extract the rotor self-noise signal of the drone to obtain the noisy speech amplitude spectrum, calculate the energy distribution of the rotor self-noise signal of the drone using the minimum statistic method according to the prediction result of the noise dynamic change label, smooth the noisy speech amplitude spectrum to obtain the smoothed spectrum, process the smoothed spectrum to obtain the spectral characteristics of the noise, subtract the spectral characteristics of the noise from the noisy speech amplitude spectrum to obtain the enhanced amplitude spectrum, process the noisy speech amplitude spectrum to obtain the phase information, combine the enhanced amplitude spectrum and the phase information to construct the enhanced spectrum, and process the enhanced spectrum to obtain the enhanced signal;
[0011] Step S5: Output the enhanced signal.
[0012] Furthermore, in step S1, the process of collecting the rotor self-noise signal of the drone is as follows:
[0013] A high-sensitivity microphone array is mounted on the drone. The high-sensitivity microphone array is connected to a high-performance digital-to-analog converter. The rotor self-noise signal of the drone is collected through the high-sensitivity microphone array. The collected rotor self-noise signal of the drone is an analog signal, and the analog signal is converted into a digital signal, which is the original sampling frequency.
[0014] Furthermore, the process of the feature vector in step S2 is as follows:
[0015] Downsample the original sampling frequency to obtain a new sampling frequency. Before downsampling, use a low-pass filter to remove the frequency components higher than half of the new sampling frequency. Take the smallest integer of the original sampling frequency and the new sampling frequency. Set the downsampling factor M to 3, which means:
[0016] M = foriginal / fnew = 3(1);
[0017] Where, foriginal represents the original sampling frequency; fnew represents the new sampling frequency; the Butterworth filter is used to filter the high-frequency noise and low-frequency interference in the new sampling frequency to obtain a filtered signal, and the order n of the Butterworth filter is 4, which means:
[0018]
[0019] Where, f represents the frequency; H(f) represents the frequency response function of the Butterworth filter; fc represents the cut-off frequency of the Butterworth filter; calculate the maximum absolute value |Xmax| of the filtered signal, and divide all the filtered signals by the maximum absolute value |Xmax| of the filtered signal to reach the fixed range [-1, 1] of signal amplitude normalization, which means:
[0020] Xnormalized = X / |Xmax| (3);
[0021] Where, Xnormalized represents the normalized signal, X represents the rotor self-noise signal of the drone; Xmax represents the maximum absolute value of the signal;
[0022] Perform short-time Fourier transform on each frame of the normalized signal to obtain the normalized signal spectrum, which means:
[0023]
[0024] Where, X(f, t) represents the transformation result of the normalized signal at time t and frequency f; x(n’) represents the n’th sample of the rotor self-noise signal; w(n’ - t) represents the value of the window function at time index n’ and window start position t; e -j2Πfn / 320 represents the complex exponential function; j represents the imaginary unit; e represents the base of the natural logarithm;
[0025] Use the peak detection algorithm to find the peak in the normalized signal spectrum, and take the peak as the fundamental frequency ffundamental;
[0026] Find the integer multiple frequency positions of the fundamental frequency in the normalized signal spectrum, and extract the harmonic components Eharmonic,i of the fundamental frequency, which means:
[0027] Eharmonic,i = |S(i * ffundamental, t)| i = 2, 3, …, 10 (5);
[0028] Where, Eharmonic,i represents the i’th harmonic component; S represents the spectrum function; ffundamental,t represents the fundamental frequency of the current frame t;
[0029] Arrange the fundamental frequency and harmonic components in sequence to form a feature vector, which is represented as F = [ffundamental,t, Eharmonic,2, Eharmonic,3, …, Eharmonic,10]; Eharmonic,2 represents the second harmonic component; Eharmonic,3 represents the third harmonic component; Eharmonic,10 represents the tenth harmonic component.
[0030] Furthermore, the prediction result of the noise dynamic change label in step S3, the specific process is as follows:
[0031] Take the feature vector F as the input data, process the input data, calculate the change value of the previous frame of the fundamental frequency and harmonic components in the input data, and generate the noise dynamic change label Y = [Δffundamental, ΔEharmonic,2, ΔEharmonic,3, …, ΔEharmonic,10], where ΔEharmonic,10 represents the dynamic change value of the tenth harmonic component, which means:
[0032] Δffundamental = ffundamental,t - ffundamental,t-1 (6);
[0033] ΔEharmonic,i = Eharmonic,i,t - Eharmonic,i,t-1 (7);
[0034] In the formula, Δffundamental represents the difference between the fundamental frequency of the current frame and the fundamental frequency of the previous frame; ffundamental,t-1 represents the fundamental frequency of the previous frame t - 1; ΔEharmonic,i represents the difference between the energy of the i-th harmonic component of the current frame and the energy of the i-th harmonic component of the previous frame; Eharmonic,i,t represents the i-th harmonic component in the previous frame t; Eharmonic,i,t-1 represents the i-th harmonic component in the previous frame t - 1;
[0035] Combine the feature vector F and the noise dynamic change label Y into a data set;
[0036] Construct a deep learning model, input the data set into the deep learning model for processing, and output the prediction result Ypred = [Δffundamental, ΔEharmonic,2, ΔEharmonic,3, …, ΔEharmonic,10] of the noise dynamic change label;
[0037] Use the mean square error as the loss function to calculate the error between the prediction result Ypred of the noise dynamic change label and the noise dynamic change label Y, which means:
[0038]
[0039] Wherein, MSE represents the mean square error; N represents the total number of noise dynamic change labels; Ypred,i represents the prediction result of the i-th noise dynamic change label; Yi represents the i-th noise dynamic change label; n' represents the sample number.
[0040] Furthermore, the enhanced signal in step S4 has the following specific process:
[0041] Extract the periodic harmonic characteristics of the rotor self-noise signal of the UAV to obtain the noisy speech amplitude spectrum |Xm(f)|;
[0042] According to the prediction result Ypred of the noise dynamic change label, use the minimum statistic method to calculate the energy distribution of the rotor self-noise signal of the UAV, perform time smoothing processing on the noisy speech amplitude spectrum |Xm(f)| to obtain the smoothed spectrum |X~m(f)|, and update the noise spectrum estimation by the recursive averaging method, expressed as:
[0043] |Nm(f)| = α·|Nm-1(f)| + (1-α)·|X~m(f)| (9);
[0044] Wherein, |Nm(f)| represents the spectral characteristics of the noise; α represents the over-subtraction factor; |Nm-1(f)| represents the amplitude of the smoothed spectrum of the previous frame;
[0045] Use the speech enhancement algorithm to subtract the spectral characteristics of the noise |Nm(f)| from the noisy speech amplitude spectrum |Xm(f)| to obtain the enhanced amplitude spectrum, expressed as:
[0046]
[0047] Wherein, |Sm(f)| represents the enhanced amplitude spectrum; max represents the maximum value;
[0048] The over-subtraction factor α is used to control the intensity of noise suppression of the noisy speech amplitude spectrum |Xm(f)|. Use the max(…, 0) function to ensure that the noisy speech amplitude spectrum |Xm(f)| is non-negative, and then process the noisy speech amplitude spectrum |Xm(f)| to retain the phase information φm(f) of the noise, expressed as:
[0049] φm(f) = ∠Xm(f) (11);
[0050] Combine the enhanced amplitude spectrum |Sm(f)| and the phase information φm(f) to obtain the enhanced spectrum Gm(f), expressed as:
[0051] Gm(f) = |Sm(f)|·φm(f)e jφm(f) (12);
[0052] where e jφm(f) represents the phase factor in complex form;
[0053] Based on Δffundamental and ΔEharmonic,i, the over-subtraction factor α of the voice enhancement algorithm is adjusted dynamically. When the values of Δffundamental and ΔEharmonic,i increase, the over-subtraction factor α will be increased, and the enhanced parameters after dynamic parameter adjustment are output. The enhanced spectrum Gm(f) is reconstructed in the time domain using the enhanced parameters, and the inverse Fourier transform is performed on the enhanced spectrum Gm(f) to obtain the time-domain signal s m (n), which means:
[0054] sm(n) = IFFT(Gm(f)) (13);
[0055] where IFFT represents the inverse Fourier transform;
[0056] Then, using the overlap-and-add method, for each frame of the time-domain signal s m (n) and the overlapping part of the previous frame of the time-domain signal s m (n) are added together, and finally the enhanced signal s(n) is output.
[0057] Furthermore, a UAV voice enhancement system based on periodic harmonic features and deep learning prediction, for the UAV voice enhancement method based on periodic harmonic features and deep learning prediction described above, includes:
[0058] Self-noise acquisition module, used to collect the rotor self-noise signal of the UAV, that is, the original sampling frequency;
[0059] Periodic harmonic feature extraction module, used to process the original sampling frequency to obtain a new sampling frequency, filter the new sampling frequency to obtain a normalized signal, extract the normalized signal to obtain the normalized signal spectrum, use the peak detection algorithm to find the peaks in the normalized signal spectrum, take the peaks as the fundamental frequency, extract the normalized signal spectrum to obtain harmonic components, and arrange the fundamental frequency and harmonic components in order to form a feature vector;
[0060] Deep learning model prediction noise dynamic change module, used to process the feature vector to generate a noise dynamic change label, combine the feature vector and the noise dynamic change label into a data set, construct a deep learning model, input the data set into the deep learning model for processing, and output the prediction result of the noise dynamic change label;
[0061] A voice enhancement module, which is used to extract the rotor self-noise signal of the drone to obtain the noisy speech amplitude spectrum, calculate the energy distribution of the rotor self-noise signal of the drone using the minimum statistic method according to the prediction result of the noise dynamic change label, smooth the noisy speech amplitude spectrum to obtain the smoothed spectrum, process the smoothed spectrum to obtain the spectral characteristics of the noise, subtract the spectral characteristics of the noise from the noisy speech amplitude spectrum to obtain the enhanced amplitude spectrum, process the noisy speech amplitude spectrum to obtain the phase information, combine the enhanced amplitude spectrum and the phase information to construct the enhanced spectrum, and process the enhanced spectrum to obtain the enhanced signal;
[0062] An output module, which is used to output the enhanced signal.
[0063] Compared with the existing technologies, the present invention has the following beneficial effects:
[0064] (1) By extracting the periodic harmonic characteristics of the rotor self-noise signal of the drone and combining with the deep learning model to predict the noise change trend, the present invention can more accurately separate the noise and the speech, thereby greatly improving the speech signal-to-noise ratio;
[0065] Compared with the traditional static noise reduction methods (such as spectral subtraction or fixed filters), the dynamic parameter adjustment of the present invention can adapt to the complex noise environment such as the change of the rotor speed and the air flow disturbance during the flight of the drone, making the enhanced speech clearer and more intelligible.
[0066] (2) By dynamically adjusting the change of the self-noise of the drone and adjusting the voice enhancement parameters in real time, the present invention reduces the voice distortion;
[0067] Traditional noise reduction methods (such as Wiener filtering or noise threshold with fixed parameters) are prone to cause excessive voice suppression or residual noise when the drone noise changes dynamically;
[0068] The present invention uses a CNN-RNN hybrid model to predict the noise change in real time (such as the fundamental frequency and harmonic energy fluctuation), and dynamically adjusts the over-subtraction factor and filter parameters to avoid voice distortion;
[0069] Experiments show that when the rotor speed changes suddenly (such as accelerating or decelerating), the voice distortion degree (PESQ score) of this method is improved by about 20% compared with the traditional method. Description of the Drawings
[0070] Figure 1 It is the flowchart of the method of the present invention. Detailed Embodiment
[0071] As Figure 1As shown, the present invention provides a technical solution: a UAV voice enhancement method based on periodic harmonic characteristics and deep learning prediction, including the following steps:
[0072] Step S1: Collect the rotor self-noise signal of the UAV, which is the original sampling frequency;
[0073] Step S2: Process the original sampling frequency to obtain a new sampling frequency, filter the new sampling frequency to obtain a normalized signal, extract the normalized signal to obtain the normalized signal spectrum, use the peak detection algorithm to find the peaks in the normalized signal spectrum, take the peaks as the fundamental frequency, extract the normalized signal spectrum to obtain the harmonic components, and arrange the fundamental frequency and harmonic components in order to form a feature vector;
[0074] Step S3: Process the feature vector to generate a noise dynamic change label, combine the feature vector and the noise dynamic change label into a data set, construct a deep learning model, input the data set into the deep learning model for processing, and output the prediction result of the noise dynamic change label;
[0075] Step S4: Extract the rotor self-noise signal of the UAV to obtain the noisy speech amplitude spectrum, calculate the energy distribution of the rotor self-noise signal of the UAV using the minimum statistic method according to the prediction result of the noise dynamic change label, smooth the noisy speech amplitude spectrum to obtain the smoothed spectrum, process the smoothed spectrum to obtain the spectral characteristics of the noise, subtract the spectral characteristics of the noise from the noisy speech amplitude spectrum to obtain the enhanced amplitude spectrum, process the noisy speech amplitude spectrum to obtain the phase information, combine the enhanced amplitude spectrum and the phase information to construct the enhanced spectrum, and process the enhanced spectrum to obtain the enhanced signal;
[0076] Step S5: Output the enhanced signal.
[0077] Among them, in step S1, the process of collecting the rotor self-noise signal of the UAV is as follows:
[0078] A highly sensitive microphone array is mounted on a drone. The highly sensitive microphone array is connected to a 24-bit high-performance analog-to-digital converter (ADC). The self-noise signal of the drone's rotor is collected through the highly sensitive microphone array. The collected self-noise signal of the drone's rotor is an analog signal, which is converted into a digital signal, that is, the original sampling frequency. The main frequency range of the self-noise signal of the drone's rotor is usually from dozens of Hz to several kHz. According to the Nyquist sampling theorem, the sampling frequency of the highly sensitive microphone array should be at least twice the highest frequency of the original sampling frequency, fsample≥2×fmax. The sampling rate of the high-performance analog-to-digital converter is set to 48 kHz to ensure the fidelity of the original sampling frequency; fsample represents the sampling frequency; fmax represents the highest frequency.
[0079] Among them, the eigenvector in step S2, the specific process is as follows:
[0080] To reduce the data volume and improve the calculation efficiency, the original sampling frequency is reduced from 48 kHz to 16 kHz. The new sampling frequency is 16 kHz, and 16 kHz meets the requirements of basic voice calls; before reducing the sampling frequency, a low-pass filter is used to remove the frequency components higher than half of the new sampling frequency to prevent aliasing and remove the frequency components higher than 8 kHz. Take the smallest integer of the original sampling frequency and the new sampling frequency, and set the decimation factor M to 3, which means:
[0081] M = foriginal / fnew = 3(1);
[0082] In the formula, foriginal represents the original sampling frequency (48 kHz); fnew represents the new sampling frequency (16 kHz);
[0083] A Butterworth filter is used to filter the high-frequency noise and low-frequency interference in the new sampling frequency to obtain a filtered signal, retaining the main frequency range of the self-noise signal of the drone's rotor: the main frequency range is 50 Hz - 5 kHz, and the order n of the Butterworth filter is 4, which means:
[0084]
[0085] In the formula, f represents the frequency; H(f) represents the frequency response function of the Butterworth filter; fc represents the cut-off frequency of the Butterworth filter, 5000 Hz; calculate the maximum absolute value |Xmax| of the filtered signal, and divide all the filtered signals by the maximum absolute value |Xmax| of the filtered signal to reach the fixed range of signal amplitude normalization [-1,1], which means:
[0086] Xnormalized = X / |Xmax|(3);
[0087] Where Xnormalized represents the normalized signal, X represents the rotor self-noise signal of the UAV; Xmax represents the maximum absolute value of the signal;
[0088] The normalized signal is divided into short-time frames for subsequent spectrum analysis. When the length of each frame is 20 ms and the sampling rate is 16 kHz, the frame length of 20 ms is 0.02 × 16000 = 320 samples;
[0089] For the extraction of periodic harmonic characteristics of the normalized signal, first perform a short-time Fourier transform on each frame of the normalized signal to obtain the normalized signal spectrum, and convert the time-domain signal of the normalized signal spectrum into a frequency-domain signal for analyzing the frequency-domain signal components. From the preprocessing, each frame of the normalized signal frequency-domain signal has 320 samples, and the window function is a Hamming window, which is expressed as:
[0090]
[0091] Where X(f,t) represents the transformation result of the normalized signal at time t and frequency f; x(n’) represents the n’th sample of the rotor self-noise signal of the UAV; w(n’ - t) represents the value of the window function at time index n’ and window start position t; e -j2Πfn / 320 represents the complex exponential function; j represents the imaginary unit; e represents the base of the natural logarithm;
[0092] Use the peak detection algorithm to find the peaks in the normalized signal spectrum, and take the peaks as the fundamental frequency ffundamental;
[0093] Find the integer multiple frequency positions of the fundamental frequency in the normalized signal spectrum, and extract the harmonic components Eharmonic,i of the fundamental frequency, which is expressed as:
[0094] Eharmonic,i = |S(i * ffundamental,t)| i = 2, 3, …, 10 (5);
[0095] Where Eharmonic,i represents the i’th harmonic component; S represents the spectrum function; ffundamental,t represents the fundamental frequency of the current frame t;
[0096] Arrange the fundamental frequency and harmonic components in order to form a feature vector, which is expressed as F = [ffundamental,t, Eharmonic,2, Eharmonic,3, …, Eharmonic,10]; Eharmonic,2 represents the second harmonic component; Eharmonic,3 represents the third harmonic component; Eharmonic,10 represents the tenth harmonic component.
[0097] Among them, the prediction result of the noise dynamic change label in step S3 is specifically as follows:
[0098] Taking the feature vector F as the input data, processing the input data, calculating the change value of the previous frame of the fundamental frequency and harmonic components in the input data, and generating the noise dynamic change label Y = [Δffundamental, ΔEharmonic,2, ΔEharmonic,3, …, ΔEharmonic,10], where ΔEharmonic,10 represents the dynamic change value of the tenth harmonic component, which means:
[0099] Δffundamental = ffundamental,t - ffundamental,t - 1 (6);
[0100] ΔEharmonic,i = Eharmonic,i,t - Eharmonic,i,t - 1 (7);
[0101] In the formula, Δffundamental represents the difference between the fundamental frequency of the current frame and the fundamental frequency of the previous frame; ffundamental,t - 1 represents the fundamental frequency of the previous frame t - 1; ΔEharmonic,i represents the difference between the energy of the i-th harmonic component of the current frame and the energy of the i-th harmonic component of the previous frame; Eharmonic,i,t represents the i-th harmonic component in the previous frame t; Eharmonic,i,t - 1 represents the i-th harmonic component in the previous frame t - 1;
[0102] Combining the feature vector F and the noise dynamic change label Y into a data set;
[0103] Constructing a deep learning model: Combining the local feature extraction ability of CNN and the time series modeling ability of RNN to construct a deep learning model based on CNN and RNN;
[0104] The deep learning model includes an input layer, a CNN layer, a pooling layer, an RNN layer, a fully connected layer, and an output layer;
[0105] The input layer is used to receive the feature vector F of the time series. The input dimension of the input layer is (T, 10), where T is the time step and 10 is the feature vector dimension (fundamental frequency + 9 harmonic components);
[0106] The CNN layer uses a 1D convolutional layer (Conv1D) to extract the local features of the feature vector F. The kernel size of the CNN layer is set to 3, the stride is 1, the number of filters is 64, the activation function of the CNN layer uses ReLU, and the output dimension of the CNN layer is (T, 64); where 64 represents the channel dimension (i.e., the number of feature maps);
[0107] The pooling layer uses a max pooling layer (MaxPooling1D) to downsample the output of the CNN layer. The pooling window size of the pooling layer is 2, and the output dimension of the pooling layer is (T / 2, 64); T / 2 represents that the time dimension is halved;
[0108] The RNN layer uses LSTM (Long Short-Term Memory Network) to model the temporal dependence relationship of the output of the pooling layer. The number of Long Short-Term Memory network units is 128, and the complete time series output is returned. The output dimension of the Long Short-Term Memory network is (T / 2, 128); LSTM processes the temporal dependence and outputs 128 LSTM units;
[0109] The fully connected layer (Dense) is used to map the output of the Long Short-Term Memory network to the output space to obtain the prediction result of the noise dynamic change label. The number of neurons in the fully connected layer is 64, and the activation function is ReLU;
[0110] The output layer is used to output the prediction result of the noise dynamic change label Ypred = [Δffundamental, ΔEharmonic,2, ΔEharmonic,3, …, ΔEharmonic,10]. The output dimension of the output layer is (T / 2, 10);
[0111] The mean squared error (MSE) is used as the loss function to calculate the error between the prediction result Ypred of the noise dynamic change label and the noise dynamic change label Y. The smaller the MSE value, the better the fitting situation, which is expressed as:
[0112]
[0113] In the formula, MSE represents the mean squared error; N represents the total number of noise dynamic change labels; Y pred,i represents the prediction result of the i-th noise dynamic change label; Y i represents the i-th noise dynamic change label; n’ represents the sample serial number.
[0114] Among them, the enhanced signal in step S4 is specifically as follows:
[0115] Extract the periodic harmonic characteristics of the rotor self-noise signal of the drone to obtain the noisy speech amplitude spectrum ∣Xm(f)∣;
[0116] According to the prediction result Ypred of the noise dynamic change label, use the minimum statistic method to calculate the energy distribution of the rotor self-noise signal of the drone, perform time smoothing processing on the noisy speech amplitude spectrum ∣Xm(f)∣ to obtain the smoothed spectrum ∣X~m(f)∣, and update the noise spectrum estimation by the recursive averaging method, which is expressed as:
[0117] |Nm(f)| = α·|Nm-1(f)| + (1 - α)·|X̃m(f)| (9);
[0118] Where, |Nm(f)| represents the spectral characteristics of the noise; α represents the over-subtraction factor, usually set between 0.9 - 0.99; |Nm-1(f)| represents the amplitude of the smoothed spectrum of the previous frame;
[0119] Using the speech enhancement algorithm to subtract the spectral characteristics of the noise |Nm(f)| from the noisy speech magnitude spectrum |Xm(f)|, the enhanced magnitude spectrum is obtained, expressed as:
[0120]
[0121] Where, |Sm(f)| represents the enhanced magnitude spectrum; max represents the maximum value;
[0122] The over-subtraction factor α is used to control the intensity of noise suppression of the noisy speech magnitude spectrum |Xm(f)|. The max(…, 0) function is used to ensure that the noisy speech magnitude spectrum |Xm(f)| is non-negative. Then, the noisy speech magnitude spectrum |Xm(f)| is processed, and the phase information φm(f) of the noise is retained, expressed as:
[0123] φm(f) = ∠Xm(f) (11);
[0124] Combining the enhanced magnitude spectrum |Sm(f)| and the phase information φm(f) to obtain the enhanced spectrum Gm(f), expressed as:
[0125] Gm(f) = |Sm(f)|·φm(f)e jφm(f) (12);
[0126] Where, e jφm(f) represents the phase factor in complex form;
[0127] Based on Δffundamental and ΔEharmonic,i, the over-subtraction factor α of the speech enhancement algorithm is adjusted dynamically. When the values of Δffundamental and ΔEharmonic,i increase, the over-subtraction factor α will be increased, and the enhanced parameters after dynamic parameter adjustment are output. Using the enhanced parameters to perform time-domain reconstruction on the enhanced spectrum Gm(f), and performing inverse Fourier transform on the enhanced spectrum Gm(f) to obtain the time-domain signal s m (n), expressed as:
[0128] sm(n) = IFFT(Gm(f))(13);
[0129] Where, IFFT represents the inverse Fourier transform;
[0130] Then, using the overlap - add method, for each frame of the time - domain signal s m (n) and the overlapping part of the previous frame of the time - domain signal s m (n) are added together, and finally the enhanced signal s(n) is output.
[0131] Among them, a UAV voice enhancement system based on periodic harmonic features and deep - learning prediction, which is used for the UAV voice enhancement method based on periodic harmonic features and deep - learning prediction, includes:
[0132] A self - noise acquisition module, which is used to collect the rotor self - noise signal of the UAV, that is, the original sampling frequency;
[0133] A periodic harmonic feature extraction module, which is used to process the original sampling frequency to obtain a new sampling frequency, filter the new sampling frequency to obtain a normalized signal, extract the normalized signal to obtain the normalized signal spectrum, use the peak - detection algorithm to find the peaks in the normalized signal spectrum, take the peaks as the fundamental frequency, extract the normalized signal spectrum to obtain harmonic components, and arrange the fundamental frequency and harmonic components in order to form a feature vector;
[0134] A deep - learning model prediction noise dynamic change module, which is used to process the feature vector to generate a noise dynamic change label, combine the feature vector and the noise dynamic change label into a data set, construct a deep - learning model, input the data set into the deep - learning model for processing, and output the prediction result of the noise dynamic change label;
[0135] A voice enhancement module, which is used to extract the rotor self - noise signal of the UAV to obtain the noisy speech amplitude spectrum, calculate the energy distribution of the rotor self - noise signal of the UAV using the minimum statistic method according to the prediction result of the noise dynamic change label, smooth the noisy speech amplitude spectrum to obtain a smoothed spectrum, process the smoothed spectrum to obtain the spectral characteristics of the noise, obtain the enhanced amplitude spectrum by subtracting the spectral characteristics of the noise from the noisy speech amplitude spectrum, process the noisy speech amplitude spectrum to obtain phase information, combine the enhanced amplitude spectrum and the phase information to construct an enhanced spectrum, and process the enhanced spectrum to obtain an enhanced signal;
[0136] An output module, which is used to output the enhanced signal.
[0137] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for drone speech enhancement based on periodic harmonic features and deep learning prediction, characterized in that: The steps include: Step S1: collecting the rotor self-noise signal of the UAV, that is, the original sampling frequency; Step S2: Process the original sampling frequency to obtain a new sampling frequency, filter the new sampling frequency to obtain a normalized signal, extract the normalized signal, obtain the normalized signal spectrum, use a peak detection algorithm to find the peak in the normalized signal spectrum, use the peak as the fundamental frequency, extract the normalized signal spectrum, obtain the harmonic component, and arrange the fundamental frequency and the harmonic component in order to form a feature vector; Step S3: Process the feature vector to generate a noise dynamic change label, combine the feature vector and the noise dynamic change label into a data set, build a deep learning model, input the data set into the deep learning model for processing, and output the prediction result of the noise dynamic change label; Step S4: extracting the rotor self-noise signal of the UAV to obtain the noisy speech amplitude spectrum, using the minimum statistics method to calculate the energy distribution of the rotor self-noise signal of the UAV according to the prediction result of the noise dynamic change label, smoothing the noisy speech amplitude spectrum to obtain a smoothed spectrum, processing the smoothed spectrum to obtain the spectrum characteristics of the noise, subtracting the spectrum characteristics of the noise from the noisy speech amplitude spectrum to obtain an enhanced amplitude spectrum, processing the noisy speech amplitude spectrum to obtain phase information, combining the enhanced amplitude spectrum with the phase information to construct an enhanced spectrum, processing the enhanced spectrum to obtain an enhanced signal; Step S5: output the enhanced signal.
2. The method for drone speech enhancement based on periodic harmonic features and deep learning prediction according to claim 1 is characterized in that: In step S1, the rotor self-noise signal of the UAV is collected, and the specific process is as follows: A high-sensitivity microphone array is mounted on a drone. The high-sensitivity microphone array is connected to a high-performance digital-to-analog converter. The rotor self-noise signal of the drone is collected through the high-sensitivity microphone array. The rotor self-noise signal of the drone is collected as an analog signal, and the analog signal is converted into a digital signal, which is the original sampling frequency.
3. The method for drone speech enhancement based on periodic harmonic features and deep learning prediction according to claim 1, characterized in that: The specific process of the feature vector in step S2 is: The original sampling frequency is downsampled to obtain a new sampling frequency. Before downsampling, a low-pass filter is used to remove frequency components higher than half of the new sampling frequency. The minimum integer between the original sampling frequency and the new sampling frequency is taken, and the downsampling factor M is set to 3, which means: M = foriginal / fnew = 3(1); In the formula, foriginal represents the original sampling frequency; fnew represents the new sampling frequency; The Butterworth filter is used to filter the high-frequency noise and low-frequency interference in the new sampling frequency to obtain a filtered signal. The Butterworth filter order n = 4, which means: Where, f represents frequency; H(f) represents the frequency response function of the Butterworth filter; fc represents the cutoff frequency of the Butterworth filter; Calculate the maximum absolute value of the filtered signal |Xmax|, and divide all filtered signals by the maximum absolute value of the filtered signal |Xmax| to reach the signal amplitude standardized fixed range [-1,1], which means: Xnormalized = X / |Xmax| (3); In the formula, Xnormalized represents the normalized signal, X represents the rotor self-noise signal of the UAV; Xmax represents the maximum absolute value of the signal; Perform short-time Fourier transform on each frame of the normalized signal to obtain the normalized signal spectrum, which is expressed as: Where X(f, t) represents the transformation result of the normalized signal at time t and frequency f; x(n') represents the n'th sample of the rotor self-noise signal; w(n'-t) represents the value of the window function at time index n' and window starting position t; e -j2Πfn / 320 represents the complex exponential function; j represents the imaginary unit; e represents the base of the natural logarithm; Use the peak detection algorithm to find the peak in the normalized signal spectrum and take the peak as the fundamental frequency ffundamental; Find the integer multiple frequency position of the fundamental frequency in the normalized signal spectrum and extract the harmonic component Eharmonic,i of the fundamental frequency, which is expressed as: Eharmonic,i=|S(i*fundamental,t)|i=2,3,…,10(5); Where, Eharmonic,i represents the i-th harmonic component; S represents the spectrum function; ffundamental,t represents the fundamental frequency of the current frame t; Arrange the fundamental frequency and harmonic components in order to form a eigenvector, which is expressed as F = [ffundamental,t,Eharmonic,2,Eharmonic,3,…,Eharmonic,10]; Eharmonic,2 represents the second harmonic component; Eharmonic,3 represents the third harmonic component; Eharmonic,10 represents the tenth harmonic component.
4. The method for drone speech enhancement based on periodic harmonic features and deep learning prediction according to claim 3 is characterized in that: The prediction result of the noise dynamic change label in step S3 is as follows: The feature vector F is used as input data, and the input data is processed to calculate the change value of the fundamental frequency and harmonic components in the input data in the previous frame, and the noise dynamic change label Y = [Δffundamental, ΔEharmonic, 2, ΔEharmonic, 3, ..., ΔEharmonic, 10] is generated. ΔEharmonic, 10 represents the dynamic change value of the tenth harmonic component, which means: Δffundamental=ffundamental,t-ffundamental,t-1(6); ΔEharmonic,i=Eharmonic,i,t-Eharmonic,i,t-1(7); Wherein, Δffundamental represents the difference between the fundamental frequency of the current frame and the fundamental frequency of the previous frame; ffundamental,t-1 represents the fundamental frequency of the previous frame t-1; ΔEharmonic,i represents the energy of the i-th harmonic component of the current frame minus the energy of the i-th harmonic component of the previous frame; Eharmonic,i,t represents the i-th harmonic component in the previous frame t; Eharmonic,i,t-1 represents the i-th harmonic component in the previous frame t-1; Combine the feature vector F and the noise dynamic change label Y into a data set; Build a deep learning model, input the data set into the deep learning model for processing, and output the prediction result of the noise dynamic change label Ypred = [Δffundamental, ΔEharmonic, 2, ΔEharmonic, 3, …, ΔEharmonic, 10]; Using the mean square error as the loss function, the error between the predicted result Ypred of the noise dynamic change label and the noise dynamic change label Y is calculated, which is expressed as: Where MSE represents mean square error; N represents the total number of noise dynamic change labels; Ypred,i represents the prediction result of the i-th noise dynamic change label; Yi represents the i-th noise dynamic change label; n' represents the sample number.
5. The method for drone speech enhancement based on periodic harmonic features and deep learning prediction according to claim 4 is characterized in that: The signal enhanced in step S4 is specifically processed as follows: The periodic harmonic characteristics of the UAV rotor self-noise signal are extracted to obtain the amplitude spectrum of the noisy speech |Xm(f)|; According to the prediction result Ypred of the noise dynamic change label, the minimum statistics method is used to calculate the energy distribution of the drone's rotor self-noise signal, and the noisy speech amplitude spectrum |Xm(f)| is time-smoothed to obtain the smoothed spectrum |X~m(f)|. The noise spectrum estimation is updated by the recursive averaging method, which is expressed as: ∣Nm(f)∣=α·∣Nm-1(f)∣+(1-α)·∣X~m(f)∣ (9); Where |Nm(f)| represents the spectral characteristics of the noise; α represents the over-subtraction factor; |Nm-1(f)| represents the amplitude of the smoothed spectrum of the previous frame; The speech enhancement algorithm is used to subtract the spectral characteristics of the noise |Nm(f)| from the amplitude spectrum of the noisy speech |Xm(f)| to obtain the enhanced amplitude spectrum, which is expressed as: In the formula, |Sm(f)| represents the enhanced amplitude spectrum; max represents the maximum value; The over-reduction factor α is used to control the intensity of noise suppression of the noisy speech amplitude spectrum |Xm(f). The max(…,0) function is used to determine that the noisy speech amplitude spectrum |Xm(f) is a non-negative value. Then the noisy speech amplitude spectrum |Xm(f) is processed to retain the phase information φm(f) of the noise, which is expressed as: φm(f)=∠Xm(f)(11); The enhanced amplitude spectrum |Sm(f)| is combined with the phase information φm(f) to obtain the enhanced spectrum Gm(f), which is expressed as: Gm(f)=∣Sm(f)∣·φm(f)e jφm(f) (12); In the formula, e jφm(f) represents the phase factor in complex form; Based on Δffundamental and ΔEharmonic,i, the over-subtraction factor α of the speech enhancement algorithm is adjusted through dynamic parameters. When the values of Δffundamental and ΔEharmonic,i increase, the over-subtraction factor α is increased, and the enhanced parameters after dynamic parameter adjustment are output. The enhanced spectrum Gm(f) is reconstructed in the time domain using the enhanced parameters, and the enhanced spectrum Gm(f) is inversely Fourier transformed to obtain the time domain signal s m (n), which means: sm(n)=IFFT(Gm(f))(13); Where, IFFT stands for inverse Fourier transform; Then use the overlap-add method to calculate each frame of the time domain signal s m (n) and the previous frame time domain signal s m The overlapping parts of the two signals are added together, and the enhanced signal s(n) is finally output.
6. A drone voice enhancement system based on periodic harmonic features and deep learning prediction, used in a drone voice enhancement method based on periodic harmonic features and deep learning prediction as claimed in any one of claims 1 to 5, characterized in that: include: The self-noise acquisition module is used to collect the self-noise signal of the drone's rotor, that is, the original sampling frequency; The periodic harmonic feature extraction module is used to process the original sampling frequency to obtain a new sampling frequency, filter the new sampling frequency to obtain a normalized signal, extract the normalized signal, obtain the normalized signal spectrum, use the peak detection algorithm to find the peak in the normalized signal spectrum, use the peak as the fundamental frequency, extract the normalized signal spectrum, obtain the harmonic component, and arrange the fundamental frequency and the harmonic component in order to form a feature vector; The deep learning model predicts the noise dynamic change module, which is used to process the feature vector, generate the noise dynamic change label, combine the feature vector and the noise dynamic change label into a data set, build a deep learning model, input the data set into the deep learning model for processing, and output the prediction result of the noise dynamic change label; The speech enhancement module is used to extract the rotor self-noise signal of the drone to obtain the noisy speech amplitude spectrum, calculate the energy distribution of the rotor self-noise signal of the drone using the minimum statistics method according to the prediction result of the noise dynamic change label, smooth the noisy speech amplitude spectrum to obtain the smoothed spectrum, process the smoothed spectrum to obtain the spectrum characteristics of the noise, subtract the spectrum characteristics of the noise from the noisy speech amplitude spectrum to obtain the enhanced amplitude spectrum, process the noisy speech amplitude spectrum to obtain the phase information, combine the enhanced amplitude spectrum with the phase information to construct the enhanced spectrum, and process the enhanced spectrum to obtain the enhanced signal; The output module is used to output the enhanced signal.
Citation Information
Patent Citations
Single-channel speech enhancement method based on amplitude estimation and phase reconstruction
CN114005457A
Unmanned aerial vehicle self-noise filtering method and system based on multi-channel time-frequency spatial filtering
CN119517066A
Non-stationary / Mixed Noise Estimation Method based on Minium Statistics and Codebook Driven Short-Term Predictor Parameter Estimation
KR1020110076315A
Method and apparatus for improved estimation of non-stationary noise for speech enhancement
US20070055508A1
Cited By
Unmanned aerial vehicle classification method and system based on noise perception model
CN120375863A
Self-adaptive noise reduction filtering method, system, medium and equipment
CN120998172A
Unmanned aerial vehicle self-noise dynamic cancellation and target sound enhancement method based on diffusion model
CN121617380A
Interference source analysis method and system for partial discharge measurement of large oil-filled equipment
CN122090877A