Lightweight speech enhancement method applied to low-altitude aircraft and model thereof
Through the lightweight voice enhancement method, the problems of high computational complexity and high delay in low-altitude vehicles are solved, and efficient speech clarity and real-time performance in strong noise environments are achieved, which is suitable for voice communication and command recognition of low-altitude vehicles.
Patent Information
- Application Number
- CN202510451241.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-08
AI Technical Summary
The existing voice enhancement technology has problems such as high computational complexity, large delay and insufficient noise robustness in low-altitude aircraft, which cannot meet the high noise, low computing power and strict real-time requirements of low-altitude aircraft.
Lightweight speech enhancement methods are adopted, including preprocessing, coordinate attention module, time convolution module, small-scale separable convolution module and mask calculation module. Voice signals are collected through microphones, time-frequency feature extraction, adaptive weighting, transient and steady-state feature extraction and noise suppression, and finally restored to a clear time-domain voice signal.
It effectively improves the speech clarity and nature of low-altitude aircraft in a strong noise environment, meets the requirements of low-altitude aircraft for real-time and lightweight, and is suitable for voice communication and command recognition scenarios.
Smart Images

Figure CN120279929A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of speech enhancement, and particularly relates to a lightweight speech enhancement method and its model applied to low-altitude aircraft. Background Art
[0002] Existing speech enhancement technologies mainly rely on deep neural networks and multi-scale convolutional noise reduction methods combined with neural network structures. Although these methods can achieve high speech quality in the PC or server environment, they have limitations in the application scenarios of low-altitude aircraft. Specifically, the parameter quantity and computational complexity of existing models are too large, far exceeding the real-time processing capabilities of airborne embedded processors, resulting in difficult deployment. At the same time, in existing solutions, such as the real-time speech enhancement algorithm based on WFST, the processing delay on the embedded platform usually exceeds 80 ms, which cannot meet the strict requirements of aviation communication for low latency (<20 ms). In addition, although traditional speech enhancement methods are effective for steady-state noise, they are insufficient in adapting to the non-stationary noise unique to low-altitude aircraft, such as rotor harmonics and airflow pulse noise, and their performance significantly degrades in an environment with a large SNR fluctuation.
[0003] Therefore, it is difficult for existing technologies to achieve efficient and reliable speech enhancement under the conditions of high noise, low computing power, and strict real-time requirements of low-altitude aircraft, and there is an urgent need for an optimized solution with lightweight, low latency, and strong noise robustness. Summary of the Invention
[0004] The purpose of the present invention is to provide a lightweight speech enhancement method and its model applied to low-altitude aircraft, aiming to solve problems such as high computational complexity, high latency, and insufficient noise robustness of existing speech enhancement technologies in the airborne embedded environment.
[0005] The technical solution adopted by the present invention is as follows: A lightweight speech enhancement method applied to low-altitude aircraft, comprising the following steps:
[0006] S1, Collect the noisy speech signal in the environment of the low-altitude aircraft through a microphone. The noisy speech signal is from the speech generated when the distance between the low-altitude aircraft and the ground speech source is within 100 meters during flight, or the speech generated during in-cockpit voice communication during the flight of the low-altitude aircraft;
[0007] S2, Input the collected noisy speech signal into the lightweight speech enhancement model, and use the preprocessing module to preprocess the noisy speech signal to obtain time-frequency features;
[0008] S3, Input the time-frequency features into the coordinate attention module, extract the spatial distribution features, and perform adaptive weighting to obtain the weighted speech signal;
[0009] S4. Input the weighted speech signal into the temporal convolutional module to extract transient features and steady-state features, and output an optimized feature signal;
[0010] S5. Process the optimized temporal feature signal through a small-scale separable convolutional module to obtain enhanced speech features;
[0011] S6. Input the enhanced speech features into the mask calculation module to perform mask calculation on the enhanced speech features, generate an enhanced speech spectrum, and restore it to a time-domain speech signal through inverse short-time Fourier transform.
[0012] Further, in S2, use the preprocessing module to preprocess the noisy speech signal. The preprocessing process includes:
[0013] S201. Perform pre-emphasis processing on the noisy speech signal;
[0014] S202. Divide the pre-emphasized audio signal into several frames, set a fixed length and frame shift for each frame, and obtain the framed data;
[0015] S203. Add a window function to the framed data. The window function is a Hanning window, and the window length is set to 10 ms;
[0016] S204. Perform short-time Fourier transform on each frame of the signal to convert the signal from the time domain to the frequency domain, and extract amplitude spectrum and phase spectrum features;
[0017] S205. Stack the short-time Fourier transform results of each frame along the time dimension to generate a spectrogram containing time-frequency features.
[0018] Further, in S3, the coordinate attention module specifically includes:
[0019] S301. Perform global average pooling on the input time-frequency features along the time axis to obtain a first aggregated feature that only varies with channels and frequencies The calculation formula is:
[0020]
[0021] where S, represents the data dimension jointly composed of C, T, and F. C is the number of channels, T is the number of time frames, F is the number of frequency points, S(c,t,f) represents the feature value of the c-th channel at the t-th frame and the f-th frequency point, represents the data dimension when T = 1;
[0022] S302. Perform global average pooling on the time-frequency features along the frequency axis to obtain a second aggregated feature that only varies with channels and time frames Its calculation formula is:
[0023]
[0024] S303. After splicing the first aggregation feature and the second aggregation feature in the spatial dimension, a fused feature is obtained, and a one-dimensional convolution operation is applied to the fused feature, followed by batch normalization and ReLU activation in sequence, and an intermediate feature is generated through the channel compression ratio r
[0025] S304. The intermediate feature is split into a time-direction component and a frequency-direction component in the spatial dimension. One-dimensional convolution and the Sigmoid activation function are respectively applied to the time-direction component and the frequency-direction component to generate a time attention weight and a frequency attention weight
[0026] S305. The time attention weight W t and the frequency attention weight W f are applied to the input time-frequency feature, and weighted processing is respectively performed on the time and frequency dimensions to generate a weighted speech signal where represents element-wise multiplication.
[0027] Furthermore, in S4, transient features and steady-state features are extracted through a time convolution module. The specific process includes:
[0028] S401. Receive the input weighted speech signal S weighted and its corresponding feature matrix where C is the number of channels and N is the feature dimension of the sequence;
[0029] S402. In the multi-scale dilated convolution layer, a set of convolution kernels [S1, S2, S3, S4] with increasing dilation rates are used to perform causal dilated convolution on the input feature matrix X respectively, and multi-scale output feature matrices The calculation formula of the input feature matrix is:
[0030]
[0031] where, represents a one-dimensional convolution operation with a kernel size of S i and S i represents the i-th convolution kernel;
[0032] S403. Align the output multi-scale feature matrices X1, X2, X3, X4 in the time dimension and retain the same time step Talign = T - S max + 1, where S max = max(S1, S2, S3, S4), and the specific operation for alignment in the time dimension is X i [:,:, -T align :] = X i [:,:,(T - S max + 1):], so that all feature matrices have the same time step. Among them, S max is the largest convolution kernel among S1, S2, S3, S4, where T align represents the number of time frames after alignment, and X i [:,:, -T align :] means that in the time dimension, only the last (T - S max + 1) time steps are retained;
[0033] S404. Concatenate the aligned feature matrices X1, X2, X3, X4 in the channel dimension to obtain the fused feature matrix
[0034] S405. Apply two fully connected layers and the ReLU non - linear activation function to the fused feature matrix X in sequence to obtain the intermediate feature matrix X L ; mid ;
[0035] S406. The intermediate feature matrix X mid generates two paths of feature outputs through two groups of parallel dilated multi - scale convolutional layers respectively. One path passes through the tanh activation function as the filter feature matrix X tanh , and the other path passes through the Sigmoid activation function as the gating feature matrix X sigmoid . Subsequently, the two paths of feature matrices combine with the gating mechanism to calculate the final time - convolutional output matrix Y = tanh(X tanh ) ⊙ σ(X sigmoid ); where ⊙ represents the element - by - element multiplication operation, X tanh is the filter feature matrix after passing through the tanh activation function, X sigmoid is the gating feature matrix after passing through the Sigmoid activation function, tanh() represents the tanh activation function, and σ() represents the Sigmoid activation function;
[0036] S407. The time - convolutional module outputs the feature matrix Y, which contains the extracted transient features and steady - state features, as the optimized feature signal for subsequent processing.
[0037] Furthermore, in S5, the small - scale separable convolution specifically includes;
[0038] S501, receive the feature matrix corresponding to the optimized feature signal output by the time convolution module where C
[0039] is the number of channels, T is the number of time frames, and F is the number of frequency points;
[0040] S502, perform a convolution operation on each channel in the feature matrix Y separately to extract the time and frequency features within the channel. The output Y d (c, t, f) of the c-th channel is:
[0041]
[0042] where Kd is the depth convolution kernel with a size of k d,t ×k d,f , k d,t represents the size of the convolution kernel in the time dimension, and k d,f represents the size in the frequency dimension. i and j are the indices of the convolution kernel in the time and frequency directions respectively;
[0043] S503, use a 1×1 convolution kernel K p on the depth convolution output Yd of S502 to perform a linear combination of the outputs of each channel to achieve cross-channel information fusion and channel compression;
[0044] Specifically, for each time frame and frequency point, the output Y p (t, f) is calculated as follows:
[0045]
[0046] where K p (c) is the convolution weight corresponding to the channel; this step compresses the number of channels from to a lower value. According to the set channel compression ratio, the final output number of channels is C / r;
[0047] S504, sequentially apply batch normalization and activation using the ReLU activation function to the output Y p obtained in step S503 to generate the final output feature matrix, which is used as the enhanced speech feature after small-scale separable convolution processing.
[0048] Furthermore, in S5, the mask calculation module specifically includes:
[0049] S601, generate a speech mask value based on the enhanced feature map;
[0050] S602, adjust the time-frequency features after noise suppression using the speech mask;
[0051] S603, Restore the adjusted time-frequency features to a clear time-domain speech signal through inverse short-time Fourier transform.
[0052] Further, a lightweight speech enhancement model applied to low-altitude aircraft, referring to the lightweight speech enhancement method applied to low-altitude aircraft, specifically:
[0053] It includes a preprocessing module, a coordinate attention module, a temporal convolution module, a separable convolution module, and a mask calculation module;
[0054] The input of the lightweight speech enhancement model is the noisy speech signal collected by the microphone;
[0055] The preprocessing module is connected to the coordinate attention module, the coordinate attention module is connected to the temporal convolution module, the temporal convolution module is connected to the small-scale separable convolution module, and the small-scale separable convolution module is connected to the mask calculation module;
[0056] The output time-domain speech signal obtained through the mask module is used as the output of the lightweight speech enhancement model, and the time-domain speech signal is played through an audio playback device.
[0057] The technical effects achieved by the present invention are as follows: The present invention can effectively improve the speech clarity and naturalness in the strong noise environment of low-altitude aircraft, while meeting the requirements of low-altitude aircraft for real-time performance and lightweight, and is applicable to the speech communication and command recognition scenarios of low-altitude aircraft. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a schematic structural diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0059] The present invention and its embodiments are described below. This description is not restrictive, and the actual embodiments are not limited thereto. Generally speaking, if those of ordinary skill in the art are inspired by it and design a structural method and embodiments similar to this technical solution without creative efforts without departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.
[0060] The technical solution adopted by the present invention is: A lightweight speech enhancement method applied to low-altitude aircraft, including the following steps:
[0061] S1, Collect the noisy speech signal in the environment of low-altitude aircraft through the microphone. The noisy speech signal comes from the speech generated when the distance between the low-altitude aircraft in flight and the ground speech sound source is within 100 meters, or the speech generated during in-cockpit speech communication during the flight of the low-altitude aircraft;
[0062] S2. Input the collected noisy speech signal into the lightweight speech enhancement model, and use the preprocessing module to preprocess the noisy speech signal to obtain time-frequency features;
[0063] S3. Input the time-frequency features into the coordinate attention module, extract the spatial distribution features, and perform adaptive weighting to obtain the weighted speech signal;
[0064] S4. Input the weighted speech signal into the temporal convolutional module, extract the transient features and steady-state features, and output the optimized feature signal;
[0065] S5. Process the optimized time feature signal through the small-scale separable convolutional module to obtain the enhanced speech features;
[0066] S6. Input the enhanced speech features into the mask calculation module, perform mask calculation on the enhanced speech features, generate the enhanced speech spectrum, and restore it to the time-domain speech signal through the inverse short-time Fourier transform.
[0067] Further, in S2, when using the preprocessing module to preprocess the noisy speech signal, the preprocessing process includes:
[0068] S201. Perform pre-emphasis processing on the noisy speech signal;
[0069] S202. Divide the pre-emphasized audio signal into several frames, set a fixed length and frame shift for each frame to obtain the framed data;
[0070] S203. Add a window function to the framed data. The window function is a Hanning window, and the window length is set to 10 ms;
[0071] S204. Perform short-time Fourier transform on each frame of the signal, convert the signal from the time domain to the frequency domain, and extract the amplitude spectrum and phase spectrum features;
[0072] S205. Stack the short-time Fourier transform results of each frame along the time dimension to generate a spectrogram containing time-frequency features.
[0073] Further, in S3, the coordinate attention module specifically includes:
[0074] S301. Perform global average pooling on the input time-frequency features along the time axis to obtain the first aggregation feature that only varies with the channel and frequency The calculation formula is:
[0075]
[0076] where S, Denote the data dimension jointly composed of C, T, and F, where C is the number of channels, T is the number of time frames, and F is the number of frequency points. S(c, t, f) represents the feature value at the c-th channel, the t-th frame, and the f-th frequency point. Denote the data dimension when T = 1;
[0077] S302, perform global average pooling on the time-frequency feature along the frequency axis to obtain a second aggregated feature that only varies with channels and time frames Its calculation formula is:
[0078]
[0079] S303, concatenate the first aggregated feature and the second aggregated feature in the spatial dimension to obtain a fused feature and apply a one-dimensional convolution operation to the fused feature, followed by batch normalization and ReLU activation in sequence, to generate an intermediate feature through the channel compression ratio r
[0080] S304, split the intermediate feature in the spatial dimension into a time-direction component and a frequency-direction component Apply one-dimensional convolution and the Sigmoid activation function to the time-direction component and the frequency-direction component respectively to generate a time attention weight and a frequency attention weight
[0081] S305, apply the time attention weight W t and the frequency attention weight W f to the input time-frequency feature, and perform weighting on the time and frequency dimensions respectively to generate a weighted speech signal where represents element-wise multiplication.
[0082] Furthermore, in S4, transient features and steady-state features are extracted through a time convolution module. The specific process includes:
[0083] S401, receive the input weighted speech signal S weighted corresponding feature matrix where C is the number of channels and N is the feature dimension of the sequence;
[0084] S402, in the multi-scale dilated convolutional layer, use a set of convolutional kernels [S1, S2, S3, S4] with increasing dilation rates to perform causal dilated convolution on the input feature matrix X respectively, and output multi-scale feature matrices The calculation formula of the input feature matrix is:
[0085] X i = Conv 1×Si (X), i ∈ {1, 2, 3, 4}
[0086] Among them, represents a one-dimensional convolution operation with a kernel size of S i and S i represents the i-th convolution kernel;
[0087] S403. Align the output multi-scale feature matrices X1, X2, X3, and X4 in the time dimension, and retain the same time step T align = T - S max + 1, where S max = max(S1, S2, S3, S4). The specific operation of alignment in the time dimension is X i [:,:,-T align :] = X i [:,:,(T - S max + 1):], so that all feature matrices have the same time step. Among them, S max is the largest convolution kernel among S1, S2, S3, and S4. Among them, T align represents the number of time frames after alignment, and X i [:,:,-T align :] means that in the time dimension, only the last (T - S max + 1) time steps are retained;
[0088] S404. Concatenate the aligned feature matrices X1, X2, X3, and X4 in the channel dimension to obtain a fused feature matrix
[0089] S405. Apply two fully connected layers and the ReLU non-linear activation function to the fused feature matrix X L in sequence to obtain an intermediate feature matrix X mid ;
[0090] S406. The intermediate feature matrix X mid generates two-way feature outputs through two groups of parallel dilated multi-scale convolution layers respectively. One path passes through the tanh activation function as the filter feature matrix X tanh , and the other path passes through the Sigmoid activation function as the gating feature matrix X sigmoid . Subsequently, the two-way feature matrices combine the gating mechanism to calculate the final time convolution output matrix Y = tanh(X tanh ) ⊙ σ(X sigmoid ); where ⊙ represents the element-wise multiplication operation, and X tanh is the filter feature matrix after passing through the tanh activation function, Xsigmoid is the gated feature matrix after passing through the Sigmoid activation function, tanh() represents the tanh activation function, and σ() represents the Sigmoid activation function;
[0091] S407, the output feature matrix Y of the temporal convolutional module contains the extracted transient and steady-state features, which are used as the optimized feature signal for subsequent processing.
[0092] Furthermore, in S5, the depthwise separable convolution specifically includes:
[0093] S501, receiving the feature matrix corresponding to the optimized feature signal output by the temporal convolutional module where C is the number of channels, T is the number of time frames, and F is the number of frequency points;
[0094] S502, performing a convolution operation on each channel in the feature matrix Y separately to extract the temporal and frequency features within the channel. The output Y d (c, t, f) is:
[0095]
[0096] where Kd is the depthwise convolution kernel with size k d,t ×k d,f , k d,t represents the size of the convolution kernel in the temporal dimension, and k d,f represents the size in the frequency dimension. i and j are the indices of the convolution kernel in the temporal and frequency directions respectively;
[0097] S503, using a 1×1 convolution kernel K p to perform a linear combination of the outputs of each channel in the depthwise convolution output Yd of S502, achieving cross-channel information fusion and channel compression;
[0098] Specifically, for each time frame and frequency point, the output Y p (t, f) is calculated as follows:
[0099]
[0100] where K p (c) is the convolution weight corresponding to the channel; this step compresses the number of channels to a lower value. According to the set channel compression ratio, the final number of output channels is C / r;
[0101] S504, applying batch normalization and activating with the ReLU activation function to the output Y p obtained in step S503 in sequence, generating the final output feature matrix, which serves as the enhanced speech feature after being processed by the depthwise separable convolution.
[0102] Further, in S5, the mask calculation module specifically includes:
[0103] S601, generating a voice mask value according to the enhanced feature map;
[0104] S602, using the voice mask to adjust the time-frequency features after noise suppression;
[0105] S603, restoring the adjusted time-frequency features to a clear time-domain voice signal through inverse short-time Fourier transform.
[0106] Further, a lightweight voice enhancement model applied to low-altitude aircraft, citing the lightweight voice enhancement method applied to low-altitude aircraft, specifically:
[0107] It includes a preprocessing module, a coordinate attention module, a temporal convolution module, a separable convolution module, and a mask calculation module;
[0108] The input of the lightweight voice enhancement model is the noisy voice signal collected by the microphone;
[0109] The preprocessing module is connected to the coordinate attention module, the coordinate attention module is connected to the temporal convolution module, the temporal convolution module is connected to the small-scale separable convolution module, and the small-scale separable convolution module is connected to the mask calculation module;
[0110] The output time-domain voice signal obtained through the mask module is used as the output of the lightweight voice enhancement model, and the time-domain voice signal is played through an audio playback device.
Claims
1. A lightweight speech enhancement method applied to low-altitude aircraft, characterized in that, It includes the following steps: S1. Collect the noisy speech signal in the environment of a low-altitude aircraft. The noisy speech signal comes from the speech generated when the distance between the low-altitude aircraft and the ground speech source is within 100 meters during flight, or the speech generated during in-cockpit voice communication during the flight of the low-altitude aircraft; S2. Input the collected noisy speech signal into the lightweight speech enhancement model, and use the preprocessing module to preprocess the noisy speech signal to obtain time-frequency features; S3. Input the time-frequency features into the coordinate attention module, extract the spatial distribution features, and perform adaptive weighting to obtain the weighted speech signal; S4. Input the weighted speech signal into the temporal convolutional module, extract the transient features and steady-state features, and output the optimized feature signal; S5. Process the optimized time feature signal through the small-scale separable convolutional module to obtain the enhanced speech features; S6. Input the enhanced speech features into the mask calculation module, perform mask calculation on the enhanced speech features to generate the enhanced speech spectrum, and restore it to the time-domain speech signal through the inverse short-time Fourier transform.
2. The lightweight voice enhancement method for low-altitude aircraft according to claim 1, wherein, In S2, when using the preprocessing module to preprocess the noisy speech signal, the preprocessing process includes: S201. Perform pre-emphasis processing on the noisy speech signal; S202. Divide the pre-emphasized audio signal into several frames, and set a fixed length and frame shift for each frame to obtain the framed data; S203. Add a window function to the framed data. The window function is a Hanning window, and the window length is set to 10 ms; S204. Perform short-time Fourier transform on each frame of the signal to convert the signal from the time domain to the frequency domain, and extract the amplitude spectrum and phase spectrum features; S205. Stack the short-time Fourier transform results of each frame in the time dimension to generate a spectrogram containing time-frequency features.
3. The lightweight voice enhancement method for a low-altitude aircraft according to claim 1, characterized in that, In S3, the coordinate attention module specifically includes: S301, perform global average pooling on the input time-frequency features along the time axis to obtain the first aggregated feature that only varies with channels and frequencies The calculation formula thereof is as follows: Among them, represents the data dimension jointly composed of C, T, and F. C is the number of channels, T is the number of time frames, F is the number of frequency points, and S(c, t, f) represents the feature S at the c-th channel, the t-th frame, and the f-th frequency point. t The av value g, ∈ represents the data dimension when T = 1; S302, perform global average pooling on the time-frequency feature along the frequency axis to obtain a second aggregated feature that only varies with channels and time frames The calculation formula thereof is: S303, concatenate the first aggregation feature and the second aggregation feature in the spatial dimension to obtain a fused feature and apply a one-dimensional convolution operation to the fused feature, followed by batch normalization and ReLU activation in sequence, and generate an intermediate feature through a channel compression ratio r S304. Split the intermediate feature into a time-direction component and a frequency-direction component in the spatial dimension. and a frequency-direction component Apply a one-dimensional convolution and a Sigmoid activation function to the time-direction component and the frequency-direction component respectively to generate a time attention weight and a frequency attention weight S305, apply the time attention weight W t and the frequency attention weight W f to the input time-frequency feature, perform weighting on the time and frequency dimensions respectively, and generate a weighted speech signal where represents element-wise multiplication.
4. The lightweight voice enhancement method for low-altitude aircraft according to claim 1, characterized in that In S4, when extracting the transient features and steady-state features through the temporal convolutional module, the specific process includes: S401, receive the input weighted speech signal S weighted The corresponding feature matrix where C is the number of channels and N is the feature dimension of the sequence; S402. In the multi-scale dilated convolutional layer, a set of convolutional kernels [S1, S2, S3, S4] with increasing dilation rates are used to perform causal dilated convolution on the input feature matrix X respectively, and the multi-scale output feature matrix The calculation formula of the input feature matrix is as follows: Among them, represents a one-dimensional convolution operation with a kernel size of S i where S i represents the i-th convolution kernel; At S403, align the output multi-scale feature matrices X1, X2, X3, X4 in the time dimension and retain the same time step T align = T - S max + 1, where S max = max(S1, S2, S3, S4). The specific operation for alignment in the time dimension is X i [:,:,-T align :] = X i [:,:,(T - S max + 1):], so that all feature matrices have the same time step. Among them, S max is the largest convolutional kernel among S1, S2, S3, S4, where T align represents the number of aligned time frames, and X i [:,:,-T align :] means that in the time dimension, only the last (T - S max + 1) time steps are retained; S404. Concatenate the aligned feature matrices X1, X2, X3, and X4 in the channel dimension to obtain the fused features Eigen Matrix S405, apply two consecutive fully-connected layers and the ReLU non-linear activation function to the fused feature matrix X L to obtain the intermediate feature matrix X mid ; S406, intermediate feature matrix X mid Generate two-way feature outputs through two groups of parallel dilated multi-scale convolutional layers respectively. One way passes through the tanh activation function and serves as the filter feature matrix X tanh , and the other way passes through the Sigmoid activation function and serves as the gating feature matrix X sigmoid . Subsequently, the two-way feature matrices combine with the gating mechanism to calculate the final temporal convolutional output matrix Y = tanh(X tanh ) ⊙ σ(X sigmoid ); where ⊙ represents the element-wise multiplication operation, and X tanh is the one that has passed through the tanh activation Filter feature matrix of the live function, X sigmoid is the gated feature matrix after passing through the Sigmoid activation function, tanh() represents the tanh activation function, and σ() represents the Sigmoid activation function; S407. The temporal convolutional module outputs the feature matrix Y containing the extracted transient features and steady-state features, which is used as the optimized feature signal for subsequent processing. Feature signal for subsequent processing.
5. A lightweight voice enhancement method for low-altitude aircraft according to claim 1, characterized in that, In S5, the small-scale separable convolution specifically includes; S501, receive the feature matrix corresponding to the optimized feature signal output by the time convolution module where C Where \(C\) is the number of channels, \(T\) is the number of time frames, and \(F\) is the number of frequency points; S502, perform a convolution operation on each channel in the feature matrix Y separately to extract the time and frequency features within that channel. The output Y d (c, t, f) of the c-th channel is as follows: d (c, t, f) is: Among them, the Kd depth convolution kernel has a size of k d,t ×k d,f , k d,t represents the size of the convolution kernel in the time dimension, and k d,f represents the size in the frequency dimension. i and j are the indices of the convolution kernel in the time and frequency directions respectively; S503, use a 1×1 convolution kernel K for the depth convolution output Yd of S502 p , perform a linear combination of the outputs of each channel to achieve cross-channel information fusion and channel compression; Specifically, the output Y p (t, f) is calculated for each time frame and frequency point as follows: Among which K p (c) is the convolution weight of the corresponding channel; this step compresses the number of channels from to a lower value, and according to the set channel compression ratio, the final output number of channels is C / r; S504, apply batch normalization and activation using the ReLU activation function to the output Y obtained in step S503 p sequentially to generate the final output feature matrix, which serves as the enhanced speech features after small-scale separable convolution processing.
6. The lightweight voice enhancement method for a low-altitude aircraft according to claim 1, wherein, In S5, the mask calculation module specifically includes: S601. Generate a speech mask value according to the enhanced feature map; S602. Use the speech mask to adjust the time-frequency features after noise suppression; S603. Restore the adjusted time-frequency features to the clear time-domain speech signal through the inverse short-time Fourier transform.
7. A lightweight voice enhancement model applied to low-altitude aircraft, characterized in that, The lightweight speech enhancement method applied to a low-altitude aircraft according to any one of claims 1-6 specifically is: It includes a preprocessing module, a coordinate attention module, a temporal convolutional module, a separable convolutional module, and a mask calculation module; The input of the lightweight speech enhancement model is the noisy speech signal collected by the microphone; The preprocessing module is connected to the coordinate attention module, the coordinate attention module is connected to the temporal convolution module, the temporal convolution module is connected to the small-scale separable convolution module, and the small-scale separable convolution module is connected to the mask calculation module; The output time-domain speech signal is obtained through the mask module and used as the output of the lightweight speech enhancement model. The time-domain speech signal is played through an audio playback device.