A Speaker Recognition Method for Aeronautical Voice Communication Facing Short Speech and Complex Noise
Through adaptive filtering and multi-scale feature fusion technology, the ECAPA-TDNN model is optimized, combined with the Bi-GRU module, the problem of noise and phrase effects in land and air calls is solved, and high-precision speaker recognition is achieved, which improves the stability and robustness of the recognition.
Patent Information
- Application Number
- CN202510609028.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Traditional voiceprint recognition technology is affected by complex noise interference and phrases in land-air call scenarios, resulting in a decrease in recognition accuracy and stability, making it difficult to adapt to individual differences and context changes, and lacks contextual correlation capabilities.
Adaptive filtered noise suppression, multi-scale feature extraction and deep learning optimization methods are adopted, combined with the ECAPA-TDNN model and Bi-GRU module, through noise suppression, multi-scale feature fusion and data enhancement technology, the speech signal quality and feature expression capabilities are improved, context information is captured, and high-precision identity recognition is achieved.
Under complex noise and phrase conditions, the stability and accuracy of identification are improved, the security risks of identity misjudgment are reduced, and the security of aviation communications is enhanced.
Smart Images

Figure CN120148524B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of voiceprint recognition, and particularly to a method for identifying speakers in land-air communication for short speech and complex noise. Background Art
[0002] Voiceprint recognition, also known as speaker recognition, is a biometric technology that identifies the identity of a speaker by analyzing the acoustic features in a voice signal. Each person's voice has unique biological characteristics, such as vocal cord vibration, pitch, intonation, etc., which provide a reliable basis for identity verification. Voiceprint recognition technology has been widely applied in fields such as security monitoring, telephone banking, identity authentication, etc., but its accuracy is often affected by factors such as environmental noise, the quality of voice devices, and individual differences of speakers.
[0003] In the land-air communication scenario, the voice signal faces complex noise interference, including radio signal interference, aircraft engine noise, background call noise, and multipath propagation interference, etc. Traditional voiceprint recognition technology usually relies on relatively long voice segments for feature extraction. However, in the actual land-air communication environment, the communication sentences are often short, and the background noise is complex and variable, such as radio interference, cabin noise, and environmental echoes. These factors will significantly affect the extraction and matching of voiceprint features, thereby reducing the recognition accuracy and increasing the risk of misrecognition or rejection. To improve the voice recognizability, traditional noise processing techniques (such as spectral subtraction, Wiener filtering) are used to suppress non-target signals, but these methods have limited effects in a dynamic noise environment. In addition, land-air communication usually occurs in the form of short speech, and the duration of the voice segment is generally less than 8 seconds, which belongs to the typical short utterance scenario. Short speech processing technology needs to extract effective information within a limited duration, which poses higher requirements for tasks such as speech recognition and speaker recognition.
[0004] The existing methods for identifying the identities of air traffic controllers and pilots mainly rely on voiceprint recognition technology, and its process includes the following steps:
[0005] 1) Voice signal acquisition and preprocessing: First, collect the voice signals of air traffic controllers or pilots, and perform processing such as frame segmentation and windowing through voice processing technology to ensure the stability of subsequent feature extraction and modeling;
[0006] 2) Voiceprint feature extraction: Extract voiceprint features from the clear voice signals collected by the voice device. These features include information such as the short-time energy, average amplitude, and short-time average zero-crossing rate of the voice. After feature extraction, a high-dimensional voiceprint feature vector with speaker identity information is formed;
[0007] 3) Voiceprint Matching and Authentication: Compare the extracted voiceprint features with the registered voiceprint database of air-ground calls to determine whether the identity of the speaker is consistent with the recorded one. Common voiceprint matching algorithms include Gaussian Mixture Model (GMM), Support Vector Machine (SVM), and Deep Neural Network (DNN). In practical applications, a similarity threshold is set during the voiceprint matching process. Only when the similarity between the input voice and the voiceprint features stored in the database exceeds the threshold will the system recognize the identity as successfully matched. Otherwise, the system may request re-verification or trigger a warning for failed authentication.
[0008] However, the above methods have significant defects in practical applications:
[0009] 1) Complex noise affects voiceprint feature extraction: Traditional voiceprint recognition methods are vulnerable to high noise interference (such as radio signal interference, background call noise, aircraft engine noise, etc.) in the air-ground call environment, resulting in a decline in the quality of voice signals. The noise masking effect makes the acoustic feature extraction process vulnerable to interference, thus reducing the accuracy and stability of speaker identification.
[0010] 2) Short speech leads to insufficient feature information: The communication in air-ground calls is usually in the form of short speech. Traditional speaker recognition methods rely on longer speech segments to extract personalized features, but the short speech duration is limited, resulting in insufficient expression of speaker identity features and affecting the discrimination ability of the model. In addition, under the condition of short speech, it is difficult for the system to establish a robust temporal pattern, further reducing the reliability of recognition.
[0011] 3) Unstable recognition: Existing technologies often have difficulty coping with the recognition instability caused by individual differences of air traffic controllers and context changes during the voiceprint recognition process. There are significant differences in the acoustic features of different air traffic controllers, such as timbre, speech rate, accent, etc. Traditional voiceprint recognition models usually rely on fixed-mode feature extraction methods and are difficult to fully adapt to the differences between individuals, resulting in a decline in recognition accuracy.
[0012] 4) Lack of context association ability: Traditional voiceprint recognition methods often only focus on local features, and the extracted speaker information is relatively limited, making it difficult to fully capture the global information of the speech content. In air-ground calls, especially in multi-round conversation scenarios, only focusing on local features may cause the model to be unable to understand the continuity and context changes in the context, resulting in a low accuracy of speaker recognition.
[0013] Therefore, there is an urgent need for a voiceprint recognition method that can adapt to short speech, complex noise, and dynamic context to improve the identity recognition accuracy and robustness in the air-ground call scenario and reduce the aviation safety risks caused by misjudgment of identity. Summary of the Invention
[0014] In view of the above problems, the purpose of the present invention is to provide a method for speaker recognition in air-ground communication for short speech and complex noise. The method improves the noise resistance of the model under complex noise by using an adaptive filtering noise suppression method, and adopts multi-scale feature enhancement and data augmentation optimization strategies to improve the ability to extract voiceprint features under short speech conditions, so as to achieve more accurate and robust identity recognition in the air-ground communication environment. The technical solution is as follows:
[0015] A method for speaker recognition in air-ground communication for short speech and complex noise, comprising the following steps:
[0016] Step 1: Input of speech signal;
[0017] The speech signals of air traffic controllers and pilots are collected in real time through aviation communication equipment;
[0018] Step 2: Preprocessing of speech signal;
[0019] Noise suppression is performed on the speech signal, including modeling the reference noise signal by using an adaptive filter, dynamically adjusting the filter weights through the least mean square algorithm, and suppressing background noise;
[0020] Step 3: Multi-scale feature extraction;
[0021] Multi-scale feature extraction and fusion are performed on the noise-reduced speech signal, including short-time Fourier transform, filter bank energy feature extraction, and multi-scale convolutional feature extraction, and data augmentation is combined with random time masking and frequency masking;
[0022] Step 4: Training the improved ECAPA-TDNN model;
[0023] The improved ECAPA-TDNN model is trained with real air-ground communication speech data. The improved ECAPA-TDNN model includes an SE-Res2Block module and a bidirectional gated recurrent unit, which are used to extract local voiceprint features and capture context time series information; the AAM-Softmax loss function is introduced during the training process to improve the identification ability of the model;
[0024] Step 5: Identity matching and determination;
[0025] The audio sample to be recognized is input into the trained improved ECAPA-TDNN model, the corresponding high-dimensional feature vector is extracted, and the cosine similarity between the voiceprint embedding vector of the speech to be recognized and the voiceprint embedding vectors of the registered speakers in the voiceprint database is calculated, and the identity corresponding to the highest similarity is selected as the recognition result.
[0026] The beneficial effects of the present invention are:
[0027] 1. The present invention effectively solves the limitations of traditional voiceprint recognition methods in complex land-air communication environments by introducing several key technologies such as noise suppression, feature enhancement, and deep learning optimization.
[0028] 2. First, regarding the impact of complex noise on voiceprint feature extraction, the present invention adopts an adaptive filtering noise reduction technology to reduce the influence of aircraft engine noise, radio interference, etc. on the quality of the speech signal, improve the quality of the speech signal, and thus improve the stability and reliability of voiceprint recognition.
[0029] 3. Second, regarding the problem of insufficient short speech feature information, the present invention proposes a multi-scale feature extraction and fusion method, combining filter bank energy features (Fbank), short-time Fourier transform (STFT), and multi-scale convolution (MSC) to enhance the expression ability of speech features. In addition, combined with data augmentation techniques (such as random time masking, frequency masking), the adaptability of the model to short speech and complex environments is improved, thereby enhancing the robustness of voiceprint recognition.
[0030] 4. To solve the recognition instability caused by context changes, the present invention adopts the ECAPA-TDNN deep learning model and adds a Bi-GRU component after SE-Res2Block to strengthen the context information modeling ability of the model, improve the stability of short speech recognition, and make it maintain higher recognition accuracy in multi-round dialogue scenarios; finally, the most similar identity is matched through cosine similarity calculation to achieve high-precision voiceprint recognition.
[0031] 5. The present invention can provide a more accurate and stable identity recognition solution under the conditions of short speech, complex noise, and dynamic context, improve the security of air traffic control communication, reduce the security risks caused by incorrect identity recognition, and provide an important guarantee for aviation communication security. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flowchart for identifying the identity of the speaker in land-air communication.
[0033] Figure 2 It is a flowchart for voiceprint recognition.
[0034] Figure 3 It is a flowchart for training the ECAPA-TDNN network.
[0035] Figure 4 It is a flowchart for the Bi-GRU module. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] The following further elaborates on the present invention in detail with reference to the drawings and specific embodiments.
[0037] In view of the limitations of traditional methods, the present invention proposes a voiceprint recognition method for short speech and complex noise environments. By using deep learning technology to optimize the identity recognition model and introducing a context correlation module Bi-GRU into the ECAPA-TDNN model, the adaptability of the system to the speaker recognition task in the air-ground communication environment is enhanced, thereby improving the recognition accuracy and robustness in complex scenarios. The specific process of the present invention is as Figure 1 shown.
[0038] The present invention proposes an improved voiceprint recognition method to address the challenges of short speech and complex noise in the air-ground communication environment and improve the accuracy and stability of controller and pilot identity recognition. First, the system collects call voices in real time and performs noise suppression through a voice preprocessing module to effectively reduce the impact of environmental factors such as radio interference, background call noise, and aircraft engine noise on the voice quality. Then, in view of the problem of insufficient feature information caused by the limited duration of short speech, the present invention introduces multi-scale feature extraction and data augmentation technologies to enhance the model's ability to extract key information from short speech. During the voiceprint recognition process, the system uses an improved ECAPA-TDNN and Bi-GRU model for feature extraction and identity matching. First, the ECAPA-TDNN model uses a multi-scale feature extraction mechanism to deeply model the local features of the voice signal and combines the ResNet network with squeeze-and-excitation (SE-Res2Block) to enhance the expression of key voiceprint information. Subsequently, the Bi-GRU module performs temporal modeling on the features extracted by the ECAPA-TDNN, and uses the context modeling ability of its bidirectional gated recurrent unit (GRU) to capture the context information in multiple rounds of conversations, ensuring that the system can more accurately identify the identity features in short speech. Finally, the system calculates the cosine similarity between the input voice and the voiceprint features stored in the database, compares it with a set threshold, and selects the matching result with the highest similarity as the final identity determination, thereby accurately obtaining the specific controller or pilot identity corresponding to the voice.
[0039] 1. Voice signal input:
[0040] The voice signal input module of the present invention is responsible for collecting the voice data of controllers and pilots in real time, providing stable data input for subsequent voiceprint recognition. The system collects the call voice data of controllers and pilots in real time through the aviation communication system and ground receiving equipment. Due to the complex air-ground communication environment, the voice signal may be affected by factors such as radio noise, aircraft engine noise, and microphone quality. Therefore, the voice collection device needs to have a high signal-to-noise ratio and be able to adapt to multiple communication channels (such as VHF, UHF, or digital data link).
[0041] 2. Voice signal preprocessing:
[0042] To improve the accuracy of voiceprint recognition, the system preprocesses the collected voice data, and the steps are as follows:
[0043] (1) Noise suppression: In a complex aviation communication environment, the voice signal is easily interfered by various background noises, such as aircraft engine noise, radio static interference, wind shear noise, and multipath propagation interference. These noises will reduce the clarity of the voice signal and affect subsequent voiceprint recognition. Therefore, the present invention adopts a noise suppression method based on adaptive filtering, and uses an adaptive filter to process the input voice signal in real time. First, the system obtains a reference noise signal, and continuously adjusts the filter weights using the least mean square (LMS) adaptive algorithm to make the filtered signal approximate the pure voice as much as possible. By iteratively updating the filter parameters, it adaptively tracks the dynamic changes of the noise, thereby effectively suppressing the background noise, improving the signal-to-noise ratio of the input voice, and ensuring the accuracy of subsequent voiceprint recognition and the robustness of the system. The specific steps are as follows:
[0044] 1) Modeling of the input signal:
[0045] Let the mixed voice signal collected in the aviation communication environment be:
[0046] ;
[0047] where, is the observed signal (noisy voice), is the pure voice signal, is the background noise.
[0048] At the same time, randomly intercept the recorded ground-air communication voice as the reference noise signal ;
[0049] 2) Modeling of the adaptive filter:
[0050] Use an adaptive filter to perform linear filtering on to generate a noise estimation signal :
[0051] ;
[0052] where, is the filter order (the number of filter weights), is the weight coefficient of the adaptive filter, is the historical sample of the reference noise signal.
[0053] 3) Calculation of the error signal:
[0054] Calculate the error signal (i.e., the estimated pure voice):
[0055] ;
[0056] Among them, As the filtered speech output signal, the goal is to be as close as possible to the real speech .
[0057] 4) Update the weights using the LMS algorithm:
[0058] To enable the filter to adapt to changes in noise, the least mean square (LMS) algorithm is used to iteratively update the weights. The weight update formula of the LMS algorithm is as follows:
[0059] ;
[0060] Among them, is the step size factor (learning rate), which controls the weight update speed and needs to be selected within an appropriate range to ensure convergence.
[0061] When the error is large, the weight adjustment amplitude is large, enabling the system to converge quickly; when the error is small, the adjustment amplitude decreases, making the system stable. Steps 3) and 4) are continuously repeated until the filter weights are stable, that is, the error signal converges, achieving the best noise suppression effect. Finally, the error signal is the output speech signal after noise suppression for use in the subsequent speaker recognition system.
[0062] (2) After noise suppression processing, in order to further improve the discriminability of the speech signal, the present invention adopts a multi-scale feature extraction and fusion method and combines data augmentation techniques to extract richer and more robust speech features. Specifically, the system extracts multi-level feature representations from different time scales and frequency ranges to capture local and global information of the speech signal. First, the filter bank energy feature (Filter Bank, Fbank) is used to extract the basic energy distribution of the speech spectrum, and at the same time, multi-scale convolution is combined to enhance the feature expression ability of different frequency bands. Second, the short-time Fourier transform (short-time Fourier transform, STFT) and Mel spectrogram (Mel-Spectrogram) are introduced for feature supplementation to enrich the time-frequency information of the speech. In addition, the system uses data augmentation techniques (random time masking, frequency masking) to improve the generalization ability of the model and enhance its robustness in different environments.
[0063] Compared with the single feature extraction method, this method can represent the speech signal more comprehensively and provide more stable and reliable feature input for subsequent speaker recognition. The specific steps are as follows:
[0064] 1) Short-Time Fourier Transform (STFT):
[0065] The reference noise signal after noise reduction After short-time Fourier transform, it can be expressed as:
[0066] ;
[0067] where represents the coefficient of the time-frequency matrix, is the window function (Hamming window), is the time frame index, is the frequency index, is the window length of STFT, j is the imaginary number; is the discrete time step index of the original speech signal, used to traverse each sampling point of the signal. It is used to represent the point-by-point sampling value of the speech signal in the time domain and participates in the short-time Fourier transform calculation. The time-frequency representation extracted by STFT serves as the basis for subsequent feature extraction, providing rich time-frequency information for multi-scale feature fusion.
[0068] 2) Filter bank energy feature (Fbank) is used to calculate the energy distribution of the speech signal in different frequency ranges. The calculation formula is as follows:
[0069] ;
[0070] where represents the filter index (Filter Index), indicating the th Mel filter. is the energy of the th filter, is the response of the Mel filter, are the lower frequency limit value and the upper frequency limit value of the th Mel filter respectively. The Fbank feature can retain more speech spectrum information and is helpful for subsequent multi-scale feature extraction.
[0071] 3) Multi-scale convolution feature extraction:
[0072] To further enhance the feature expression ability, the system uses multi-scale convolution (Multi-Scale Convolution, MSC) to model the features in different frequency bands. The calculation formula is as follows:
[0073] ;
[0074] where is the input feature, is the The weights of the layer convolution kernels, is the bias term, is the non-linear activation function (ReLU), represents different scales of multi-scale convolution. This method learns local and global time-frequency information through convolution kernels of different sizes to achieve feature fusion.
[0075] 4) To improve the robustness of the model, the system adopts random time masking and frequency masking, and their formulas are as follows:
[0076] Time Masking:
[0077] ;
[0078] where, is the starting time frame index, is the masking range.
[0079] Frequency Masking:
[0080] ;
[0081] where, is the starting frequency index, is the masking range. This method can effectively simulate the loss of speech signals and improve the generalization ability of the model.
[0082] 5) Comprehensive feature representation:
[0083] Finally, by combining Fbank features, multi-scale convolution, STFT and data augmentation methods, the final feature representation is obtained:
[0084] ;
[0085] where, represents the short-time Fourier transform, represents the filter bank energy feature extraction, represents the multi-scale convolution feature extraction, represents the data augmentation method. Compared with the single feature extraction method, this method can represent speech signals more comprehensively and provide more stable and reliable feature inputs for subsequent speaker recognition.
[0086] 3. Speaker recognition:
[0087] The present invention uses voiceprint recognition technology as one of the core means of identity verification for controllers and pilots to ensure the security and accuracy of voice communication. Voiceprint recognition is a biometric-based identity recognition technology that uses the speaker's personalized voice features for identity verification. Compared with traditional identity verification methods (such as passwords or card recognition), voiceprint recognition has the advantages of being contactless, highly real-time, and suitable for remote communications. However, due to the particularity of the land-to-air call communication environment, such as factors such as aircraft engine noise, radio interference, and signal attenuation, the accuracy of voiceprint recognition may be affected. In addition, there are significant individual differences between controllers and pilots, which also increases the difficulty of recognition. To this end, the present invention introduces the ECAPA-TDNN model in the voiceprint recognition process to extract voiceprint features to improve the robustness and recognition accuracy of the system in complex environments. The specific process is as follows Figure 2 shown.
[0088] (1) Training phase:
[0089] During the training phase, the system first pre-trains the model using about 1,000 hours of real land and air call voice data (with speaker labels annotated) to obtain a speaker recognition model with good recognition performance. This model serves as an initialization model to provide stable feature representation capabilities for subsequent training, ensuring that the voiceprint recognition system can converge efficiently. In order to optimize the performance of the model and enhance its ability to distinguish different speakers, the AAM-Softmax Loss (Additive Angular Margin-Softmax Loss) loss function is introduced during the training process for training guidance. This loss function can effectively reduce the intra-class distance of the voice features of the same speaker, while increasing the inter-class gap between different speakers, thereby improving the recognition ability of the model. The entire training process continues, and the system continuously evaluates the performance of the model on the validation set until the loss function reaches the optimal solution and tends to be stable, the training is terminated, and a high-performance voiceprint recognition model is finally obtained.
[0090] The initial neural network selects the ECAPA-TDNN network for training. The training process is as follows Figure 3 shown.
[0091] In the training phase, we first input the 80-dimensional multi-scale speech features ( , The one-dimensional convolution operation (Conv1D) is performed with the time step, and the ReLU activation function and batch normalization (BN) are used for feature transformation to ensure stable data distribution. The calculation process is as follows:
[0092] ;
[0093] in, is the weight of the 1D convolutional kernel, is the bias term, the convolutional kernel size is 5; the dilation coefficient is 1. The output , , where represents the number of channels (Channels), ReLU is the activation function, and Conv1D is the one-dimensional convolution operation.
[0094] Next, deep feature extraction is performed through the SE-Res2Block. This module contains three convolutional units, and the convolutional kernel size is set to , and dilated convolutions with dilation rates are used respectively to enhance the receptive field. At the same time, the Squeeze-and-Excitation (SE) mechanism is combined to adaptively adjust the channel weights and improve the feature expression ability. The output of the i-th SE-Res2Block ( ) is calculated as follows:
[0095] ;
[0096] where the dilated convolution uses unchanged, and the dilation coefficient d increases with the layer. X i and X i+1 are the input and output of the i-th SE-Res2Block module respectively; SE-Res2Block represents the processing of the SE-Res2Block module.
[0097] The outputs of the three-layer SE-Res2Block are concatenated: , where Concat represents the concatenation operation.
[0098] After the concatenated SE-Res2Block, in order to further improve the temporal modeling ability and global context dependence of the speech features, the present invention introduces a bidirectional gated recurrent unit (Bi-GRU) to process the feature sequence, as shown in Figure 4 . The Bi-GRU is composed of a forward GRU and a backward GRU, and can capture the forward and backward dependence information of the speech signal in the time dimension at the same time, so as to effectively model the temporal features across frames. Among them, X1, X2, X3, X4 represent the input at each time step; X1 corresponds to the input at the first time step, X2 corresponds to the input at the second time step, and so on; Y1, Y2, Y3, Y4 represent the final output at each time step.
[0099] 1) After the SE-Res2Block is concatenated, is sent to the Bi-GRU for bidirectional sequence modeling. The Bi-GRU is composed of GRUs in two directions, and the forward and backward hidden states are calculated respectively:
[0100] ;
[0101] ;
[0102] Among them, is the hidden state of the forward and backward GRUs, H is the number of GRU hidden units. GRU fwd is the forward GRU process, and GRU bwd is the backward GRU process.
[0103] 2) Bi-GRU concatenates the forward and backward information and outputs :
[0104] ;
[0105] Among them, 2H represents the feature dimension after concatenating the forward and backward GRU hidden states.
[0106] The output passes through the convolutional activation normalization layer:
[0107] ;
[0108] Among them, and are the weights and biases of this layer.
[0109] In order to extract the temporal information, the Attentive Statistics Pooling (ASP) layer is used to calculate the global features:
[0110] ;
[0111] Next, the vector after passing through the pooling layer enters the fully connected normalization layer (FC+BN) to output a 192-dimensional voiceprint embedding code vector :
[0112] ;
[0113] Finally, the AAM-Softmax loss function is used to guide the training until the model training converges to obtain a seed model with better performance.
[0114] ;
[0115] Among them, represents the true label, is the average loss value of the training samples; AAMLoss is the applied additive angular margin loss function.
[0116] (2)Testing Phase:
[0117] In the testing phase, first, L registered audio samples are input into the trained ECAPA-TDNN network to extract their high-dimensional feature vectors to represent the unique voiceprint information of the speaker. For multiple audio samples of the same speaker, the mean of their feature vectors is calculated to obtain the mean feature vector (speaker template) of the speaker. These speaker templates together constitute the voiceprint database in the testing phase, which is used to store the feature information of the registered speakers.
[0118] Subsequently, the audio sample to be recognized (Speaker A) is input into the same neural network, the corresponding high-dimensional feature vector is extracted, and the cosine similarity between it and each speaker template in the voiceprint database is calculated. The speaker corresponding to the template with the highest similarity is taken, and the speaker identification task is completed. The specific calculation process is as follows:
[0119] ;
[0120] where, represents the voiceprint embedding vector of the speech to be measured, represents the voiceprint embedding vector of a certain speaker in the voiceprint library, represents the calculated cosine similarity, ranging from [-1, 1]. The closer the value is to 1, the more similar the two are; represents the norm of the vector.
[0121] To sum up, in the testing phase, by calculating the cosine similarity between the audio sample to be recognized and the registered speaker templates, the system can effectively match the most similar identity and achieve high-precision speaker recognition. This method uses ECAPA-TDNN to extract high-dimensional voiceprint features and constructs a voiceprint database through the mean feature vector, thereby improving the stability and robustness of recognition. In addition, due to the use of cosine similarity measurement, the system can maintain strong discrimination ability in different noise environments and ensure the accuracy of identity matching. This testing process lays a reliable technical foundation for the practical application of speaker identification, enabling the system to be applicable to complex land-air communication scenarios and achieve efficient and secure identity authentication.
Claims
1. A method for speaker recognition in land-air communication facing short speech and complex noise, characterized in that Including the following steps: Step 1: Voice signal input; Real-time collection of voice signals between air traffic controllers and pilots through aviation communication equipment; Step 2: Voice signal preprocessing; Perform noise suppression on the voice signal, including modeling the reference noise signal using an adaptive filter, dynamically adjusting the filter weights through the least mean square algorithm, and suppressing background noise; Step 3: Multi-scale feature extraction; Perform multi-scale feature extraction and fusion on the denoised voice signal, including short-time Fourier transform, filter bank energy feature extraction, and multi-scale convolutional feature extraction, and perform data augmentation by combining random time masking and frequency masking; Step 4: Train the improved ECAPA-TDNN model; Train the improved ECAPA-TDNN model with real air-ground communication voice data. The improved ECAPA-TDNN model includes an SE-Res2Block module and a bidirectional gated recurrent unit, which are used to extract local speaker feature and capture context time series information; the AAM-Softmax loss function is introduced during the training process to improve the identification ability of the model; Step 5: Identity matching and determination; Input the audio sample to be recognized into the trained improved ECAPA-TDNN model, extract the corresponding high-dimensional feature vector, calculate the cosine similarity between the speaker feature embedding vector of the voice to be recognized and the speaker feature embedding vectors of the registered speakers in the speaker database, and select the identity corresponding to the highest similarity as the recognition result; In the said Step 4, training the improved ECAPA-TDNN model specifically includes: Step 4.1: Perform one-dimensional convolution operation on the input multi-scale voice feature X0, perform feature transformation using the ReLU activation function and batch normalization, and the calculation process is as follows: ; Among them, is the weight of the 1D convolution kernel, is the bias term; the output signal , represents the number of channels, is the time step; is for batch normalization processing; ReLU is the activation function, and Conv1D is the one-dimensional convolution operation; Step 4.2: Perform deep feature extraction through the SE-Res2Block module; the SE-Res2Block module includes three convolutional units, and the convolutional kernel size is set to , and dilated convolutions with dilation rates of are respectively used to enhance the receptive field, and at the same time, the SE mechanism is combined to adaptively adjust the channel weights; the output of the th SE-Res2Block module is calculated as follows: ; Among them, the dilated convolution adopts unchanged, and the dilation coefficient d increases with the layer; ; X i and X i+1 are the input and output of the i th SE-Res2Block module respectively; SE-Res2Block represents the processing of the SE-Res2Block module; The outputs of three layers of SE-Res2Block are concatenated: ; Where Concat represents the concatenation operation; Step 4.3: Introduce a bidirectional gated recurrent unit to process the feature sequence, where the bidirectional gated recurrent unit consists of a forward GRU and a backward GRU; send the concatenated output to the introduced bidirectional gated recurrent unit for bidirectional sequence modeling, and calculate the forward and backward hidden states respectively: ; ; wherein, , and are the hidden states of the forward GRU and the backward GRU, respectively; is the number of GRU hidden units; GRU fwd is processed by the forward GRU, and GRU bwd is processed by the backward GRU; The bidirectional gated recurrent unit outputs after concatenating the hidden states of the forward GRU and the backward GRU : ; Among them, represents the feature dimension after concatenating the hidden states of the forward GRU and the backward GRU; T is the time step; Output after passing through the convolutional activation normalization layer: ; Among them, , and are the output, weight, and bias of the convolutional activation normalization layer, respectively; Calculate the global feature using the attention pooling layer to obtain a vector : ; Among them, ASP is for the calculation of the attention pooling layer; The vector that then passes through the pooling layer enters the fully connected normalization layer and outputs the voiceprint embedding code vector : ; Where FC is the full connection layer processing; Step 4.5: Use the AAM-Softmax loss function to guide the training until the model training converges to obtain a seed model that meets the performance requirements: ; Among them, represents the true label, is the average loss value of the training samples; AAMLoss is the application of the additive angular margin loss function.
2. The method for identifying a speaker in an air-ground call for short speech and complex noise according to claim 1, wherein The noise suppression in the said Step 2 specifically includes the following sub-steps: Step 2.1: Input signal modeling; Assume that the mixed voice signal collected in the aviation communication environment is: ; Among them, is the observed signal of the noisy speech, is the clean speech signal, is the background noise; Meanwhile, randomly intercept the already recorded voice of the ATC communication as the reference noise signal ; Step 2.2: Adaptive filter modeling; An adaptive filter is used to linearly filter the reference noise signal to generate a noise estimation signal : ; Among them, is the filter order, that is, the number of weights of the filter; is the weight coefficient of the adaptive filter, is the historical sample of the reference noise signal; Step 2.3: Error signal calculation; Calculate the error signal, that is, the estimated clean voice: ; Step 2.4: Update the weights through the least mean square algorithm; The weight update formula of the least mean square algorithm is as follows: ; Among them, is the step factor, that is, the learning rate; Step 2.5: Continuously repeat Step 2.3 and Step 2.4 until the error signal converges, then use the error signal as the output speech signal after noise suppression.
3. A method for identifying speakers in land-air communication for short speech and complex noise according to claim 1, characterized in that The multi-scale feature extraction in the said Step 3 specifically includes: Step 3.1: Short-time Fourier transform; Reference noise signal after noise reduction processing After short-time Fourier transform, it is expressed as: ; Among them, represents the coefficient of the time-frequency matrix, is the window function, is the time frame index, is the frequency index, is the window length of the short-time Fourier transform; is the discrete time step index; j is the imaginary number; Step 3.2: Filter bank energy feature extraction, and the calculation formula is as follows: ; Among them, represents the filter index, indicating the th Mel filter; is the energy of the th filter, is the response of the Mel filter; are respectively the th lower and upper frequency limit values of the Mel filter; Step 3.3: Multi-scale convolutional feature extraction F l ; Use multi-scale convolution to model the features of different frequency bands, and its calculation formula is as follows: ; Among them, is the input feature, is the weight of the layer convolutional kernel, is the bias term, is the non-linear activation function, represents different scales of multi-scale convolution; Step 3.4: Perform data augmentation using random time masking and frequency masking, and its formula is as follows: Time masking: ; Among them, is the starting time frame index, is the occlusion range; Frequency masking: ; Among them, is the starting frequency index, is the masking range; Step 3.4: Comprehensive feature representation; Combining the energy features of filter banks, multi-scale convolutional features, short-time Fourier transform, and data augmentation methods to obtain the final feature representation F : ; Among them, represents the short-time Fourier transform, represents the filter bank energy feature extraction, represents the multi-scale convolutional feature extraction, represents the data augmentation method.
4. A method for identifying a speaker in an air-ground communication facing short speech and complex noise according to claim 1, characterized in that, The cosine similarity calculation in the said Step 5 is as follows: ; Among them, represents the voiceprint embedding vector of the voice to be measured, represents the voiceprint embedding vector of a certain speaker in the voiceprint database; represents the norm of the vector; represents the calculated cosine similarity, with a range of [-1, 1]. The closer the value is to 1, the more similar the two are.
Citation Information
Patent Citations
Identity feature extraction method and device based on acoustic feature generation and storage medium
CN116631406A
Audio processing method and device, equipment, medium and product
CN117153175A