Land-air communication speaker recognition method facing short voice and complex noise

By adopting adaptive filtered noise suppression, multi-scale feature extraction and deep learning optimization methods in land-air call environments, the recognition accuracy and stability of traditional voiceprint recognition technology in complex noise and phrase speech environments is solved, and a more efficient and robust identity recognition effect is achieved.

CN120148524AActive Publication Date: 2025-06-13CIVIL AVIATION FLIGHT UNIV OF CHINA

Patent Information

Application Number
CN202510609028.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-13
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

Traditional voiceprint recognition technology is affected by complex noise, insufficient information on phrase speech characteristics and identification instability in land-to-air call environments, resulting in a decrease in recognition accuracy and stability.

Method used

Adaptive filtering noise suppression technology is used to reduce noise interference, combine multi-scale feature extraction and data enhancement technology to improve the vocalprint feature extraction capability under phrase speech conditions, and introduce Bi-GRU module into the ECAPA-TDNN model to enhance context information modeling capabilities.

Benefits of technology

It improves noise resistance and recognition stability in complex noise environments, enhances the vocalprint feature extraction capability under phrase speech conditions, and achieves more accurate and robust identity recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148524A_ABST
    Figure CN120148524A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voiceprint recognition, and discloses an air-ground conversation speaker recognition method for short voice and complex noise, which dynamically suppresses background noise through a self-adaptive filter in combination with a minimum mean square algorithm, and improves the voice signal quality. A multi-scale feature fusion technology is adopted, short-time Fourier transform, filter bank energy features and multi-scale convolution are combined, a random time and frequency shielding data enhancement strategy is introduced, and the short speech feature expression ability is enhanced; bi-GRU is integrated in the ECAPA-TDNN model, and context information of multiple rounds of conversations is captured through context time sequence modeling, so that the recognition stability in a dynamic scene is improved; and an AAM-Softmax loss function is adopted to strengthen the distinguishing capability of the model, and finally identity judgment is realized through cosine similarity matching. According to the method, the voiceprint recognition precision and robustness under complex noise and short voice conditions are remarkably improved, and the safety risk caused by identity misjudgment in aviation communication is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voiceprint recognition, and particularly to a method for identifying speakers in air-ground calls for short speech and complex noise. Background Technique

[0002] Voiceprint recognition, also known as speaker recognition, is a biometric technology that identifies the identity of a speaker by analyzing the acoustic features in a voice signal. Each person's voice has unique biological characteristics, such as vocal cord vibration, pitch, intonation, etc., which provide a reliable basis for identity verification. Voiceprint recognition technology has been widely used in fields such as security monitoring, telephone banking, and identity authentication, but its accuracy is often affected by factors such as environmental noise, the quality of voice devices, and individual differences of speakers.

[0003] In the air-ground call scenario, the voice signal faces complex noise interference, including radio signal interference, aircraft engine noise, background call noise, and multipath propagation interference, etc. Traditional voiceprint recognition technology usually relies on relatively long voice segments for feature extraction. However, in the actual air-ground communication environment, the communication sentences are often short, and the background noise is complex and variable, such as radio interference, cabin noise, and environmental echo, etc. These factors will significantly affect the extraction and matching of voiceprint features, thereby reducing the recognition accuracy and increasing the risk of misrecognition or rejection. To improve the voice recognizability, traditional noise processing technologies (such as spectral subtraction, Wiener filtering) are used to suppress non-target signals, but these methods have limited effects in a dynamic noise environment. In addition, air-ground calls are usually made in the form of short speech, and the duration of the voice segment is generally less than 8 seconds, which belongs to the typical short utterance scenario. Short speech processing technology needs to extract effective information within a limited duration, which poses higher requirements for tasks such as speech recognition and speaker recognition.

[0004] The existing methods for identifying the identities of air traffic controllers and pilots are mainly based on voiceprint recognition technology, and its process includes the following steps: 1) Voice signal acquisition and preprocessing: First, collect the voice signals of air traffic controllers or pilots, and perform processing such as frame division and windowing through voice processing technology to ensure the stability of subsequent feature extraction and modeling; 2) Voiceprint feature extraction: Extract voiceprint features from the clear voice signals collected by the voice device. These features include information such as the short-time energy, average amplitude, and short-time average zero-crossing rate of the voice. After feature extraction, a high-dimensional voiceprint feature vector with speaker identity information is formed; 3) Voiceprint Matching and Authentication: Compare the extracted voiceprint features with the registered voiceprint database of air-ground calls to determine whether the identity of the speaker matches the recorded one. Common voiceprint matching algorithms include Gaussian Mixture Model (GMM), Support Vector Machine (SVM), and Deep Neural Network (DNN). In practical applications, a similarity threshold is set during the voiceprint matching process. Only when the similarity between the input voice and the voiceprint features stored in the database exceeds the threshold will the system recognize the identity as successfully matched. Otherwise, the system may request re-verification or trigger a warning for failed authentication.

[0005] However, the above methods have significant drawbacks in practical applications: 1) Complex noise affects voiceprint feature extraction: Traditional voiceprint recognition methods are vulnerable to high noise interference in the air-ground call environment (such as radio signal interference, background call noise, aircraft engine noise, etc.), resulting in a decline in the quality of voice signals. The noise masking effect makes the acoustic feature extraction process vulnerable to interference, thus reducing the accuracy and stability of speaker recognition.

[0006] 2) Short speech leads to insufficient feature information: Communication in air-ground calls usually occurs in the form of short speeches. Traditional speaker recognition methods rely on longer speech segments to extract personalized features, but the limited duration of short speeches leads to insufficient expression of speaker identity features, affecting the discrimination ability of the model. In addition, under the condition of short speeches, it is difficult for the system to establish a robust temporal pattern, further reducing the reliability of recognition.

[0007] 3) Unstable recognition: Existing technologies often struggle to handle the recognition instability caused by individual differences among air traffic controllers and context changes during voiceprint recognition. There are significant differences in the acoustic features of different air traffic controllers, such as timbre, speech rate, accent, etc. Traditional voiceprint recognition models usually rely on fixed-mode feature extraction methods, making it difficult to fully adapt to the differences between individuals, resulting in a decline in recognition accuracy.

[0008] 4) Lack of context association ability: Traditional voiceprint recognition methods often only focus on local features, extracting limited speaker information and making it difficult to fully capture the global information of the speech content. In air-ground calls, especially in multi-turn conversation scenarios, focusing only on local features may cause the model to fail to understand the continuity and context changes in the context, resulting in a lower accuracy of speaker recognition.

[0009] Therefore, there is an urgent need for a voiceprint recognition method that can adapt to short speeches, complex noise, and dynamic contexts to improve the identity recognition accuracy and robustness in air-ground call scenarios and reduce the aviation safety risks caused by misjudgment of identity. Summary of the Invention

[0010] In view of the above problems, the purpose of the present invention is to provide a method for identifying speakers in air-ground communication for short speech and complex noise. The method improves the noise resistance of the model under complex noise by using an adaptive filtering noise suppression method, and adopts multi-scale feature enhancement and data enhancement optimization strategies to improve the ability to extract voiceprint features under short speech conditions, so as to achieve more accurate and robust identity recognition in the air-ground communication environment. The technical solution is as follows: A method for identifying speakers in air-ground communication for short speech and complex noise, comprising the following steps: Step 1: Input of speech signals; The speech signals of air traffic controllers and pilots are collected in real time through aviation communication equipment; Step 2: Preprocessing of speech signals; Noise suppression is performed on the speech signals, including modeling the reference noise signal using an adaptive filter, dynamically adjusting the filter weights through the least mean square algorithm, and suppressing background noise; Step 3: Multi-scale feature extraction; Multi-scale feature extraction and fusion are performed on the noise-reduced speech signals, including short-time Fourier transform, filter bank energy feature extraction, and multi-scale convolutional feature extraction, and data enhancement is combined with random time masking and frequency masking; Step 4: Training the improved ECAPA-TDNN model; The improved ECAPA-TDNN model is trained with real air-ground communication speech data. The improved ECAPA-TDNN model includes an SE-Res2Block module and a bidirectional gated recurrent unit, which are used to extract local voiceprint features and capture context time series information; the AAM-Softmax loss function is introduced during the training process to improve the discrimination ability of the model; Step 5: Identity matching and determination; The audio sample to be identified is input into the trained improved ECAPA-TDNN model, the corresponding high-dimensional feature vector is extracted, and the cosine similarity between the voiceprint embedding vector of the speech to be identified and the voiceprint embedding vectors of the registered speakers in the voiceprint database is calculated. The identity corresponding to the highest similarity is selected as the recognition result.

[0011] The beneficial effects of the present invention are: 1. By introducing a number of key technologies such as noise suppression, feature enhancement, and deep learning optimization, the present invention effectively solves the limitations of traditional voiceprint recognition methods in complex air-ground communication environments.

[0012] 2. First, regarding the impact of complex noise on voiceprint feature extraction, the present invention adopts an adaptive filtering noise reduction technique to reduce the impact of aircraft engine noise, radio interference, etc. on the quality of the voice signal, improve the quality of the voice signal, and thus enhance the stability and reliability of voiceprint recognition.

[0013] 3. Second, to address the problem of insufficient feature information in short speech, the present invention proposes a multi-scale feature extraction and fusion method, combining filter bank energy features (Fbank), short-time Fourier transform (STFT), and multi-scale convolution (MSC) to enhance the expression ability of speech features. In addition, combined with data augmentation techniques (such as random time masking, frequency masking), the adaptability of the model to short speech and complex environments is improved, thereby enhancing the robustness of voiceprint recognition.

[0014] 4. To solve the recognition instability caused by context changes, the present invention uses an ECAPA-TDNN deep learning model and adds a Bi-GRU component after SE-Res2Block to strengthen the model's context information modeling ability, improve the stability of short speech recognition, and maintain higher recognition accuracy in multi-round dialogue scenarios; finally, the most similar identity is matched through cosine similarity calculation to achieve high-precision voiceprint recognition.

[0015] 5. The present invention can provide a more accurate and stable identity recognition solution under the conditions of short speech, complex noise, and dynamic context, improve the security of air traffic control communication, reduce the security risks caused by incorrect identity recognition, and provide an important guarantee for aviation communication security. Description of the Drawings

[0016] Figure 1 It is a flow chart for identifying the identity of the speaker in the ground-air communication.

[0017] Figure 2 It is a flow chart for voiceprint recognition.

[0018] Figure 3 It is a flow chart for training the ECAPA-TDNN network.

[0019] Figure 4 It is a flow chart for the Bi-GRU module. Detailed Embodiments

[0020] The following further elaborates on the present invention in detail in conjunction with the drawings and specific embodiments.

[0021] In view of the limitations of traditional methods, the present invention proposes a voiceprint recognition method for short speech and complex noise environments. The identity recognition model is optimized through deep learning technology, and a context correlation module Bi-GRU is introduced into the ECAPA-TDNN model to enhance the adaptability of the system to the speaker recognition task in the air-ground communication environment, thereby improving the recognition accuracy and robustness in complex scenarios. The specific process of the present invention is as Figure 1 shown.

[0022] The present invention proposes an improved voiceprint recognition method to address the challenges of short speech and complex noise in the air-ground communication environment and improve the accuracy and stability of controller and pilot identity recognition. First, the system collects call voices in real time, and performs noise suppression through a voice preprocessing module to effectively reduce the impact of environmental factors such as radio interference, background call noise, and aircraft engine noise on the voice quality. Then, in view of the problem of insufficient feature information caused by the limited duration of short speech, the present invention introduces multi-scale feature extraction and data augmentation techniques to improve the model's ability to extract key information from short speech. During the voiceprint recognition process, the system uses an improved ECAPA-TDNN and Bi-GRU model for feature extraction and identity matching. First, the ECAPA-TDNN model uses a multi-scale feature extraction mechanism to deeply model the local features of the voice signal, and combines the ResNet network with squeeze and excitation (SE-Res2Block) to enhance the expression of key voiceprint information. Subsequently, the Bi-GRU module performs temporal modeling on the features extracted by the ECAPA-TDNN, and uses the context modeling ability of its bidirectional gated recurrent unit (GRU) to capture the context information in multiple rounds of conversations, ensuring that the system can more accurately identify the identity features in short speech. Finally, the system calculates the cosine similarity between the input voice and the voiceprint features stored in the database, compares it with a set threshold, and selects the matching result with the highest similarity as the final identity determination, so as to accurately obtain the specific controller or pilot identity corresponding to the voice.

[0023] 1. Voice signal input: The voice signal input module of the present invention is responsible for collecting the voice data of controllers and pilots in real time, providing stable data input for subsequent voiceprint recognition. The system collects the call voice data of controllers and pilots in real time through the aviation communication system and ground receiving equipment. Due to the complex air-ground communication environment, the voice signal may be affected by factors such as radio noise, aircraft engine noise, and microphone quality. Therefore, the voice collection device needs to have a high signal-to-noise ratio and be able to adapt to multiple communication channels (such as VHF, UHF, or digital data link).

[0024] 2. Voice signal preprocessing: To improve the accuracy of voiceprint recognition, the system preprocesses the collected voice data, and the steps are as follows: (1) Noise suppression: In a complex aviation communication environment, the voice signal is easily interfered by various background noises, such as aircraft engine noise, radio static interference, wind shear noise, and multipath propagation interference. These noises will reduce the clarity of the voice signal and affect subsequent voiceprint recognition. Therefore, the present invention adopts a noise suppression method based on adaptive filtering, and uses an adaptive filter to process the input voice signal in real time. First, the system obtains a reference noise signal and continuously adjusts the filter weights using the least mean square (LMS) adaptive algorithm to make the filtered signal approximate the pure voice as much as possible. By iteratively updating the filter parameters, it adaptively tracks the dynamic changes of the noise, thereby effectively suppressing the background noise, improving the signal-to-noise ratio of the input voice, and ensuring the accuracy of subsequent voiceprint recognition and the robustness of the system. The specific steps are as follows: 1) Modeling of the input signal: Let the mixed voice signal collected in the aviation communication environment be: ; where, is the observed signal (noisy voice), is the pure voice signal, is the background noise.

[0025] At the same time, randomly intercept the recorded ground-air communication voice as the reference noise signal ; 2) Modeling of the adaptive filter: Use an adaptive filter to perform linear filtering on to generate a noise estimation signal : ; where, is the filter order (the number of filter weights), is the weight coefficient of the adaptive filter, is the historical sample of the reference noise signal.

[0026] 3) Calculation of the error signal: Calculate the error signal (i.e., the estimated pure voice): ; where, as the filtered voice output signal, the goal is to make it as close as possible to the real voice .

[0027] 4) Update the weights using the LMS algorithm: To enable the filter to adapt to the changes in noise, the least mean square (LMS) algorithm is used to iteratively update the weights. The weight update formula of the LMS algorithm is as follows: ; Wherein, is the step size factor (learning rate), which controls the weight update speed and needs to be selected within an appropriate range to ensure convergence.

[0028] When the error is large, the weight adjustment amplitude is large, enabling the system to converge quickly; when the error is small, the adjustment amplitude decreases, making the system stable. Steps 3) and 4) are continuously repeated until the filter weights are stable, that is, the error signal converges to achieve the best noise suppression effect. Finally, the error signal is the output speech signal after noise suppression and is used for the subsequent speaker recognition system.

[0029] (2) After the noise suppression process, in order to further improve the identifiability of the speech signal, the present invention adopts a multi-scale feature extraction and fusion method and combines data augmentation techniques to extract richer and more robust speech features. Specifically, the system extracts multi-level feature representations from different time scales and frequency ranges to capture the local and global information of the speech signal. First, the filter bank energy feature (Filter Bank, Fbank) is used to extract the basic energy distribution of the speech spectrum, and at the same time, multi-scale convolution is combined to enhance the feature expression ability of different frequency bands. Secondly, the short-time Fourier transform (short-time Fourier transform, STFT) and Mel spectrogram are introduced for feature supplementation to enrich the time-frequency information of the speech. In addition, the system uses data augmentation techniques (random time masking, frequency masking) to improve the generalization ability of the model and enhance its robustness in different environments.

[0030] Compared with the single feature extraction method, this method can represent the speech signal more comprehensively and provide a more stable and reliable feature input for subsequent speaker recognition. The specific steps are as follows: 1) Short-time Fourier transform (STFT): The reference noise signal after noise reduction processing After short-time Fourier transform, it can be expressed as: ; Wherein, represents the coefficient of the time-frequency matrix, is the window function (Hamming window), is the time frame index, is the frequency index, is the window length of the STFT, j is the imaginary number; is the discrete-time step index of the original speech signal, used to traverse each sampling point of the signal. It represents the point-by-point sampling value of the speech signal in the time domain and participates in the short-time Fourier transform calculation. The time-frequency representation extracted by the STFT serves as the basis for subsequent feature extraction, providing rich time-frequency information for multi-scale feature fusion.

[0031] 2) Filter bank energy features (Fbank) are used to calculate the energy distribution of the speech signal in different frequency ranges. The calculation formula is as follows: ; where, represents the filter index (Filter Index), indicating the th Mel filter. is the energy of the th filter, is the response of the Mel filter, are the lower and upper frequency limit values of the th Mel filter respectively. The Fbank features can retain more speech spectrum information, which is helpful for subsequent multi-scale feature extraction.

[0032] 3) Multi-scale convolution feature extraction: To further enhance the feature expression ability, the system uses multi-scale convolution (Multi-Scale Convolution, MSC) to model the features in different frequency bands. The calculation formula is as follows: ; where, is the input feature, is the weight of the th convolutional kernel, is the bias term, is the non-linear activation function (ReLU), represents different scales of the multi-scale convolution. This method learns local and global time-frequency information through convolutional kernels of different sizes to achieve feature fusion.

[0033] 4) To improve the robustness of the model, the system uses random time masking (Time Masking) and frequency masking (Frequency Masking). The formulas are as follows: Time Masking: ; where, is the starting time frame index, is the masking range.

[0034] Frequency Masking: ; where is the starting frequency index, is the masking range. This method can effectively simulate the loss of speech signals and improve the generalization ability of the model.

[0035] 5) Comprehensive feature representation: Finally, combining Fbank features, multi-scale convolution, STFT, and data augmentation methods, the final feature representation is obtained: ; where represents the short-time Fourier transform, represents the filter bank energy feature extraction, represents the multi-scale convolution feature extraction, represents the data augmentation method. Compared with a single feature extraction method, this method can represent speech signals more comprehensively and provide a more stable and reliable feature input for subsequent speaker recognition.

[0036] 3. Speaker recognition: The present invention uses speaker recognition technology as one of the core means for the authentication of air traffic controllers and pilots to ensure the security and accuracy of voice communication. Speaker recognition is a biometric-based identity recognition technology that uses the personalized voice characteristics of speakers for identity authentication. Compared with traditional identity authentication methods (such as password or card recognition), speaker recognition has the advantages of non-contact, strong real-time performance, and suitability for remote communication. However, due to the special nature of the air-ground communication environment, such as aircraft engine noise, radio interference, and signal attenuation, the accuracy of speaker recognition may be affected. In addition, there are significant individual differences between air traffic controllers and pilots, which also increases the difficulty of recognition. Therefore, the present invention introduces the ECAPA-TDNN model in the speaker recognition process for speaker feature extraction to improve the robustness and recognition accuracy of the system in complex environments. The specific process is as follows Figure 2 as shown.

[0037] (1) Training stage: During the training phase, the system first pre-trains the model using approximately 1000 hours of real air-ground communication voice data (with speaker labels already annotated) to obtain a speaker recognition model with good recognition performance. This model serves as the initialization model, providing stable feature representation capabilities for subsequent training and ensuring that the voiceprint recognition system can converge efficiently. To optimize the model performance and enhance its ability to distinguish different speakers, the AAM-Softmax Loss (Additive Angular Margin-Softmax Loss) function is introduced during the training process for training guidance. This loss function can effectively reduce the intra-class distance of the voice features of the same speaker while increasing the inter-class gap between different speakers, thereby improving the model's recognition ability. The entire training process continues, and the system continuously evaluates the model's performance on the validation set until the loss function reaches the optimal solution and stabilizes, at which point the training terminates, and a high-performance voiceprint recognition model is finally obtained.

[0038] The initial neural network selects the ECAPA-TDNN network for training, and the training process is as follows Figure 3 as shown.

[0039] During the training phase, first perform a one-dimensional convolution operation (Conv1D) on the input 80-dimensional multi-scale voice features ( , where is the time step), and use the ReLU activation function and batch normalization (BN) for feature transformation to ensure stable data distribution. The calculation process is as follows: ; where, is the 1D convolution kernel weight, is the bias term, the convolution kernel size is 5; the dilation coefficient is 1. The output , represents the number of channels (Channels), ReLU is the activation function, and Conv1D is the one-dimensional convolution operation.

[0040] Next, perform deep feature extraction through the SE-Res2Block. This module contains three convolutional units, and the convolution kernel size is set to , and dilated convolutions with dilation rates are used respectively to enhance the receptive field, and at the same time, the Squeeze-and-Excitation (SE) mechanism is combined to adaptively adjust the channel weights to improve the feature expression ability. The output of the i-th SE-Res2Block ( ) is calculated as follows: ;

[0041] where the dilated convolution uses Remain unchanged, the expansion coefficient d increases with the layer. X i and X i+1 are the input and output of the i-th SE-Res2Block module respectively; SE-Res2Block represents the processing of the SE-Res2Block module.

[0042] The outputs of the three-layer SE-Res2Block are concatenated: , where Concat represents the concatenation operation.

[0043] After the concatenated SE-Res2Block, in order to further improve the temporal modeling ability and global context dependence of speech features, the present invention introduces a bidirectional gated recurrent unit (Bi-GRU) to process the feature sequence, as Figure 4 shown. Bi-GRU consists of a forward GRU and a backward GRU, which can capture the forward and backward dependence information of the speech signal in the time dimension simultaneously, so as to effectively model the temporal features across frames. Among them, X1, X2, X3, X4 represent the input at each time step; X1 corresponds to the input at the 1st time step, X2 corresponds to the input at the 2nd time step, and so on; Y1, Y2, Y3, Y4 represent the final output at each time step.

[0044] 1) After the SE-Res2Block is concatenated, is sent to the Bi-GRU for bidirectional sequence modeling. Bi-GRU consists of GRUs in two directions, and the forward and backward hidden states are calculated respectively: ; ; Among them, are the hidden states of the forward and backward GRUs, H is the number of GRU hidden units. GRU fwd is the processing of the forward GRU, and GRU bwd is the processing of the backward GRU.

[0045] 2) Bi-GRU concatenates the forward and backward information and outputs : ; Among them, 2H represents the feature dimension after concatenating the forward and backward GRU hidden states.

[0046] The output passes through a convolutional activation normalization layer: ; Among them, and are the weights and biases of this layer.

[0047] To extract the temporal information, the Attentive Statistics Pooling (ASP) layer is adopted to calculate the global features: ; Then, the vector passing through the pooling layer enters the fully connected normalization layer (FC + BN) to output a 192-dimensional voiceprint embedding code vector : ; Finally, the AAM-Softmax loss function is used to guide the training until the model training converges to obtain a seed model with better performance.

[0048] ; Among them, represents the true label, is the average loss value of the training samples; AAMLoss is the applied additive angular margin loss function.

[0049] (2) Testing stage: In the testing stage, first, L registered audio samples are input into the trained ECAPA-TDNN network to extract their high-dimensional feature vectors to represent the unique voiceprint information of the speaker. For multiple audio samples of the same speaker, the mean of their feature vectors is calculated to obtain the mean feature vector (speaker template) of the speaker. These speaker templates together form the voiceprint database in the testing stage, which is used to store the feature information of the registered speakers.

[0050] Subsequently, the audio sample to be recognized (Speaker A) is input into the same neural network, the corresponding high-dimensional feature vector is extracted, and the cosine similarity between it and each speaker template in the voiceprint database is calculated. The speaker corresponding to the template with the highest similarity is selected, and thus the speaker identification task is completed. The specific calculation process is as follows: ; Among them, represents the voiceprint embedding vector of the speech to be measured, represents the voiceprint embedding vector of a certain speaker in the voiceprint library, represents the calculated cosine similarity, with a range of [-1, 1]. The closer the value is to 1, the more similar the two are; represents the norm of the vector.

[0051] In summary, during the testing phase, by calculating the cosine similarity between the audio samples to be recognized and the registered speaker templates, the system can effectively match the most similar identity and achieve high-precision speaker recognition. This method uses ECAPA-TDNN to extract high-dimensional voiceprint features and constructs a voiceprint database through the mean feature vectors, thereby improving the stability and robustness of recognition. In addition, due to the use of cosine similarity measurement, the system can maintain strong discrimination ability in different noise environments, ensuring the accuracy of identity matching. This testing process lays a reliable technical foundation for the practical application of speaker identification, enabling the system to be applicable to complex land-air communication scenarios and achieve efficient and secure identity authentication.

Claims

1. A method for speaker recognition in land-air conversations with short speech and complex noise, characterized in that: The following steps are included: 1: Voice signal input; Collect voice signals from controllers and pilots in real time through aviation communication equipment; Step 2: Speech signal preprocessing; Noise suppression is performed on the speech signal, including using an adaptive filter to model a reference noise signal, dynamically adjusting the filter weights through a least mean square algorithm, and suppressing background noise; Step 3: Multi-scale feature extraction; Perform multi-scale feature extraction and fusion on the denoised speech signal, including short-time Fourier transform, filter bank energy feature extraction and multi-scale convolution feature extraction, and combine random time masking and frequency masking for data enhancement; Step 4: Train the improved ECAPA-TDNN model; The improved ECAPA-TDNN model is trained using real land-air call voice data. The improved ECAPA-TDNN model includes a SE-Res2Block module and a bidirectional gated recurrent unit for extracting local voiceprint features and capturing contextual timing information. The AAM-Softmax loss function is introduced during the training process to improve the recognition ability of the model. Step 5: Identity matching and determination; The audio sample to be recognized is input into the trained improved ECAPA-TDNN model, the corresponding high-dimensional feature vector is extracted, and the cosine similarity between the voiceprint embedding vector of the speech to be recognized and the voiceprint embedding vector of the registered speaker in the voiceprint database is calculated, and the identity corresponding to the highest similarity is selected as the recognition result.

2. According to claim 1, a method for speaker recognition in land-air conversations for short speech and complex noise, characterized in that: The noise suppression in step 2 specifically includes the following sub-steps: Step 2.1: Input signal modeling; Assume that the mixed speech signal collected in the aviation communication environment is: ; in, is the observed signal of noisy speech, For a pure speech signal, is the background noise; At the same time, the recorded land-air conversation voice is randomly intercepted as the reference noise signal ; Step 2.2: Adaptive filter modeling; Adaptive filter is used to filter the reference noise signal Perform linear filtering to generate a noise estimation signal : ; in, is the filter order, that is, the number of filter weights; is the weight coefficient of the adaptive filter, is the historical sample of the reference noise signal; Step 2.3: Error signal calculation; Calculate the error signal, which is the estimated clean speech: ; Step 2.4: Update weights using the least mean square algorithm; The weight update formula of the least mean square algorithm is as follows: ; in, is the step size factor, i.e., the learning rate; Step 2.5: Repeat steps 2.3 and 2.4 until the error signal Convergence, then the error signal As the output speech signal after noise suppression.

3. The method for speaker recognition in land-air conversations for short speech and complex noise according to claim 1, characterized in that: The multi-scale feature extraction in step 3 specifically includes: Step 3.1: Short-time Fourier transform; The reference noise signal after noise reduction After short-time Fourier transform, it is expressed as: ; in, represents the coefficients of the time-frequency matrix, is a window function, is the timeframe index, is the frequency index, is the window length of short-time Fourier transform; is the discrete time step index; j is an imaginary number; Step 3.2: Filter bank energy feature extraction, the calculation formula is as follows: ; in, Represents the filter index, indicating the Mel filters; For the The energy of the filter, is the response of the Mel filter; Respectively The lower and upper frequency limits of the Mel filter; Step 3.3: Multi-scale convolutional feature extraction F l ; Multi-scale convolution is used to model the features of different frequency bands. The calculation formula is as follows: ; in, is the input feature, For the The weights of the convolution kernels of the layers, is the bias term, is a nonlinear activation function, Represents different scales of multi-scale convolution; Step 3.4: Use random time masking and frequency masking for data enhancement. The formula is as follows: Time masking: ; in, is the starting timeframe index, For the shielding range; Frequency Masking: ; in, is the starting frequency index, For the shielding range; Step 3.4: Comprehensive feature representation; Combining filter bank energy features, multi-scale convolution features, short-time Fourier transform and data enhancement methods, the final feature representation is obtained. F : ; in, represents the short-time Fourier transform, represents the filter bank energy feature extraction, represents multi-scale convolution feature extraction, Represents data augmentation method.

4. The method for speaker recognition in land-air conversations with short speech and complex noise according to claim 1, characterized in that: In step 4, training the improved ECAPA-TDNN model specifically includes: Step 4.1: Perform a one-dimensional convolution operation on the input multi-scale speech feature X0, and use the ReLU activation function and batch normalization to transform the feature. The calculation process is as follows: ; in, is the 1D convolution kernel weight, is the bias term; the output signal , Represents the number of channels, is the time step; is batch normalization processing; ReLU is the activation function, and Conv1D is a one-dimensional convolution operation; Step 4.2: Perform deep feature extraction through the SE-Res2Block module; the SE-Res2Block module contains three layers of convolution units, and the convolution kernel size is set to , and the expansion rate The dilated convolution is used to enhance the receptive field, and the SE mechanism is combined to adaptively adjust the channel weights; The output of each SE-Res2Block module is calculated as follows: ; Among them, the dilated convolution adopts The expansion coefficient d increases with each layer. ; X i and X i+1 Respectively i SE-Res2Block module input and output; SE-Res2Block represents SE-Res2Block module processing; The outputs of three layers of SE-Res2Block are concatenated: ; Among them, Concat represents the concatenation operation; Step 4.3: Introduce a bidirectional gated recurrent unit to process the feature sequence. The bidirectional gated recurrent unit consists of a forward GRU and a backward GRU. The bidirectional gated recurrent unit is introduced for bidirectional sequence modeling, and the forward and backward hidden states are calculated separately: ; ; in, , and are the hidden states of the forward GRU and the backward GRU respectively; is the number of GRU hidden units; fwd For forward GRU processing, GRU bwd For backward GRU processing; The bidirectional gated recurrent unit concatenates the hidden states of the forward GRU and the backward GRU and outputs : ; in, Represents the feature dimension after concatenating the hidden states of the forward GRU and the backward GRU; T is the time step; Step 4.4: Output after convolution activation normalization layer: ; in, , and are the output, weight, and bias of the convolutional activation normalization layer, respectively; The attention pooling layer is used to calculate the global features and obtain the vector : ; in, ASP Calculated for the attention pooling layer; Then the vector that passes through the pooling layer Enter the fully connected normalized layer to output the voiceprint embedded code vector : ; Among them, FC is the fully connected layer processing; Step 4.5: Use the AAM-Softmax loss function to guide training until the model training converges to obtain a seed model that meets the performance requirements: ; in, represents the true label, is the average loss value of the training samples; AAMLoss is the application of the additive angular margin loss function.

5. The method for speaker recognition in land-air conversations for short speech and complex noise according to claim 1, characterized in that: The cosine similarity in step 5 is calculated as follows: ; in, represents the voiceprint embedding vector of the speech to be tested, Represents the voiceprint embedding vector of a speaker in the voiceprint database; represents the norm of a vector; Represents the calculated cosine similarity, ranging from [-1,1]. The closer the value is to 1, the more similar the two are.

Citation Information

Patent Citations

  • Identity feature extraction method and device based on acoustic feature generation and storage medium

    CN116631406A

  • Audio processing method and device, equipment, medium and product

    CN117153175A

  • Education scene-oriented voiceprint recognition model training method

    CN117854508A

  • Abnormal sound detection method for fault diagnosis of belt conveyor

    CN119360891A

  • Converter valve fault identification method based on voiceprint small sample learning

    CN119517089A

Cited By

  • Audio authentic identification method, device, equipment and medium

    CN120581038A

  • English pronunciation correction method using AI speech recognition

    CN120783751A