Air traffic control dialogue type speech recognition method and device based on multi-task learning
By building a multi-task learning model, combining speech feature representation and activity detection modules, and using a dynamic local window attention mechanism, the problem of segmentation error and insufficient context utilization in air-controlled speech recognition is solved, and efficient air-controlled dialogue recognition is achieved.
Patent Information
- Application Number
- CN202510456545.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-05
AI Technical Summary
The existing air-tube speech recognition technology has problems such as accumulation of segmentation errors and failure to effectively utilize the context in complex dialogue scenarios, resulting in insufficient recognition accuracy and robustness.
A multi-task learning model is built, combining the speech feature representation learning module and the speech activity detection module, and a dynamic local window attention mechanism is adopted to optimize speech segmentation and recognition through self-supervised pre-training and dual-branch joint training to achieve end-to-end collaborative optimization.
It significantly improves the accuracy and system efficiency of air-tube speech recognition, and can provide highly robust speech resolution capabilities in complex noise and variable air-tube dialogue scenarios.
Smart Images

Figure CN120431907A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence speech recognition technology, and in particular to an air traffic control conversational speech recognition method and device based on multi-task learning. Background Art
[0002] As the core interaction method for air traffic control, ground-to-air communication carries the real-time transmission of key instructions between controllers and pilots, and is a vital hub for ensuring aviation safety and efficiency. With the increasing demand for automated and intelligent air traffic control, speech recognition technology has gradually become the technical cornerstone for parsing unstructured communication data and building situational awareness systems. However, air traffic control speech is highly domain-specific, with its dialogue patterns characterized by short interactions, frequent role switching, and dense technical terminology. It is also often affected by factors such as environmental noise, accent differences, and distortion from communication equipment. This poses severe challenges to the accuracy and real-time performance of traditional general-purpose speech recognition algorithms.
[0003] Currently, air traffic control speech recognition technology primarily relies on general speech recognition frameworks and relies on large-scale annotated data to train end-to-end models. While these approaches demonstrate basic performance in common scenarios, they face significant limitations in complex conversational scenarios. First, air traffic control conversations contain a large number of specialized terms (such as flight call signs and navigation point names) and similarly pronounced control instructions. Traditional acoustic-language models, lacking domain adaptability, are prone to misrecognition. Second, the dynamic segmentation of speech and non-speech segments in continuous conversational streams (such as when the controller and pilot alternately speak) has not been effectively modeled. The cascaded architecture of independent speech activity detection and recognition modules can easily lead to segmentation error accumulation, reducing system robustness. Third, air traffic control conversations typically have strict temporal logic and contextual constraints, making it difficult for existing frameworks to effectively capture cross-turn semantic dependencies in long-term conversational scenarios.
[0004] To address the above issues, existing technologies attempt to optimize from the perspectives of multi-task learning and integration of prior knowledge. In terms of multi-task learning, existing research focuses on the joint training of acoustic models and language models, while auxiliary tasks such as voice activity detection and role separation are often handled independently, resulting in a lack of coordination between time series segmentation and speech recognition tasks, making it difficult for the model to learn globally consistent representations from the dialogue structure level. In terms of integrating prior knowledge, it mainly relies on static domain knowledge bases or expert rules, which lack flexibility in the face of the variability and complexity of air traffic control dialogues. With the emergence of new terminology and interaction patterns, static knowledge bases are difficult to update adaptively, limiting the performance of the model in practical applications.
[0005] In summary, existing air traffic control speech recognition technology has significant defects in segmentation error accumulation and failure to effectively utilize context. There is an urgent need for a multi-task learning framework that can simultaneously optimize speech segmentation detection and speech recognition and adaptively integrate dialogue state information to improve recognition accuracy and system efficiency in complex air traffic control dialogue scenarios. Summary of the Invention
[0006] To address the above-mentioned problems in the prior art, this application proposes an air traffic control conversational speech recognition method based on multi-task learning, comprising the following steps:
[0007] S1: Build and annotate an air traffic control conversational multi-task speech dataset.
[0008] S2: Construct a speech feature representation learning module, and input the continuous speech signal waveform in the air traffic control conversational multi-task speech dataset into the speech representation learning module for self-supervised pre-training.
[0009] S3: Construct a voice activity detection module and integrate the voice representation learning module to build a multi-task learning model to form an end-to-end trainable network that includes voice recognition and voice segmentation detection.
[0010] S4: Design a dual-branch joint training architecture, in which the main branch introduces a dynamic local window attention mechanism to fine-tune the parameters of the speech representation learning module, and the auxiliary branch is responsible for optimizing the speech activity detection module.
[0011] S5: Inputting the real-time collected air traffic control voice stream into the fully trained multi-task learning model, and synchronously outputting text information and voice activity detection identification.
[0012] As a preferred embodiment of the present invention, S1 includes:
[0013] S11: By collecting original recordings of air traffic control calls, a continuous voice signal waveform library is constructed, including two-way conversation scenarios between controllers and pilots;
[0014] S12: removing silent segments from the continuous speech signal waveform library using a preprocessing algorithm based on an energy threshold and a short-time zero-crossing rate;
[0015] S13: Mark the start and end time points of the voices of different characters in each valid speech segment, generate a structured temporal index, and transcribe the speech content into standardized text sentence by sentence;
[0016] S14: The processed data are stored in an associated manner to build a triple database (continuous speech signal waveform, speech time sequence index and corresponding text label) that can support end-to-end model training;
[0017] S15: Perform stratified sampling according to the air traffic control call scene type (takeoff, landing, approach, etc.), and divide the triple database into a training set, a validation set, and a test set to form an air traffic control conversational multi-task speech dataset.
[0018] As a preferred embodiment of the present invention, the speech representation learning module includes a feature extractor, a speech encoder and a quantizer;
[0019] The feature extractor FeatureExtractor(·) consists of multiple layers of one-dimensional convolutional modules, using Gaussian error linear units as activation functions between layers, supplemented by layer normalization. Through layered convolution operations, the original audio waveform is converted into high-dimensional temporal acoustic features, where shallow convolutions capture short-term spectral characteristics, while deep convolutions extract more discriminative speech unit boundary information. The expression is as follows:
[0020] H=FeatureExtractor(W)
[0021] Where H={h1,h2,…,h s} represents the high-dimensional temporal acoustic features extracted from the original speech waveform, W = {w1,w2,…,w t} is the original speech waveform;
[0022] The speech encoder SpeechEncoder(·) contains multiple layers of self-attention modules stacked on top of the Transformer architecture. This multi-head self-attention mechanism captures long-range dependencies in speech signals and generates richer contextual information. The expression is as follows:
[0023] C=SpeechEncoder(H)
[0024] Where C={c1,c2,…,c s} represents the context-sensitive speech representation extracted from high-dimensional temporal acoustic features;
[0025] The quantizer Quantization (·) uses a product quantization strategy and includes multiple learnable discrete codebooks. Vector quantization is used to map continuous speech features into discrete representations (pseudo-labels) that serve as supervisory signals for contrastive learning. The expression is as follows:
[0026] Q=Quantization(H)
[0027] Where Q = {q1,q2,…,q k} represents the discrete codebook generated from high-dimensional temporal acoustic features.
[0028] As a preferred embodiment of the present invention, the speech representation learning module pre-training step includes:
[0029] (1) performing random masking on the high-dimensional temporal acoustic features output by the feature extractor, and adopting a segmented continuous masking mode as the masking strategy;
[0030] (2) Construct a contrast prediction task, and at each masking position t, convert the speech encoder output c t As the query vector, calculate the candidate codeword q corresponding to the quantizer position t Similarity distribution L m The expression is as follows:
[0031]
[0032] in, represents the set of candidate quantized representations for position t, consisting of one positive sample and K negative samples. sim(·) refers to the cosine similarity function;
[0033] (3) Introducing the diversity regularization term L based on codebook entropy maximization div , by maximizing the selection probability entropy of the codewords in each codebook to avoid codebook collapse. The expression is as follows:
[0034]
[0035] Where M is the number of codebooks, p m,k Refers to the codeword selection probability, count(m,k) represents the probability that the kth codeword in the mth codebook is activated in the batch;
[0036] (4) The Adam optimizer is used to optimize the objective function L, and the trainable parameters of the feature extractor, speech encoder, and quantizer are updated synchronously through the back-propagation algorithm. The expression is as follows:
[0037] L=L m +μL div
[0038] Where μ is the regularization coefficient.
[0039] As a preferred embodiment of the present invention, the voice activity detection module includes a channel-temporal attention unit, a self-attention module based on a Transformer architecture stack, and an index mapping layer;
[0040] The channel-temporal attention unit comprises a parallel channel attention submodule and a temporal attention submodule. The channel attention submodule selectively enhances discriminative acoustic features, while the temporal attention submodule highlights the salient features of valid speech segments. The outputs of the two submodules achieve cross-dimensional feature interaction through tensor addition. The expression is as follows:
[0041] X CS =X C +X S
[0042] Among them, X C represents the channel attention feature map, X S Represents a temporal attention feature map.
[0043] The self-attention module SelfAttention(·) maintains sensitivity to changes in local acoustic events by establishing global contextual associations between frame-level features. The expression is as follows:
[0044] Z = SelfAttention(X CS )
[0045] The index mapping layer IndexProjection(·) consists of a fully connected network and a decision unit. This layer ultimately outputs the precise time-series index containing the start and end points of the speech segment. The expression is as follows:
[0046] Index=IndexProjection(Z)
[0047] As a preferred solution of the present invention, the collaborative logic of each module of the multi-task learning model is as follows:
[0048] (1) The pre-trained speech representation learning module achieves a paradigm shift from self-supervised learning to supervised learning by removing the quantizer, and at the same time adds a linear classification layer to construct a mapping channel from speech features to vocabulary space.
[0049] (2) The high-dimensional temporal acoustic features output by the feature extractor are routed using a three-branch architecture: the main path feature H main Input speech encoder, capture cross-round long-term dependencies (such as the temporal logic of control instructions) through the global self-attention mechanism, and establish the global semantic representation of the speech signal; auxiliary path feature H aux The access voice activity detection module VADModule(·) extracts the temporal segmentation information and provides a structured segmentation prior for the main path; the dynamic path feature H dyn Combined with the real-time segmentation index output by the voice activity detection module, dynamic local window attention calculation is performed in the speech encoder to focus on the fine-grained features of the effective speech segments and suppress cross-turn interference caused by role switching. The expression is as follows:
[0050] H=FeatureExtractor(W)
[0051] H main ,H aux ,H dyn=clone(H)
[0052] C main =SpeechEncoder(H main )
[0053] Index=VADModule(H aux )
[0054] C dyn =SpeechEncoder(H dyn ,Index)
[0055] Here, clone(·) represents a cloning operation.
[0056] (3) The encoding outputs of the main path and the dynamic path are fused through weighted summation and projected into the vocabulary space through the linear classification layer Classification(·), outputting the vocabulary probability matrix.
[0057] C=λC main +(1-λ)C dyn
[0058] Y=Classification(C)
[0059] Among them, λ∈(0,1) is an adjustable parameter.
[0060] As a preferred embodiment of the present invention, the dynamic local window attention calculation process includes:
[0061] (1) Dynamic attention window generation phase: Based on the real-time segmentation index windows output by the voice activity detection module, a time-limited attention window is constructed for each valid speech segment. The segmentation index set is defined as:
[0062] windows=[[s0,e0],[s1,e1],…,[s i ,e i ]]
[0063] where s i ,e i They represent the start frame index and end frame index of the i-th valid speech segment, respectively, and constitute the time domain attention constraint boundary.
[0064] (2) Window-constrained self-attention calculation stage: In the Transformer architecture of the speech encoder, feature extraction is achieved through the window-based self-attention mechanism. i =[s i ,e i ], and its self-attention calculation process is expressed as:
[0065]
[0066] Among them, Q i ,K i ,V i Respectively represent the i-th window window i The query matrix, key matrix and value matrix of , d is the feature dimension. The final attention output is obtained by concatenating features across windows.
[0067] Attention=concat(Attention0,Attention1,…,Attention i )
[0068] concat(·) represents the tensor concatenation operation along the time dimension to ensure the temporal consistency of the processing results of each window.
[0069] As a preferred embodiment of the present invention, the dual-branch joint training architecture includes:
[0070] (1) The main branch adopts a parameter inheritance and fine-tuning strategy, retaining all parameters of the pre-trained speech representation learning module while using the connection temporal classification loss function to align and guide the frame-level mapping learning of acoustic features to text symbols;
[0071] (2) The auxiliary branch constructs a voice activity detection module with random initialization, and learns boundary-sensitive features through a binary cross-entropy loss function, focusing on strengthening the ability to distinguish between speech and non-speech segments;
[0072] (3) The joint loss function L is optimized by Adam J Perform back propagation and gradient descent, the L J By connecting the temporal classification loss L CTC With binary cross entropy loss L BCE Perform weighted fusion. The expression is as follows:
[0073] L J =(1-Δ)L CTC +αL BCE
[0074] Where α∈(0,1) is the dynamic weight adjustment parameter.
[0075] As a preferred embodiment of the present invention, the joint training process includes a dynamic task weight adjustment mechanism:
[0076] (1) The loss function weight distribution is achieved through the learnable task importance coefficient α. In the initial stage, a higher α value is set to prioritize the voice activity detection task. As the training rounds increase, the α value is linearly attenuated to improve the learning intensity of the speech recognition task.
[0077] (2) Construct a loss surface curvature perception module to monitor the Hessian matrix eigenvalue distribution of the loss functions of the two tasks in real time. When a conflict in the optimization direction between tasks is detected, the gradient update amplitude in the conflicting direction is automatically reduced.
[0078] An electronic device comprises at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any of the methods described above.
[0079] Compared with the prior art, the present invention has the following beneficial effects:
[0080] This paper addresses existing issues such as segmentation error accumulation and ineffective use of context in existing technologies by proposing an end-to-end recognition framework with multi-task collaborative optimization. By constructing a joint learning model for acoustic semantic representation and speech activity detection, this framework achieves collaborative optimization of speech segmentation and content recognition. This significantly improves the performance and efficiency of air traffic control speech recognition in complex conversations and noise interference scenarios, providing highly robust speech parsing capabilities for air traffic control automation systems.
[0081] The present invention proposes a dynamic local window attention mechanism, which focuses on the context modeling of valid speech segments based on real-time segmentation indexes, thereby achieving fast and effective air traffic control conversational speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 This is a flow chart of an air traffic control conversational speech recognition method based on multi-task learning according to the present invention;
[0083] Figure 2 This is a diagram showing the pre-training structure of the speech representation learning module in the multi-task learning-based air traffic control conversational speech recognition method described in the present invention;
[0084] Figure 3 This is a diagram of the fine-tuning structure of the multi-task learning model in the multi-task learning-based air traffic control conversational speech recognition method described in the present invention;
[0085] Figure 4 This is the inference logic diagram of the multi-task learning model in the air traffic control conversational speech recognition method based on multi-task learning described in the present invention. DETAILED DESCRIPTION
[0086] The present invention will be further described below with reference to the accompanying drawings.
[0087] The present invention provides
[0088] In one embodiment, this example is a specific application example of the method described in the present invention, wherein the research data in this example is derived from the real-time recording database of the actual air traffic control system. After a standardized preprocessing process (including data cleaning, quality inspection and time series alignment), a benchmark data set containing 210,000 voice samples was constructed, and at the same time, it was used for pre-training of the voice representation learning module according to the 9:1 division rule. In addition, one-tenth of the data was labeled, and the multi-task learning model was fine-tuned according to the 8:1:1 division rule. The detailed information of the air traffic control conversational multi-task voice data set is shown in the following table:
[0089] Table 1 Dataset information
[0090]
[0091] The fine-tuning process of this example uses a mixed character set as the basic recognition unit: the total character set contains 712 elements, covering typical pronunciation characteristics in the civil aviation field. Specifically, it includes:
[0092] (1) 681 high-frequency Chinese characters, basically covering the core Chinese characters of civil aviation terminology;
[0093] (2) 26 English letters to meet the recognition needs of special scenarios such as aircraft call signs and navigation points;
[0094] (3) 5 special function symbols, including the silent segment identifier <blank>, word separator <space>, statement starter <sos>, statement terminator <eos>and unregistered word identifiers <unk>.
[0095] S1: Build and annotate an air traffic control conversational multi-task speech dataset.
[0096] S2: Construct a speech feature representation learning module, and input the continuous speech signal waveform in the air traffic control multi-dialogue multi-task speech dataset into the speech representation learning module for self-supervised pre-training.
[0097] The specific parameters of the speech feature representation learning module in this embodiment are set as follows:
[0098] Feature Extractor: The feature extractor in this embodiment uses a cascaded convolutional architecture, consisting of seven sequential one-dimensional convolutional layers. The first layer has a convolution kernel with 512 channels, a kernel size of 10, and a stride of 5. The following four layers maintain the same number of channels, with a kernel size of 3 and a stride of 2. The final two layers maintain the same number of channels, with a kernel size of 2 and a stride of 2. Batch normalization is performed after each convolutional layer.
[0099] Quantizer: The quantizer in this example uses a dual-codebook discrete representation scheme, consisting of two independent codebooks, each containing 320 learnable discrete feature vectors.
[0100] Speech Encoder: The speech encoder in this example consists of 12 layers of self-attention modules stacked based on the Transformer architecture. The dimension of the hidden layer is 768, the number of self-attention heads is 8, and the feedforward network dimension is expanded to 3072 dimensions.
[0101] S3: Construct a voice activity detection module and integrate the voice representation learning module to build a multi-task learning model to form an end-to-end trainable network including acoustic modeling and voice segmentation detection.
[0102] The specific parameters of the voice activity detection module in this embodiment are set as follows:
[0103] Channel-Temporal Attention Unit: The channel-temporal attention unit in this example uses a dual-path attention mechanism. The channel attention submodule consists of two one-dimensional convolutional layers, each activated by a ReLU function and a Sigmoid function, respectively; the temporal attention submodule consists of one one-dimensional convolutional layer activated by a Sigmoid function.
[0104] Self-attention module: The self-attention module in this example consists of a two-layer streamlined Transformer structure, where the dimension of the hidden layer is 256, the number of self-attention heads is 1, and the feedforward network dimension is expanded to 1024.
[0105] Index mapping layer: The index mapping layer in this example uses a linear transformation structure to map 256-dimensional features into a single-channel output and generates frame-level speech activity probabilities through Sigmoid activation.
[0106] S4: Design a dual-branch joint training architecture, in which the main branch introduces a dynamic local window attention mechanism to fine-tune the parameters of the speech representation learning module, and the auxiliary branch is responsible for optimizing the speech activity detection module.
[0107] S5: Inputting the real-time collected air traffic control voice stream into the fully trained multi-task learning model, and synchronously outputting text information and voice activity detection identification.
[0108] To verify the effectiveness of the proposed method, this example conducted two comparative experiments, Groups A and B. Group A compared speech recognition performance, including mainstream speech recognition models such as Deep Speech 2, Speech Transformer, Conformer CTC, Conformer Transducer, Whisper, and Wav2Vec2. Group B compared voice activity detection, using representative detection algorithms such as WebRTC VAD, rVAD, Self-Attentive VAD, and CLDNN-VAD as benchmarks.
[0109] Table 2 Speech recognition model effect comparison
[0110] Model Word error rate (%) Speech Transformer 14.49 Deep Speech 2 5.50 Conformer CTC 4.21 Conformer Transducer 3.67 Whisper 2.67 Wav2Vec2 2.75 The method of this embodiment 2.06
[0111] Table 3 Comparison of voice activity detection model effects
[0112] Model Accuracy (%) Recall rate (%) F1 score (%) WebRTC VAD 86.12 80.63 83.28 rVAD 73.06 64.86 68.72 Self-Attentive VAD 98.35 99.02 98.69 CLDNN-VAD 90.92 90.99 90.96 The method of this embodiment 99.87 99.91 99.89
[0113] Although the present invention is described herein with reference to specific embodiments, it should be understood that these embodiments are merely illustrative of the principles and applications of the invention. It should be understood that many modifications may be made to the illustrative embodiments, and that other arrangements may be devised, without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that the various dependent claims and features described herein may be combined in ways other than those described in the original claims. It should also be understood that features described in conjunction with individual embodiments may be employed in conjunction with other described embodiments.< / unk> < / eos> < / sos> < / space> < / blank>
Claims
1. A multi-task learning-based air traffic control conversational speech recognition method, characterized in that: The following steps are involved: S1: Build and annotate an ATC conversational multi-task speech dataset; S2: Constructing a speech feature representation learning module and inputting the continuous speech signal waveform in the ATC conversational multi-task speech dataset into the speech representation learning module for self-supervised pre-training; S3: Construct a voice activity detection module and integrate it with the voice representation learning module to build a multi-task learning model, forming an end-to-end trainable network that includes speech recognition and speech segmentation detection; S4: Design a two-branch joint training architecture, where the main branch introduces a dynamic local window attention mechanism to fine-tune the parameters of the speech representation learning module, and the auxiliary branch is responsible for optimizing the voice activity detection module; S5: Input the real-time collected air traffic control voice stream into the fully trained multi-task learning model, and simultaneously output text information and voice activity detection identification.
2. The multi-task learning-based air traffic control conversational speech recognition method according to claim 1, characterized in that: Said S1 comprises: S11: By collecting original recordings of air traffic control calls, a continuous voice signal waveform library is constructed, including two-way conversation scenarios between controllers and pilots; S12: A preprocessing algorithm based on energy threshold and short-time zero-crossing rate is used to remove silent segments from the continuous speech signal waveform library; S13: Mark the start and end time points of the voices of different characters in each valid speech segment, generate a structured temporal index, and transcribe the speech content into standardized text sentence by sentence; S14: The processed data are stored in an associated manner to build a triple database that can support end-to-end model training; S15: Perform stratified sampling according to the air traffic control call scenario type, and divide the triple database into a training set, a validation set, and a test set to form an air traffic control conversational multi-task speech dataset.
3. The multi-task learning-based air traffic control conversational speech recognition method according to claim 1, characterized in that: The speech representation learning module includes a feature extractor, a speech encoder and a quantizer: the feature extractor contains a multi-layer one-dimensional convolution module, Gaussian error linear unit is used as the activation function between layers, and is supplemented by layer normalization; through layered convolution operations, the original audio waveform is converted into high-dimensional time series acoustic features, where shallow convolution captures short-time spectrum characteristics and deep convolution abstracts more discriminative speech unit boundary information.
4. The multi-task learning-based air traffic control conversational speech recognition method according to claim 3, characterized in that: The speech representation learning module pre-training step includes: (41) performing random masking on the high-dimensional temporal acoustic features output by the feature extractor, wherein the masking strategy adopts a segmented continuous masking mode; (42) Construct a contrast prediction task, and at each masking position t, convert the speech encoder output c t As the query vector, calculate the candidate codeword q corresponding to the quantizer position t Similarity distribution L m ; (43) Introducing the diversity regularization term L based on codebook entropy maximization div , codebook collapse is avoided by maximizing the selection probability entropy of codewords in each codebook.
5. The multi-task learning-based air traffic control conversational speech recognition method according to claim 1, characterized in that: The voice activity detection module includes a channel-temporal attention unit, a self-attention module based on the Transformer architecture stack, and an index mapping layer; The channel-temporal attention unit includes a parallel channel attention submodule and a temporal attention submodule; wherein, the channel attention submodule realizes the selective enhancement of discriminative acoustic features; the temporal attention submodule highlights the salient features of valid speech segments; the outputs of the two submodules realize cross-dimensional feature interaction through tensor addition.
6. The multi-task learning-based air traffic control conversational speech recognition method according to claim 1, characterized in that: The collaborative logic of each module of the multi-task learning model is as follows: (61) The pre-trained speech representation learning module achieves a paradigm shift from self-supervised learning to supervised learning by removing the quantizer and adding a linear classification layer to construct a mapping channel from speech features to vocabulary space; (62) The high-dimensional temporal acoustic features output by the feature extractor are routed using a three-branch architecture: the main path features are input into the speech encoder, and the global self-attention mechanism is used to capture the long-term dependencies across turns and establish a global semantic representation of the speech signal; the auxiliary path features are connected to the voice activity detection module to extract temporal segmentation information and provide a structured segmentation prior for the main path; the dynamic path features are combined with the real-time segmentation index output by the voice activity detection module to perform dynamic local window attention calculation in the speech encoder, focusing on the fine-grained features of the effective speech segments and suppressing the cross-turn interference caused by role switching.
7. The multi-task learning-based air traffic control conversational speech recognition method according to claim 6, characterized in that: The dynamic local window attention calculation process includes: (71) Dynamic attention window generation stage: Based on the real-time segmentation index windows output by the voice activity detection module, a time-limited attention window is constructed for each valid speech segment; (72) Window-constrained self-attention calculation stage: In the Transformer architecture of the speech encoder, feature extraction is achieved through the windowed self-attention mechanism.
8. The multi-task learning-based air traffic control conversational speech recognition method according to claim 1, characterized in that: The dual-branch joint training architecture includes: (81) The main branch adopts a parameter inheritance and fine-tuning strategy, which retains all parameters of the pre-trained speech representation learning module and uses the connection temporal classification loss function to align and guide the frame-level mapping learning of acoustic features to text symbols; (82) The auxiliary branch constructs a voice activity detection module with random initialization, and learns boundary-sensitive features through a binary cross-entropy loss function, focusing on strengthening the ability to distinguish between speech and non-speech segments; (83) The joint loss function L is optimized by Adam J Perform back propagation and gradient descent, the L J By connecting the temporal classification loss L CTC With binary cross entropy loss L BCE Perform weighted fusion.
9. The multi-task learning-based air traffic control conversational speech recognition method according to claim 8, characterized in that: The joint training process includes a dynamic task weight adjustment mechanism: (91) The loss function weight distribution is achieved through the learnable task importance coefficient α. In the initial stage, a higher α value is set to prioritize the voice activity detection task. As the training rounds increase, the α value is linearly attenuated to improve the learning intensity of the speech recognition task. (92) Construct a loss surface curvature perception module to monitor the Hessian matrix eigenvalue distribution of the loss functions of the two tasks in real time. When a conflict in the optimization direction between tasks is detected, the gradient update amplitude in the conflicting direction is automatically reduced.
10. An electronic device, characterized in that: The invention comprises at least one processor and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1 to 9.
Citation Information
Cited By
Speech recognition method and device based on artificial intelligence, and medium
CN121999782A