Aviation control speech recognition method and system

By combining a dual-stream generative adversarial network and a multimodal feature fusion encoder, the problem of voice command recognition accuracy in complex flight environments was solved, accurate understanding and safe control of aviation control instructions were achieved, and the safety and efficiency of air transportation were improved.

CN120510841BActive Publication Date: 2025-09-12NAVAL AVIATION UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510998607.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-09-12
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

Existing aviation control speech recognition methods are easily affected by background noise in complex and changing flight environments, resulting in reduced clarity and recognizability of voice commands. In particular, commands containing complex semantics and numerical information cannot be accurately understood, posing a safety hazard.

Method used

A two-stream generative adversarial network is used to enhance speech signals. The multimodal feature fusion encoder is combined to extract and fuse features. The parameter range prediction head and scene adaptation classification head are used for prediction. The decoder is used to generate instruction text sequences to implement constraints on instruction parameters and action types.

Benefits of technology

It improves the accuracy and stability of voice recognition, ensures that the parameters and action types of instructions comply with aviation control rules, reduces the generation of illegal instructions, and improves the safety and operational efficiency of air transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120510841B_ABST
    Figure CN120510841B_ABST
Patent Text Reader

Abstract

This application discloses an aviation control speech recognition method and system, relating to the field of speech recognition. The method comprises acquiring air traffic controller speech data; classifying and enhancing the air traffic controller speech data using a two-stream generative adversarial network to obtain an enhanced speech signal; extracting and fusing features from the enhanced speech signal using a multimodal feature fusion encoder to obtain a context embedding vector and fused features; predicting the context embedding vector using a parameter range prediction head and a scene adaptation classification head to obtain a predicted parameter legal value range and permitted instruction action types; and generating an instruction text sequence using a decoder based on constraints imposed by the fused features based on the predicted parameter legal value range and permitted instruction action types. This application can improve recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition, and in particular to an aviation control speech recognition method and system. Background Art

[0002] In the training system for air traffic controller trainees, scenario-based training using highly realistic flight simulators has become a core means of enhancing trainees' professional capabilities. Trainees are required to control the flight status of a virtual aircraft in real time using voice commands within a simulated cockpit environment. The standardization and accuracy of these commands directly determine the effectiveness of the training.

[0003] With the development of virtual reality (VR) and artificial intelligence technologies, modern flight simulators can replicate the multi-dimensional interference factors found in complex control scenarios, such as airport tower background noise, communication link distortion, high-density flight traffic, and sudden emergency events. These factors place higher demands on the compliance of trainees' instructions. However, current simulator training primarily relies on manual evaluation, which is inefficient and highly subjective. There is an urgent need to establish automated methods for voice command recognition and compliance assessment.

[0004] While several research methods have been proposed for automated voice command recognition, their accuracy remains low due to several factors. Firstly, the complex and ever-changing flight environment makes voice communications susceptible to interference from various background noises, such as the roar of engines near airports and the sound of wind and rain in inclement weather. This severely impacts the clarity and legibility of voice commands. Secondly, air traffic control instructions not only include fixed-format commands like "OK to land" and "OK to take off," but also include numerous commands with specific numerical changes, such as "Adjust altitude to 10,000 feet" and "Descend 8,000 feet." Accurate recognition of these commands is crucial for safe aircraft operation. However, existing recognition methods often fail to accurately understand the semantic integrity of air traffic control instructions containing complex semantics and numerical information, which can easily lead to incorrect command execution and potentially pose safety risks. Summary of the Invention

[0005] The purpose of this application is to provide an aviation control speech recognition method and system to improve the recognition accuracy.

[0006] To achieve the above objectives, this application provides the following solutions:

[0007] In a first aspect, the present application provides an aviation control speech recognition method, comprising:

[0008] Obtain air traffic controller voice data;

[0009] Classifying and enhancing the air traffic controller voice data using a dual-stream generative adversarial network to obtain an enhanced voice signal;

[0010] Extracting and fusing features using a multimodal feature fusion encoder based on the enhanced speech signal to obtain a context embedding vector and a fusion feature;

[0011] Performing prediction using a parameter range prediction head and a scene adaptation classification head based on the context embedding vector to obtain a legal value range of the prediction parameter and an allowed instruction action type;

[0012] Constraints are performed based on the fusion features, the legal value range of the prediction parameters and the allowed instruction action types, and an instruction text sequence is generated using a decoder.

[0013] In one embodiment, the air traffic controller voice data is classified and enhanced using a dual-stream generative adversarial network to obtain an enhanced voice signal, specifically including:

[0014] Based on the air traffic controller voice data, the dual-path convolution structure of the two-stream adversarial generative network is used to perform depth convolution and point-by-point convolution to obtain the local features and point-by-point convolution results of each channel;

[0015] The local features of each channel and the point-by-point convolution results are concatenated and convolved using a dual-path convolution structure to obtain an enhanced speech signal.

[0016] In one embodiment, feature extraction and fusion are performed using a multimodal feature fusion encoder based on the enhanced speech signal to obtain a context embedding vector and fusion features, specifically including:

[0017] Extracting speech time-frequency features using a multi-layer convolutional Transformer model of a multimodal feature fusion encoder according to the enhanced speech signal to obtain an acoustic embedding vector;

[0018] Based on the flight status, the LSTM model of the multimodal feature fusion encoder is used to encode the temporal changes of the flight status and obtain the context embedding vector;

[0019] The acoustic embedding vector and the context embedding vector are weightedly fused using the gating weights of the multimodal feature fusion encoder to obtain a fusion feature.

[0020] In one embodiment, prediction is performed using a parameter range prediction head and a scene adaptation classification head based on the context embedding vector to obtain a legal value range of the predicted parameter and an allowed instruction action type, specifically including:

[0021] According to the context embedding vector, a parameter range prediction head is used to predict and obtain a legal value range of the prediction parameter; the parameter range prediction head is an MLP model including 5 fully connected layers;

[0022] The scene adaptation classification head is used to perform prediction based on the context embedding vector to obtain the allowed instruction action type; the scene adaptation classification head includes a 5-layer MLP model and a Softmax function connected to the MLP model.

[0023] In one embodiment, the decoder includes a feature extraction network, an instruction prediction branch and a numerical prediction branch; the feature extraction network is connected to the instruction prediction branch and the numerical prediction branch respectively; the instruction prediction branch and the numerical prediction branch both include three fully connected layers; the instruction text sequence includes fixed instructions and numerical instructions.

[0024] In one embodiment, the total loss function of the decoder is a weighted sum of cross entropy loss and numerical loss.

[0025] In a second aspect, the present application provides an aviation control speech recognition system, comprising:

[0026] Acquisition module, used to obtain air traffic controller voice data;

[0027] A classification and enhancement module, configured to classify and enhance the air traffic controller voice data using a dual-stream adversarial generative network to obtain an enhanced voice signal;

[0028] A feature extraction and fusion module is used to extract and fuse features based on the enhanced speech signal using a multimodal feature fusion encoder to obtain a context embedding vector and a fusion feature;

[0029] A prediction module, configured to perform prediction based on the context embedding vector using a parameter range prediction head and a scene adaptation classification head to obtain a legal value range of the prediction parameter and an allowed instruction action type;

[0030] A generation module is used to constrain the prediction parameter legal value range and the allowed instruction action type according to the fusion feature, and use a decoder to generate an instruction text sequence.

[0031] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aviation control speech recognition method.

[0032] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the aviation control speech recognition method when executed by a processor.

[0033] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the aviation control speech recognition method when executed by a processor.

[0034] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0035] The present application provides an aviation control speech recognition method and system, which uses a dual-stream adversarial generative network to perform classification and enhancement based on the aviation controller's speech data to obtain an enhanced speech signal; uses a multimodal feature fusion encoder to perform feature extraction and fusion based on the enhanced speech signal to obtain a context embedding vector and a fusion feature; uses a parameter range prediction head and a scene adaptation classification head to perform prediction based on the context embedding vector to obtain the legal value range of the predicted parameters and the allowed instruction action type; uses a decoder to generate an instruction text sequence based on the legal value range of the predicted parameters and the allowed instruction action type according to the fusion feature. The parameter range prediction head and the scene adaptation classification head can limit the parameters and action type of the instruction. When generating the instruction text sequence, the legal value range of the predicted parameters and the allowed instruction action type are used for constraint, thereby reducing the generation of illegal instructions at the source, reducing the deviation of the instructions, and improving stability, reliability and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0037] Figure 1 This is an application environment diagram for the aviation control speech recognition method.

[0038] Figure 2 The flowchart of the aviation control speech recognition method is shown in FIG.

[0039] Figure 3 This is a schematic diagram of the functional modules of the aviation control speech recognition system.

[0040] Figure 4 A schematic diagram of the structure of a computer device provided in one embodiment of the present application.

[0041] Figure 5 This is the network structure diagram of the algorithm. DETAILED DESCRIPTION

[0042] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0043] This application introduces a hybrid recognition framework enhanced by multimodal adversarial analysis, integrating a dynamic environment perception module and a command structure constraint mechanism into the traditional speech recognition process, focusing on addressing the two core challenges of noise robustness and command semantic integrity. By conducting a comprehensive and in-depth analysis of the commander's input voice, accurate voice signal recognition is achieved, and the recognition results are used to precisely control the flight status of the virtual aircraft, providing more reliable and efficient technical support for air traffic control and improving the safety and operational efficiency of air transportation.

[0044] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0045] The aviation control speech recognition method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the air traffic controller voice data to be processed to the server 104. After the server 104 receives the air traffic controller voice data to be processed, the server 104 classifies and enhances the air traffic controller voice data to be processed based on the air traffic controller voice data using a dual-stream adversarial generation network to obtain an enhanced voice signal; based on the enhanced voice signal, a multimodal feature fusion encoder is used to extract and fuse features to obtain a context embedding vector and a fusion feature; based on the context embedding vector, a parameter range prediction head and a scene adaptation classification head are used to predict to obtain the legal value range of the prediction parameter and the allowed instruction action type; based on the fusion feature, the legal value range of the prediction parameter and the allowed instruction action type are constrained, and a decoder is used to generate an instruction text sequence. The server 104 can feed back the obtained instruction text sequence to the terminal 102. In addition, in some embodiments, the air traffic control speech recognition method can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly perform air traffic control speech recognition on the air traffic controller speech data to be processed, or the server 104 can obtain the air traffic controller speech data to be processed from the data storage system and perform air traffic control speech recognition on the air traffic controller speech data to be processed.

[0046] Terminal 102 may include, but is not limited to, various desktop computers, laptops, smartphones, tablet computers, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. Server 104 may be implemented as a standalone server or a server cluster consisting of multiple servers, or may be a cloud server.

[0047] In an exemplary embodiment, Figure 2 As shown, a method for recognizing aviation control speech is provided. The method is executed by a computer device, specifically a computer device such as a terminal or a server, or a terminal and a server. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used for explanation, including the following steps 201 to 205.

[0048] Step 201: Acquire air traffic controller voice data.

[0049] Step 202: Classify and enhance the air traffic controller voice data using a dual-stream adversarial generative network to obtain an enhanced voice signal.

[0050] Step 203: performing feature extraction and fusion using a multimodal feature fusion encoder based on the enhanced speech signal to obtain a context embedding vector and fusion features.

[0051] Step 204: perform prediction using the parameter range prediction head and the scene adaptation classification head according to the context embedding vector to obtain the legal value range of the prediction parameter and the allowed instruction action type.

[0052] Step 205: Constraints are performed based on the fusion features, the legal value range of the prediction parameters, and the allowed instruction action types, and a decoder is used to generate an instruction text sequence.

[0053] The parameter range prediction header and the scenario adaptation classification header can limit the parameters and action types of instructions. When generating the instruction text sequence, the legal value range of the predicted parameters and the allowed instruction action types are used for constraints, thereby reducing the generation of illegal instructions at the source, reducing the deviation of instructions, and improving stability, reliability and accuracy.

[0054] In an exemplary embodiment, the voice data of air traffic controllers is collected and the instructions therein are labeled. Instructions can be divided into two parts, fixed instructions and numerical instructions. For example, fixed format instructions such as "OK to land" and "OK to take off" and numerical instructions such as "adjust altitude to 10,000 feet" and "descend altitude 8,000 feet". Construct a fixed instruction set as a classification type, assuming there are N categories. At the same time, record the numerical instruction V corresponding to each type of instruction, as well as the range R of the value min , R max . Thus, a classification and regression problem is constructed.

[0055] In an exemplary embodiment, a dual-stream adversarial generative network is used to classify and enhance the air traffic controller voice data to obtain an enhanced voice signal, specifically comprising: performing depth convolution and point-by-point convolution on the air traffic controller voice data using a dual-path convolution structure of a dual-stream adversarial generative network to obtain local features and point-by-point convolution results of each channel; and splicing and convolving the local features and point-by-point convolution results of each channel using a dual-path convolution structure to obtain an enhanced voice signal.

[0056] In air traffic control, the background noise in the environment where voice is collected is high and the communication quality is poor. The original voice signal may be severely interfered with by the noise, affecting subsequent voice recognition. To this end, this application designs a Two-Stream Generative Adversarial Network (TS-AGN) to enable the generator to generate enhanced speech that is closer to clean speech through adversarial training. At the same time, a discriminator is used to classify noise types to assist in training, improving the enhancement effect in specific noisy scenarios. Specifically, it includes the following two steps:

[0057] First, a dual-path convolution structure is used to process wide-band and narrow-band features in parallel to better preserve key speech components. The input is the original speech signal x with noise. n , the output is the enhanced speech signal x c The original noisy voice signal collected is the air traffic controller voice data. The specific formula is:

[0058]

[0059] in, It is a depthwise convolution operation that performs convolution on each input channel independently to extract local features of each channel; is a point-by-point convolution operation, i.e. Convolution is used to combine features extracted by depthwise convolution and adjust the number of channels. It is a splicing operation that splices the results of depthwise convolution and pointwise convolution. Finally, The spliced ​​features are further processed to output enhanced speech features. For input feature map or speech signal The output after the convolution operation. A discriminator is built to distinguish between clean speech and generated enhanced speech, improving speech enhancement. After training the GAN, only the generator is retained, not the discriminator.

[0060] Specifically, the discriminator is used to assist in training the noise type, and its loss function The formula is:

[0061]

[0062] Among them, the discriminator consists of 4 G (x) and 1 fully connected layer containing 64 nodes. Finally, the Sigmoid output is used to judge whether the input signal is a clean speech signal. Represents the log-likelihood expectation of clean speech passing through the discriminator, so that the discriminator can correctly identify clean speech; It represents the inverse of the log-likelihood expectation of the enhanced speech generated by the generator with noisy speech passing through the discriminator, which is used to train the generator to generate enhanced speech that is closer to clean speech.

[0063] Through these two tasks, the input original speech signal x can be n Extract enhanced speech signal x c The above model is trained 100 times to obtain the trained generator model. Then, the output of the generator is used as the final extracted enhanced speech signal x c .

[0064] In an exemplary embodiment, a multimodal feature fusion encoder is used to perform feature extraction and fusion based on the enhanced speech signal to obtain a context embedding vector and a fusion feature, specifically including: extracting speech time-frequency features based on the enhanced speech signal using a multi-layer convolutional Transformer model of a multimodal feature fusion encoder to obtain an acoustic embedding vector; encoding the temporal changes of the flight status based on the flight status using an LSTM model of a multimodal feature fusion encoder to obtain a context embedding vector; and performing weighted fusion based on the acoustic embedding vector and the context embedding vector using the gating weights of the multimodal feature fusion encoder to obtain a fusion feature.

[0065] like Figure 5 As shown in the figure, a multimodal feature fusion encoder is designed to enhance the speech signal x c Perform feature extraction and fusion to obtain fusion features containing speech time-frequency features and flight scene context information, so that the model can better understand the environment in which the voice command is located, improve the ability to recognize complex aviation control commands, and more accurately judge the intention of the command. , whose output is the fused feature Specifically, we design a dual-branch Transformer structure:

[0066] ① The acoustic branch uses multi-layer convolutional Transformer to enhance the speech signal Processing is performed to extract the time-frequency features of speech. Formula:

[0067]

[0068] in, Represents the multi-layer convolutional Transformer model of the acoustic branch, which transforms the input enhanced speech signal Convert to acoustic embedding vector , which contains the acoustic feature information of the speech.

[0069] ② The context branch realizes spatiotemporal fusion and uses LSTM to encode the temporal changes of flight status (flight phase, altitude, speed, weather conditions at the locations passed), captures dynamic information during the flight, and outputs a context embedding vector ,formula:

[0070]

[0071] This formula represents the input of contextual information such as flight phase, airspace rules, and weather conditions into LSTM for time modeling to obtain the context embedding vector , which contains the contextual feature information of the flight scene, where stage is the flight stage, v is the speed, and s is the passed position. For weather conditions.

[0072] Flight stages are discrete categories, encompassing five fixed stages: takeoff preparation, takeoff, cruise, landing, and docking. These are represented by a one-hot encoding of length 5. Speed ​​and position are first normalized to [0, 1] and then represented by encodings of length 3, representing speed and position in the x, y, and z directions, respectively. Weather conditions are categorized into six categories: sunny, cloudy, precipitation, thunderstorm, strong wind, and fog, and are represented by a one-hot encoding of length 6.

[0073] ③ Dynamic gated feature fusion:

[0074] First calculate the gating weights:

[0075]

[0076] in is the Sigmoid function, The result is mapped to interval, is a learnable parameter, Represents the acoustic embedding vector and context embedding vector Concatenate by dimension. Gating weight Used to control the contribution of acoustic features and context features in the fusion process.

[0077] Then by the formula Perform feature fusion, is element-by-element multiplication. The formula is based on the gate weight Acoustic embedding vector and context embedding vector Perform weighted combination and finally obtain fusion features .

[0078] In an exemplary embodiment, prediction is performed using a parameter range prediction head and a scene adaptation classification head based on the context embedding vector to obtain a legal value range of the predicted parameters and an allowed instruction action type, specifically including: prediction is performed using a parameter range prediction head based on the context embedding vector to obtain a legal value range of the predicted parameters; the parameter range prediction head is an MLP model including 5 fully connected layers; prediction is performed using the scene adaptation classification head based on the context embedding vector to obtain an allowed instruction action type; the scene adaptation classification head includes a 5-layer MLP model and a Softmax function connected to the MLP model.

[0079] During training, the parameter range prediction head and scene adaptation classification head receive the context embedding vector The parameter range prediction head outputs the legal value range through a 5-layer fully connected layer MLP, and the scene adaptation classification head outputs the probability distribution of the command action type. P These outputs are used in the calculation of the loss function. Together with the model's predicted instruction text and other results, they are compared with the true label. The error is calculated and backpropagated, adjusting the model parameters. This allows the model to learn instruction patterns and parameter ranges that comply with aviation control regulations, improving the model's ability to identify and generate legal instructions.

[0080] When the model is trained and put into practical use, the input data related to the aviation control voice command will also pass through the parameter range prediction head and the scene adaptation classification head. The parameter range prediction head predicts the legal value range of the parameters in the current scenario, which can help determine whether the parameters in the command are within a reasonable range. For example, when the command involves adjusting the aircraft's flight altitude, the prediction head can determine whether the given altitude value is within the legal altitude range in the current airspace, flight phase, and other scenarios. The scene adaptation classification head outputs the probability distribution of the command action type, which is used to determine what type of action the command belongs to (such as take-off, landing, cruising, etc.), assisting the model to accurately understand the intent of the command, filter out illegal or unreasonable commands, ensure that the generated command text complies with aviation control rules, and improve the safety and accuracy of speech recognition and command execution.

[0081] like Figure 5 As shown in Figure 2, we design a parameter range prediction head and a scene adaptation classification head to implement implicit rule constraint learning. Its input is the following embedding vector , output the legal value range of the prediction parameters and the probability distribution of instruction action types .

[0082] ① Parameter range prediction head: Use 5 layers of fully connected layers to form MLP, the last layer outputs two nodes, and embeds the context vector Process and predict the legal value range of the parameters in the current scenario The formula is: .

[0083] Represents embedding the context into a vector Input to the MLP layer, after linear transformation and activation function processing, the legal value range of the output parameter ( Represents the lowest, represents the highest value).

[0084] ②Scene adaptation classification head: embedding the context vector through the classification layer Process and predict the currently allowed instruction action type. The formula is:

[0085]

[0086] After passing through 5 layers of MLP and then processed by the Softmax function, the probability distribution of each instruction action type is obtained , indicating the possibility of executing each command action in the current scenario.

[0087] In an exemplary embodiment, the decoder includes a feature extraction network, an instruction prediction branch and a numerical prediction branch; the feature extraction network is connected to the instruction prediction branch and the numerical prediction branch respectively; the instruction prediction branch and the numerical prediction branch each include three fully connected layers; the instruction text sequence includes fixed instructions and numerical instructions.

[0088] Aviation control commands have strict parameter ranges and command type restrictions. Violating these rules can lead to serious safety issues. By predicting the legal parameter value range and permitted command action types, and applying implicit constraints, we can reduce the generation of illegal commands at the source and improve the accuracy and safety of command recognition.

[0089] To this end, a decoder is designed based on Transformer, using fusion features Generates instruction text sequences and constrains the generation process through parameter range prediction head and scene adaptation classification head to ensure that the generated instructions comply with the rules and restrictions of the current flight scene. Its input is the fusion feature , the output is a sequence of instruction text.

[0090] The designed decoder is a 5-layer Transformer as a feature extraction network, which is then connected to the instruction prediction branch and the numerical prediction branch. Both branches are composed of 3 fully connected layers. The instruction prediction branch has N nodes, representing the N types collected in step 201, and the numerical prediction branch has 1 node for outputting the predicted numerical instructions.

[0091] In an exemplary embodiment, the total loss function of the decoder is a weighted sum of cross entropy loss and numerical loss.

[0092] For the N nodes output by the instruction prediction branch, the N-classification cross entropy loss is used for training, which is recorded as , which is the cross entropy loss of instruction type classification, is used to classify instruction types so that the model can accurately identify different instruction action types. Matrix, that is:

[0093]

[0094] Here, M represents the number of training samples in a batch. During training, multiple samples are used to train the model. Here, M is the total number of samples input into the model for training. N represents the number of instruction types. is a matrix Middle The samples correspond to The probability value of the instruction type is obtained by processing the context embedding vector through the scene adaptation classification head and passing it through the Softmax function, which reflects the probability of the first instruction type in the current scene. The samples belong to The probability of the type of instruction. Is the instruction prediction branch output The samples correspond to The predicted probability value of each node. The model predicts the probability of each sample belonging to different instruction types, and these predicted probabilities constitute the output of the instruction prediction branch. This formula uses the matrix By calculating the difference between the model prediction results and the probability distribution of instruction types in actual scenarios through cross entropy, the model is prompted to output instruction types that are more in line with the actual situation during training, thereby improving the accuracy of instruction type recognition.

[0095] For the output of the numerical prediction branch prediction, convert it to a decimal r in [0,1] p , then according to [R min ,R max ]Get the final value R p :

[0096] .

[0097] Then calculate the numerical loss :

[0098] .

[0099] Among them, R p are the parameter values ​​predicted by the decoder, Is to Rp Clip to the legal value range By calculating their mean squared error, the model parameters are continuously adjusted during training to make the generated parameters more consistent with the legal range.

[0100] Through multi-task joint training, the model's performance in multiple tasks such as speech recognition, parameter range constraints, and instruction type classification is optimized, thereby improving the model's generalization ability and robustness. Total loss function:

[0101] .

[0102] Among them, the weight coefficient , They respectively represent the importance of each loss function in the total loss. By adjusting these weight coefficients, the training intensity of different tasks can be balanced.

[0103] This application has the following advantages.

[0104] The problem of identifying air traffic control instructions is transformed into a classification and regression problem. For a fixed set of instructions, a classification task is used to identify the category of each instruction. For numerical instructions, a regression task is used to predict their value and unit. This transformation makes the complex instruction recognition problem more structured and easier to handle, enabling the model to more accurately output the instruction type and specific value, providing a basis for subsequent air traffic control operations.

[0105] A two-stream generative adversarial network (TS-AGN) is designed to process speech signals. Its dual-path convolutional architecture processes wideband and narrowband features separately. This design fully leverages information from different frequency bands in the speech signal to effectively capture speech characteristics. Using the generative adversarial network mechanism, TS-AGN learns the characteristics of noise and generates enhanced speech that more closely resembles clean speech. This enables the model to accurately recognize voice commands even in noisy environments, improving the robustness of speech recognition and ensuring accurate transmission of air traffic control instructions in harsh acoustic environments.

[0106] A multimodal feature fusion encoder is designed, in which a dual-branch Transformer structure extracts features from the acoustic and contextual branches, respectively. The acoustic branch focuses on the time-frequency characteristics of speech, capturing information such as speech pitch and rhythm; the contextual branch focuses on the multimodal characteristics of the flight scene, such as flight phase, airspace restrictions, and weather conditions. By fusing the features of these two branches, the model considers not only the meaning of the speech itself but also the current flight phase and airspace conditions when determining whether a command is reasonable. This allows the model to comprehensively utilize both speech and flight scene information to more comprehensively understand air traffic control commands, thereby improving the accuracy and reliability of command recognition.

[0107] A dynamic gated feature fusion mechanism is designed to adaptively adjust the fusion ratio of acoustic and contextual features based on different flight scenarios and speech characteristics. In some cases, speech features may be more important, such as when a command is clearly expressed but has little contextual relevance. In other cases, contextual features may play a key role, such as when the speech signal is subject to some interference but the context allows the command's meaning to be inferred. This flexible fusion approach enables the model to appropriately allocate feature weights based on actual conditions, fully leveraging the strengths of both acoustic and contextual features to improve model performance and adaptability.

[0108] Implicit rule-constrained learning constructs a parameter range prediction head and a scenario-adaptive classification head, enabling the model to predict the legal parameter value range and permitted instruction action types in the current scenario. In aviation control, instruction parameters and action types are strictly regulated and restricted. Through this implicit rule-constrained learning, the model can automatically follow these rules when generating instructions, reducing the generation of illegal instructions at the source. Furthermore, in practical applications, the model may be affected by various interferences and uncertainties, resulting in deviations in the generated instructions. Implicit rule-constrained learning provides the model with a self-correcting and constraint mechanism, making the model more stable and reliable in the face of these interferences.

[0109] A Transformer-based decoder is designed to fuse features, fully leveraging the multimodal information extracted by the encoder to generate accurate instruction text sequences. A rule-based instruction generation system is then designed, combining the outputs of the parameter range prediction head and the scenario adaptation classification head to constrain the generation process. This ensures that the generated instructions are not only semantically accurate but also comply with aviation control regulations and the requirements of current flight scenarios in terms of parameters and action types.

[0110] A loss function is designed and multi-task collaborative optimization is performed. Cross-entropy loss and numerical loss are designed for the instruction prediction branch and the numerical prediction branch, respectively. These are incorporated into the overall loss function for multi-task joint training. This design enables the model to optimize multiple tasks simultaneously, improving performance across different tasks. Regression of numerical values ​​of varying magnitudes is then achieved using the minimum and maximum values ​​of numerical instructions, ultimately achieving more accurate numerical instruction predictions.

[0111] Based on the same inventive concept, embodiments of the present application also provide an air traffic control speech recognition system for implementing the aforementioned air traffic control speech recognition method. The solution provided by this system is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more of the air traffic control speech recognition system embodiments provided below can be found in the limitations of the air traffic control speech recognition method described above and will not be further elaborated here.

[0112] In an exemplary embodiment, Figure 3 As shown, an aviation control speech recognition system is provided, comprising:

[0113] The acquisition module 301 is used to acquire the voice data of the air traffic controller.

[0114] The classification and enhancement module 302 is used to classify and enhance the air traffic controller voice data using a dual-stream adversarial generative network to obtain an enhanced voice signal.

[0115] The feature extraction and fusion module 303 is used to perform feature extraction and fusion based on the enhanced speech signal using a multimodal feature fusion encoder to obtain a context embedding vector and a fusion feature.

[0116] The prediction module 304 is configured to perform prediction based on the context embedding vector using a parameter range prediction head and a scene adaptation classification head to obtain a legal value range of the prediction parameter and an allowed instruction action type.

[0117] The generation module 305 is used to generate an instruction text sequence using a decoder based on the constraints of the fusion feature, the legal value range of the prediction parameter and the allowed instruction action type.

[0118] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store aviation control speech recognition data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an aviation control speech recognition method is implemented.

[0119] Those skilled in the art will understand that Figure 4The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present application and does not constitute a limitation on the computer device to which the solution of the present application is applied. A specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned method embodiments when executing the computer program.

[0120] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the above-mentioned method embodiments when executed by a processor.

[0121] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the above method embodiments are implemented.

[0122] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0123] In this application, all actions to obtain signals, information or data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0124] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0125] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0126] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0127] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for recognizing aviation control speech, characterized in that: The aviation control speech recognition method comprises: Obtain air traffic controller voice data; Classifying and enhancing the air traffic controller voice data using a dual-stream generative adversarial network to obtain an enhanced voice signal; According to the enhanced speech signal, a multimodal feature fusion encoder is used to perform feature extraction and fusion to obtain a context embedding vector and a fusion feature; according to the enhanced speech signal, a multimodal feature fusion encoder is used to perform feature extraction and fusion to obtain a context embedding vector and a fusion feature, specifically including: according to the enhanced speech signal, a multi-layer convolutional Transformer model of the multimodal feature fusion encoder is used to extract speech time-frequency features to obtain an acoustic embedding vector; according to the flight status, an LSTM model of the multimodal feature fusion encoder is used to encode the time series change of the flight status to obtain a context embedding vector; according to the acoustic embedding vector and the context embedding vector, a gating weight of the multimodal feature fusion encoder is used to perform weighted fusion to obtain a fusion feature; Performing prediction using a parameter range prediction head and a scene adaptation classification head based on the context embedding vector to obtain a legal value range of the prediction parameter and an allowed instruction action type; Constraints are performed based on the fusion features, the legal value range of the prediction parameters and the allowed instruction action types, and an instruction text sequence is generated using a decoder.

2. The method for recognizing aviation control speech according to claim 1, characterized in that: The air traffic controller voice data is classified and enhanced using a dual-stream adversarial generative network to obtain an enhanced voice signal, specifically including: Based on the air traffic controller voice data, the dual-path convolution structure of the two-stream adversarial generative network is used to perform depth convolution and point-by-point convolution to obtain the local features and point-by-point convolution results of each channel; The local features of each channel and the point-by-point convolution results are concatenated and convolved using a dual-path convolution structure to obtain an enhanced speech signal.

3. The method for recognizing aviation control speech according to claim 1, wherein: According to the context embedding vector, the parameter range prediction head and the scene adaptation classification head are used to perform prediction to obtain the legal value range of the prediction parameter and the allowed instruction action type, specifically including: According to the context embedding vector, a parameter range prediction head is used to predict and obtain a legal value range of the prediction parameter; the parameter range prediction head is an MLP model including 5 fully connected layers; The scene adaptation classification head is used to perform prediction based on the context embedding vector to obtain the allowed instruction action type; the scene adaptation classification head includes a 5-layer MLP model and a Softmax function connected to the MLP model.

4. The method for recognizing aviation control speech according to claim 1, wherein: The decoder includes a feature extraction network, an instruction prediction branch and a numerical prediction branch; the feature extraction network is connected to the instruction prediction branch and the numerical prediction branch respectively; the instruction prediction branch and the numerical prediction branch both include three fully connected layers; the instruction text sequence includes fixed instructions and numerical instructions.

5. The method for recognizing aviation control speech according to claim 4, characterized in that: The total loss function of the decoder is the weighted sum of the cross entropy loss and the numerical loss.

6. An aviation control speech recognition system, characterized in that: The aviation control speech recognition system includes: Acquisition module, used to obtain air traffic controller voice data; A classification and enhancement module, configured to classify and enhance the air traffic controller voice data using a dual-stream adversarial generative network to obtain an enhanced voice signal; A feature extraction and fusion module is configured to perform feature extraction and fusion based on the enhanced speech signal using a multimodal feature fusion encoder to obtain a context embedding vector and fusion features; perform feature extraction and fusion based on the enhanced speech signal using a multimodal feature fusion encoder to obtain a context embedding vector and fusion features, specifically comprising: extracting speech time-frequency features based on the enhanced speech signal using a multi-layer convolutional Transformer model of the multimodal feature fusion encoder to obtain an acoustic embedding vector; encoding flight status temporal changes based on flight status using an LSTM model of the multimodal feature fusion encoder to obtain a context embedding vector; and performing weighted fusion based on the acoustic embedding vector and the context embedding vector using the gating weights of the multimodal feature fusion encoder to obtain a fusion feature; A prediction module, configured to perform prediction based on the context embedding vector using a parameter range prediction head and a scene adaptation classification head to obtain a legal value range of the prediction parameter and an allowed instruction action type; A generation module is used to constrain the prediction parameter legal value range and the allowed instruction action type according to the fusion feature, and use a decoder to generate an instruction text sequence.

7. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the aviation control speech recognition method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the aviation control speech recognition method according to any one of claims 1 to 5 is implemented.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the aviation control speech recognition method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Method and device for controlling unmanned aerial vehicle based on voice

    CN111554286A

  • Speech enhancement method of deep generative adversarial network

    CN114446314A