Aviation voice transliteration method and device based on multi-mode integration and medium
By employing a multimodal integrated aviation speech transcription method, which utilizes a multimodal large model and a comprehensive loss function to optimize speech and text conversion, the problems of speech interference and language barriers in aviation intercom are solved, achieving efficient and accurate air-to-ground communication.
Patent Information
- Application Number
- CN202511070449.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-10-28
AI Technical Summary
Existing aviation intercom mechanisms are prone to voice signal interference in complex environments, resulting in unclear conversations. Key instructions may be drowned out by noise. Differences in individual expression habits can lead to communication deviating from industry standards. Language barriers increase the risk of information transmission errors. Furthermore, manual repetition and verification are time-consuming and prone to errors.
采用基于多模态集成的航空语音转写方法,通过多模态大模型将初始语音信号转换为目标语言文本,并生成纠正建议,利用语音识别、语音合成和多模态匹配任务,结合综合损失函数优化模型性能,实现双向纠错机制,确保语音和文本的准确转换和规范表达。
It has improved the accuracy and efficiency of air-to-ground communication, lowered the language barrier, reduced duplicate communication caused by misunderstandings, and ensured the accurate transmission of instructions and the security of aviation communication.
Smart Images

Figure CN120853573A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and in particular to an aviation speech transcription method, apparatus and medium based on multimodal integration. Background Technology
[0002] In the aviation field, communication between the control tower and pilots is the "nerve center" for maintaining flight safety. Its importance permeates the entire process from takeoff to landing, and every correct flight instruction is a core pillar ensuring safe and reliable flight operations, directly determining whether an aircraft can avoid risks and accurately execute missions in a complex and ever-changing airspace environment. The existing aviation intercom system is a system based on traditional radio communication, relying on manual execution of procedures and lacking intelligent assistance. Its design logic is deeply rooted in the standardization framework of the International Civil Aviation Organization (ICAO). This system still mainly relies on VHF analog signals to transmit voice, depending on the controllers and pilots' manual mastery and strict adherence to standard terminology.
[0003] However, existing aviation intercom headsets face numerous challenges in practical applications: in complex environments with engine noise and interspersed radio static, voice signals are easily interfered with, leading to often unclear conversations and critical instructions being drowned out by noise; in conversations between control towers and pilots, differences in personal expression habits often cause communication to deviate from industry-standard terminology, such as simplified expressions or accents, easily resulting in misunderstandings; the manual repetition and verification process is not only time-consuming but also prone to errors. According to statistics from the International Civil Aviation Organization, the error rate of critical information transmission in traditional aviation intercom methods exceeds 3%. Furthermore, with the continuous increase in the number of international flights, language barriers between pilots and control towers in different regions not only increase the communication burden but also further raise the risk of errors in information transmission. Summary of the Invention
[0004] This invention provides a method, device, and medium for aviation speech transcription based on multimodal integration, in order to solve the problem of difficulty in reducing the language communication barrier between the control tower and pilots, and to improve the accuracy and efficiency of air-to-ground communication.
[0005] To achieve the above objectives, this application provides an aviation speech transcription method based on multimodal integration, comprising:
[0006] Acquire the initial speech signal from the input party;
[0007] Convert the initial speech signal into target language text;
[0008] Correction suggestions are generated for the target language text based on a multimodal large model and output to the input. The multimodal large model is obtained by training an initial model using a comprehensive loss function. The initial model is established by setting task heads including speech recognition, speech synthesis, and multimodal matching tasks based on the multimodal features of speech. The comprehensive loss function is established by quantifying the difference between generated data and real data in each task of the task head based on the loss function.
[0009] The system acquires the final speech signal after the input party has corrected its expression according to the correction suggestions, generates instructions in the target language based on the final speech signal, and sends them to the receiver; wherein the input party and the receiver are the control tower and the pilot, respectively, forming a two-way communication mechanism.
[0010] This invention converts the initial speech signal from the input party into text, avoiding interference from cabin noise and improving readability. The initial model is configured with three tasks: speech recognition, speech synthesis, and multimodal matching, corresponding to speech-to-text, text-to-speech, and speech-to-text alignment, respectively, and can collaboratively address issues such as accent discrepancies and terminology misuse. A comprehensive loss function, by quantifying the differences between the output of each task and real data, optimizes the overall model performance, ensuring recognition and conversion accuracy in complex scenarios. Furthermore, multi-task training allows the model to simultaneously master multi-dimensional features of speech, text, and context, adapting to complex communication scenarios. Based on this, a multimodal large-scale model generates correction suggestions for the target language text, directly standardizing non-standard expressions and reducing repetitive communication due to misunderstandings, thus improving instruction accuracy and shortening the time required for a single communication. The two-way error correction mechanism, with the control tower and pilot as input parties, enables end-to-end standardization verification, further ensuring the accuracy and efficiency of air-to-ground communication. In addition, presenting instructions to the receiver in the target language overcomes language barriers.
[0011] Compared to existing technologies, this invention, through a multimodal large model-driven closed-loop error correction mechanism and a two-way real-time communication framework, can generate accurate instructions in the target language from the corrected speech of the input party and send them to the receiver. Therefore, it can solve the problem of difficulty in reducing the language communication barrier between the control tower and the pilot, as well as improving the accuracy and efficiency of air-to-ground communication.
[0012] As a preferred embodiment, the initial model is established by setting task heads including speech recognition, speech synthesis, and multimodal matching tasks based on the multimodal features of speech, specifically:
[0013] Acquire multimodal data in an aviation communication environment;
[0014] The speech data in the multimodal data is feature-encoded using a convolutional neural network to obtain a speech feature sequence;
[0015] The text data in the multimodal data is feature-encoded using a bidirectional long short-term memory network to obtain a text feature sequence;
[0016] The speech feature sequence and the text feature sequence are weighted and fused to obtain multimodal features;
[0017] The task head, which includes a speech recognition task, a speech synthesis task, and a multimodal matching task, is set according to the multimodal features. The initial model is established based on the task head. The speech recognition task converts the feature sequence into a text sequence through a decoder. The speech synthesis task generates the corresponding speech waveform based on the text features through a vocoder. The multimodal matching task outputs the matching probability that the speech and text have the same semantics through contrastive learning.
[0018] This preferred solution, through weighted fusion of speech and text features, enables the model to simultaneously grasp the correlation information of both modalities, avoiding the limitations of a single modality. Speech recognition and synthesis tasks enable bidirectional conversion between speech and text, meeting the needs of command transmission and repetition in aviation communications; the multimodal matching task ensures semantic consistency between speech and text through contrastive learning, technically avoiding the risk of discrepancies between spoken and written language.
[0019] As a preferred embodiment, the comprehensive loss function is established by quantifying the difference between generated data and real data based on the loss function in each task of the task header, specifically as follows:
[0020] In the speech recognition task, a first loss function is established based on the CTC loss function to measure the difference between the text transcription results output by the model and the preset real labels;
[0021] In the speech synthesis task, a second loss function is established by calculating the difference between the speech signal generated by the model and the preset real speech signal based on the mean square error loss function.
[0022] In the multimodal matching task, a third loss function is established to narrow the distance between speech and text features with the same semantic information based on the contrastive loss function;
[0023] The comprehensive loss function is established based on the first loss function, the second loss function, and the third loss function.
[0024] This preferred solution uses the CTC loss function in the speech recognition task to solve the alignment ambiguity between speech and text, accurately measuring the difference between the transcription result and the real label. In the speech synthesis task, it uses the mean squared error loss function, which is sensitive to continuous signal errors and effectively measures the acoustic feature differences between generated speech and real speech, ensuring the naturalness and accuracy of the synthesized speech. In the multimodal matching task, it uses the contrastive loss function, which, by narrowing the feature distance of similar semantics, ensures semantic consistency between speech and text, avoiding information misalignment. The comprehensive loss method binds and optimizes these three tasks, improving the performance of a single task while also considering other tasks, avoiding getting trapped in local optima.
[0025] As a preferred embodiment, the initial speech signal is converted into target language text, specifically as follows:
[0026] The initial speech signal is subjected to feature extraction and encoding operations using a one-dimensional convolutional neural network to obtain processed audio features.
[0027] The processed audio features are then masked to generate mask information; wherein the mask information is used to distinguish between speech and noise.
[0028] The audio features are decoded based on the mask information to obtain the enhanced speech signal;
[0029] The enhanced speech signal is converted into text content corresponding to the target language to obtain the target language text.
[0030] This preferred solution effectively captures local features and global trends in speech signals using a one-dimensional convolutional neural network, preserving more key speech information compared to traditional methods. By explicitly distinguishing speech from noise through masking, essentially labeling audio features, it accurately locates noise regions and prioritizes filtering them during decoding, preventing noise from being misidentified as valid speech.
[0031] As a preferred embodiment, the enhanced speech signal is converted into text content corresponding to the target language to obtain the target language text, specifically as follows:
[0032] The enhanced speech signal is converted into the corresponding source language text to obtain the first language text;
[0033] The first language text is preprocessed according to the multimodal large model to obtain a second language text that conforms to the preset aviation dialogue standard rule set; wherein, the preprocessing includes locating and removing non-semantic characters introduced during speech transcription based on the noise feature information in the first language text, and mapping the spoken expression text in the first language text to terminology expression text that conforms to the preset aviation dialogue standard rule set.
[0034] The second language text is converted into the text content corresponding to the target language to obtain the target language text.
[0035] This preferred solution effectively filters out background noise or transcription errors remaining after speech signal enhancement by locating and removing non-semantic characters introduced during speech transcription, ensuring the purity of the first-language text. First, the source language text is standardized into a second-language text conforming to industry rules, and then translated into the target language. This ensures that the target language text accurately conveys the semantics while strictly adhering to aviation dialogue standards, avoiding imprecise expressions in cross-language translation.
[0036] As a preferred approach, correction suggestions are generated for the target language text based on a multimodal large model and output to the input, specifically as follows:
[0037] The target language text is structurally decomposed based on the multimodal large model, and a preset aviation dialogue norm rule set is invoked to compare the structurally decomposed target language text with the corresponding entries in the preset aviation dialogue norm rule set in multiple dimensions.
[0038] If the target language text contains expressions that deviate from the preset aviation dialogue specification rule set, then corresponding correction suggestions are generated according to the preset aviation dialogue specification rule set;
[0039] The correction suggestions can be converted into voice prompts and played to the input device via an audio channel, or the correction suggestions can be displayed on the input device's screen in structured text form.
[0040] This preferred solution decomposes the target language text into structured units, breaking down complex dialogue content into multiple analyzable units. These units are then compared across multiple dimensions with a pre-defined set of aviation dialogue norms, comprehensively identifying potential problems and avoiding overlooked details. Furthermore, this comparison method ensures that corrective suggestions fully align with professional aviation dialogue standards. It offers both voice prompts and structured text display as output options, allowing for flexible selection based on the input user's real-time status, enhancing ease of use.
[0041] This application also provides an aviation speech transcription device based on multimodal integration, including a signal module, a text module, a correction module and a communication module;
[0042] The signal module is used to acquire the initial voice signal from the input party.
[0043] The text module is used to convert the initial speech signal into target language text;
[0044] The correction module is used to generate correction suggestions for the target language text based on a multimodal large model and output them to the input. The multimodal large model is obtained by training an initial model using a comprehensive loss function. The initial model is established by setting task heads including speech recognition, speech synthesis, and multimodal matching tasks based on the multimodal features of speech. The comprehensive loss function is established by quantifying the difference between generated data and real data in each task of the task head based on the loss function.
[0045] The communication module is used to acquire the final speech signal after the input party has corrected its expression according to the correction suggestions, generate instructions presented in the target language based on the final speech signal, and send them to the receiver; wherein the input party and the receiver are the control tower and the pilot, respectively, forming a two-way communication mechanism.
[0046] As a preferred embodiment, the correction module includes a data unit, a speech unit, a text unit, a fusion unit, and a model unit;
[0047] The data unit is used to acquire multimodal data in an aviation communication environment;
[0048] The speech unit is used to perform feature encoding on the speech data in the multimodal data according to the convolutional neural network to obtain a speech feature sequence;
[0049] The text unit is used to encode the text data in the multimodal data according to the bidirectional long short-term memory network to obtain a text feature sequence;
[0050] The fusion unit is used to perform weighted fusion of the speech feature sequence and the text feature sequence to obtain multimodal features;
[0051] The model unit is used to set the task head, which includes a speech recognition task, a speech synthesis task, and a multimodal matching task, according to the multimodal features, and to establish the initial model according to the task head; wherein, the speech recognition task converts the feature sequence into a text sequence through a decoder, the speech synthesis task generates the corresponding speech waveform according to the text features through a vocoder, and the multimodal matching task outputs the matching probability that the speech and text have the same semantics through contrastive learning.
[0052] As a preferred embodiment, the correction module includes an identification unit, a synthesis unit, a matching unit, and a comprehensive unit;
[0053] The recognition unit is used to establish a first loss function in the speech recognition task by measuring the difference between the text transcription result output by the model and the preset real label according to the CTC loss function.
[0054] The synthesis unit is used to calculate the difference between the speech signal generated by the model and the preset real speech signal according to the mean square error loss function in the speech synthesis task, and establish a second loss function.
[0055] The matching unit is used to establish a third loss function in the multimodal matching task by narrowing the distance between speech and text features with the same semantic information based on the contrast loss function.
[0056] The synthesis unit is used to establish the synthesis loss function based on the first loss function, the second loss function, and the third loss function.
[0057] This application also provides a storage medium storing a computer program, which is called and executed by a computer to implement the aviation speech transcription method based on multimodal integration as described above. Attached Figure Description
[0058] Figure 1 This is a flowchart illustrating an aviation speech transcription method based on multimodal integration provided in an embodiment of this application.
[0059] Figure 2 This is a flowchart of the speech enhancement process provided in an embodiment of this application;
[0060] Figure 3 This is a flowchart of the multilingual conversion process provided in the embodiments of this application;
[0061] Figure 4 This is a flowchart of the dialogue standard calibration provided in the embodiments of this application;
[0062] Figure 5 This is a system architecture diagram provided in the embodiments of this application;
[0063] Figure 6 This is a schematic diagram of the structure of an aviation speech transcription device based on multimodal integration provided in an embodiment of this application. Detailed Implementation
[0064] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0065] In the description of this application, it should be understood that the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first," "second," and "third" may explicitly or implicitly include one or more of that feature.
[0066] The aviation speech transcription method based on multimodal integration provided in this application aims to reduce the language communication barrier between the control tower and pilots. While improving the convenience of communication, it also achieves a dual improvement in the accuracy and efficiency of air-to-ground communication by correcting erroneous instructions from the speaker in real time, thus building a more efficient and reliable aviation communication system.
[0067] Example 1:
[0068] Please see Figure 1 The embodiments of this application provide an aviation speech transcription method based on multimodal integration, including S1 to S4, and the specific implementation steps are as follows:
[0069] S1. Obtain the initial speech signal from the input party.
[0070] Step S1 in this embodiment of the application is specifically as follows:
[0071] The initial voice signal emitted by the input party (such as the control tower or pilot) is captured in real time through the microphone array or voice acquisition module of the aviation intercom system. This signal contains the original voice waveform, background environmental noise and possible transmission interference, forming the original data input source for the multimodal processing flow.
[0072] S2. Convert the initial speech signal into target language text.
[0073] Step S2 in this embodiment of the application is specifically as follows:
[0074] After the initial speech signal is input into the Conv-TasNet speech enhancement algorithm model, preprocessing operations such as normalization are first performed on the initial speech signal to ensure that it meets the model's input criteria. Then, the preprocessed initial speech signal is fed into a one-dimensional convolutional neural network. By setting a convolutional kernel size of 40 and a convolutional stride of 20, feature extraction and encoding are performed on the signal to obtain the processed audio features. The Conv-TasNet speech enhancement algorithm model is an end-to-end speech separation and enhancement framework based on deep convolutional neural networks, specifically designed to solve the problem of speech signal extraction in complex environments, and is particularly suitable for scenarios with strong noise and multi-source interference.
[0075] The processed audio features are input into a 6-dimensional, 256-layer Transformer model for masking. Combined with a multimodal large model, mask information is generated to accurately distinguish speech signals from noise components. The multimodal large model is used to provide feedback on environmental features and identify noise patterns, thus providing cross-modal references for the Transformer to generate masks and compensating for the information loss of a single audio modality in complex environments. The specific construction process of the multimodal large model will be described in detail in subsequent sections.
[0076] Subsequently, a one-dimensional convolutional layer with a kernel size of 40 and a stride of 20 is used to decode the processed audio features based on the generated mask information. By reconstructing and restoring the effective speech features identified by the mask, the final output is one-dimensional decoded audio data, thus obtaining the enhanced speech signal.
[0077] The ASR engine is used to convert the enhanced speech signal into the corresponding source language text to obtain the first language text.
[0078] Leveraging the multilingual understanding and generation capabilities of a multimodal large-scale model, the first-language text is preprocessed to obtain second-language text conforming to a pre-defined aviation dialogue norm rule set. This preprocessing includes locating and removing non-semantic characters introduced during speech-to-text transcription based on noise features in the first-language text, and mapping colloquial expressions in the first-language text to terminological expressions conforming to the pre-defined aviation dialogue norm rule set. The "pre-defined aviation dialogue norm rule set" is a standardized language interaction framework built upon established dialogue standards in the aviation industry. Its core is to transform industry-standard aviation terminology, command formats, and dialogue processes into a computer-recognizable rule system, ensuring the accuracy, standardization, and professionalism of aviation dialogues through systematic standardization. This rule set achieves standardization from multiple dimensions: clarifying standard expressions for aviation-specific terminology to avoid colloquial or ambiguous expressions; standardizing sentence structure and logical order in dialogues to ensure clear information transmission; defining prohibited content such as vague words and non-standard abbreviations; and providing emergency dialogue templates for special scenarios, forming a comprehensive language interaction norm system covering all scenarios.
[0079] In response to the frequent language conversion needs in aviation scenarios, such as the translation between English and other languages, the target language is pre-defined, and a multi-language conversion model is built or an appropriate language conversion interface is called to achieve accurate and efficient conversion between different languages and meet the standardization requirements of cross-language interaction in aviation communications.
[0080] Relying on a multilingual conversion model or a compatible language conversion interface, second-language text is converted into corresponding text content in the target language, resulting in target language text. This conversion process strictly adheres to the principle of semantic consistency, ensuring that the converted text is highly semantically consistent with the original speech information, fluent and accurate, and precisely conveys the key information in the original speech, ultimately generating target language text that conforms to the expression norms of the target language.
[0081] For examples of this application, please refer to [link / reference]. Figure 2 , Figure 2 This is a flowchart of speech enhancement provided in this application embodiment, which shows the specific processing flow of processing the initial speech signal (original audio) to obtain the enhanced speech signal, including the encoding stage, the masking stage and the decoding stage.
[0082] This embodiment S2 utilizes a one-dimensional convolutional neural network to effectively capture local features and global trends in speech signals, preserving key speech information better than traditional methods. By explicitly distinguishing speech from noise through masking, it essentially labels audio features, accurately locating noise regions and prioritizing filtering during decoding to prevent noise from being misidentified as valid speech.
[0083] Furthermore, by locating and removing non-semantic characters introduced during speech transcription, it is possible to effectively filter out background noise or transcription errors remaining after speech signal enhancement, ensuring the purity of the first-language text. First, standardizing the source language text into a second-language text that conforms to industry rules, and then translating it into the target language, ensures that the target language text accurately conveys the semantics while strictly adhering to aviation dialogue standards, avoiding issues of imprecise expression in cross-language translation.
[0084] S3. Generate correction suggestions for the target language text based on a multimodal large model and output them to the input; wherein, the multimodal large model is obtained by training the initial model according to the comprehensive loss function; the initial model is established by setting task heads including speech recognition task, speech synthesis task and multimodal matching task according to the multimodal features of speech; the comprehensive loss function is established by quantifying the difference between generated data and real data based on the loss function in each task of the task head.
[0085] Step S3 in this embodiment includes S3.1 to S3.3, specifically as follows:
[0086] S3.1 Acquire multimodal data in the aviation communication environment, including voice samples, associated text data, and aviation dialogue-related images and other data that are interfered with by aircraft engine noise and wind noise in complex aviation environments;
[0087] The speech data in the multimodal data is feature-encoded using a convolutional neural network, and the long-term dependencies in the speech data are captured using a self-attention mechanism to obtain the speech feature sequence.
[0088] The text feature sequence is obtained by encoding the text data in the multimodal data using a bidirectional long short-term memory network.
[0089] First, pooling operations are used to adjust the length of the feature sequences, ensuring precise alignment of the speech and text feature sequences in the time dimension. Then, the two are concatenated along the feature dimension. Next, a multi-head self-attention mechanism is used to dynamically weight and fuse the concatenated hybrid features, ultimately yielding multimodal features. In this process, the attention mechanism autonomously learns and assigns corresponding weight coefficients based on the actual importance of different modal features in the current task. Specifically, in aviation voice command analysis scenarios, if the speech signal has high clarity, the system assigns greater weight to the speech features; if the text description contains key information such as "altitude instructions" or "runway number," the weight of the text features is increased accordingly, effectively highlighting core features and suppressing interference from irrelevant information.
[0090] Based on multimodal features, set task heads including speech recognition, speech synthesis, and multimodal matching tasks, and build an initial model based on the task heads;
[0091] The speech recognition task converts feature sequences into text sequences using a CTC decoder or attention decoder. The speech synthesis task generates corresponding speech waveforms based on text features using a vocoder. The multimodal matching task outputs the matching probability of speech and text sharing the same semantic meaning through contrastive learning, such as verifying whether a speech command and a text command are consistent, thereby outputting the matching probability. Furthermore, the CTC decoder (Connectionist Temporal Classification Decoder) is an algorithm specifically designed for processing time-series data, playing a crucial role, especially in the field of speech recognition.
[0092] In this embodiment S3.1, by weighted fusion of speech and text features, the model can simultaneously grasp the correlation information of the two modalities, avoiding the limitations of a single modality. The speech recognition and synthesis tasks can achieve bidirectional conversion between speech and text, meeting the needs of command transmission and repetition in aviation communications; the multimodal matching task ensures the semantic consistency between speech and text through contrastive learning, technically avoiding the risk of discrepancies between spoken and written language.
[0093] S3.2 In speech recognition tasks, the first loss function is established by measuring the difference between the text transcription results output by the model and the preset real labels based on the CTC loss function or the attention loss function. Among them, the CTC loss function (Connectionist Temporal Classification Loss) is a loss function specifically used for temporal sequence modeling and is widely used in scenarios such as speech recognition and handwriting recognition where the lengths of the input and output sequences are inconsistent.
[0094] In speech synthesis tasks, a second loss function is established by calculating the difference between the speech signal generated by the model and the preset real speech signal based on the mean square error loss function or the spectral convergence loss function.
[0095] In multimodal matching tasks, a third loss function is established to narrow the distance between speech and text features with the same semantic information while widening the distance between features with different semantic information, based on the contrastive loss function. It should be noted that this setting is based on the speech and text matching task as an example. In practical applications, it can be adjusted and set according to other specific task types (such as speech and image matching, text and image matching, etc.).
[0096] A comprehensive loss function is established based on the first loss function, the second loss function, and the third loss function.
[0097] Based on multimodal data, a contrastive learning method is used to construct sample pairs; positive sample pairs consist of speech features and text features carrying the same semantic information, while negative sample pairs consist of feature combinations belonging to different semantic information.
[0098] The Adam optimization algorithm and its variants are employed, with the comprehensive loss function as the optimization objective. Initial model training is conducted using multimodal data and constructed sample pairs. During training, a learning rate decay strategy is simultaneously introduced, gradually reducing the learning rate as the training epochs increase. This allows the model to more accurately capture subtle features hidden in the data in later iterations, further improving the model's learning accuracy and generalization ability, thus obtaining a large multimodal model. Furthermore, to prevent overfitting and enhance generalization ability, multiple regularization strategies are employed: firstly, L2 regularization, which limits the parameter size by adding the L2 norm term of the model parameters to the loss function to avoid excessive model complexity; secondly, Dropout technology, which randomly discards some neuron outputs during training, forcing the model to rely on distributed feature learning and improving robustness against interference; and thirdly, data augmentation techniques, such as adding noise, adjusting speech rate and pitch, and performing synonym replacement and sentence restructuring on text data, to expand the diversity of training data and strengthen the model's adaptability to variant samples. Among them, the Adam (Adaptive Moment Estimation) optimization algorithm is a deep learning optimization algorithm that combines momentum and adaptive learning rate characteristics, and is widely used in neural network training.
[0099] After the multimodal large-scale model is built, it is necessary to conduct systematic evaluation and validation of the model regularly: using validation set data, comprehensively test the model's performance metrics on core tasks such as speech enhancement, multilingual conversion, and dialogue standard calibration. Based on the evaluation results, iterative optimization of the model should be carried out in a targeted manner, such as adjusting the inter-layer connection structure of the network and optimizing hyperparameter configurations such as the learning rate decay coefficient, so as to continuously improve the model's prediction accuracy and robustness against interference in complex scenarios.
[0100] The trained and optimized multimodal large-scale model is embedded into the software architecture of the aviation intercom headset system, ensuring seamless integration and efficient data interaction with hardware devices such as the headset's voice input module and OLED display, as well as software modules such as voice enhancement, multilingual conversion, and dialogue standard calibration. Simultaneously, compatible software interfaces and data transmission protocols are developed to construct a real-time processing and analysis workflow for multimodal data such as speech and text using the multimodal large-scale model. Specifically, after beamforming the raw audio, the model sequentially performs speech recognition to text conversion, Conv-TasNet noise reduction, and text recognition analysis and optimization, ultimately converting the standard-compliant text content into speech for output to the receiver, forming a fully automated processing loop.
[0101] In this embodiment, S3.2 uses the CTC loss function in the speech recognition task to solve the alignment ambiguity problem between speech and text, accurately measuring the difference between the transcription result and the real label. In the speech synthesis task, the mean squared error loss function is used. This function is sensitive to continuous signal errors and can effectively measure the acoustic feature difference between generated speech and real speech, ensuring the naturalness and accuracy of synthesized speech. In the multimodal matching task, the contrastive loss function is used. By narrowing the feature distance of similar semantics, it can ensure the semantic consistency between speech and text, avoiding information misalignment. The comprehensive loss combines and optimizes these three tasks, which can improve the performance of a single task while taking into account other tasks, avoiding getting trapped in local optima. Furthermore, training based on multimodal data from the aviation communication environment can accurately capture the professional features of speech and text in this scenario, enabling the model to adapt to industry needs from the bottom up.
[0102] S3.3. Based on the multimodal large model, the target language text is structurally decomposed, and the preset aviation dialogue norm rule set is called. The target language text after structural decomposition is compared with the corresponding items in the preset aviation dialogue norm rule set in multiple dimensions.
[0103] If the target language text contains expressions that deviate from the preset aviation dialogue standard rules set, such as the use of non-standard aviation terminology or errors in the instruction structure, then corresponding correction suggestions will be generated based on the preset aviation dialogue standard rules set.
[0104] Corrective suggestions are converted into voice prompts and played to the input party via an audio channel, or displayed as structured text on the input party's screen to guide them in following the corrective guidelines. Specifically, this might involve playing a voice prompt such as "Your terminology is incorrect; it should be XX" through headphones, or displaying the corrective guidelines on an OLED screen, ensuring the input party clearly understands and adopts the corrective information. During the conversion of corrective suggestions into voice prompts or text displays, a multi-dimensional, multi-modal fusion analysis is performed on the corrective suggestions themselves, the input party's dialogue history, and the current flight status. This accurately generates voice prompts or structured text displays tailored to the specific scenario, ensuring the corrective information is both relevant to the real-time communication context and matches the urgency and priority of the flight operation phase, thereby improving the input party's acceptance and execution efficiency of the corrective suggestions.
[0105] This embodiment, S3.3, decomposes the target language text using structured decomposition, breaking down complex dialogue content into multiple analyzable units. These units are then compared across multiple dimensions with a pre-defined aviation dialogue specification rule set, comprehensively identifying potential problems and avoiding overlooked details. Furthermore, this comparison method ensures that corrective suggestions fully align with professional aviation dialogue standards. It provides two output modes: voice prompts and structured text display, allowing for flexible selection based on the input user's real-time status, enhancing ease of use.
[0106] For embodiments S2 and S3, when the control tower or pilots communicate using a non-standard language, the system can achieve rapid and accurate target language conversion, completely eliminating language barriers to ensure barrier-free communication between the two parties. This reduces the risk of communication errors caused by language barriers at the source, and simultaneously improves work efficiency and operational safety. At the same time, the multimodal big data model performs real-time monitoring and intelligent analysis of each round of dialogue. Once it identifies expressions that deviate from the standard, it immediately triggers a correction mechanism, pushing standardized suggestions to the user through voice prompts or text displays, guiding them to complete the dialogue interaction according to standard scripts, thereby further enhancing the accuracy and professional standardization of information transmission.
[0107] S4. Obtain the final speech signal after the input party has corrected its expression based on the correction suggestions, generate instructions in the target language based on the final speech signal and send them to the receiver; wherein, the input party and the receiver are the tower and the pilot, respectively, forming a two-way communication mechanism.
[0108] Step S4 in this embodiment of the application is specifically as follows:
[0109] Obtain the final speech signal after the input party corrects their expression based on the correction suggestions;
[0110] Based on the above processing methods, the final voice signal undergoes further voice processing, language conversion, and standard calibration to ensure it strictly conforms to the expression requirements of the preset aviation dialogue specification rule set. On this basis, instructions presented in the target language are generated from the processed final voice signal. This not only ensures clear and intelligible speech and a speech rate within the standard range for aviation communication, but also that the intonation conforms to industry standards, comprehensively adapting to the communication needs of aviation scenarios.
[0111] Furthermore, the instruction is simultaneously sent to the receiver's visual interaction layer, such as being displayed on a miniature OLED display on the headset; and the display interface adopts a minimalist design logic to highlight key instruction content such as altitude parameters and runway number, so that the receiver can quickly complete visual confirmation.
[0112] It should be noted that in this first embodiment, the input party and the receiver correspond to the control tower and the pilot, respectively. The two form a two-way communication mechanism in which the tower can act as the input party to send instructions and as the receiver to receive feedback from the pilot, and vice versa, thereby realizing a closed loop of information interaction.
[0113] For examples of this application, please refer to [link / reference]. Figure 3-5 ;
[0114] Figure 3 This is a multilingual conversion flowchart provided in this application embodiment, illustrating the voice data conversion process of this embodiment one, including the logical relationships and execution steps of each stage from data input to data output.
[0115] Figure 4 This is a flowchart of the dialogue standard calibration process provided in this application embodiment, which shows the dialogue calibration process of this embodiment one, including the logical relationships and execution steps of each stage from data input to specification output.
[0116] Figure 5 This is a system architecture diagram provided in the embodiments of this application. Specifically, it takes "obtaining tower instructions and outputting instructions to the pilot" as the application scenario, and clearly shows the overall system architecture involved in the process in Embodiment 1.
[0117] Overall, this embodiment has the following beneficial effects:
[0118] This application converts the initial speech signal from the input party into text, avoiding interference from cabin noise and improving readability. The initial model is configured with three tasks: speech recognition, speech synthesis, and multimodal matching, corresponding to speech-to-text, text-to-speech, and speech-to-text alignment, respectively, and can collaboratively address issues such as accent discrepancies and terminology misuse. The comprehensive loss function optimizes the overall model performance by quantifying the differences between the output of each task and real data, ensuring recognition and conversion accuracy in complex scenarios. Furthermore, multi-task training allows the model to simultaneously master multi-dimensional features of speech, text, and context, adapting to complex communication scenarios. Based on this, correction suggestions are generated for the target language text using a large multimodal model, directly standardizing non-standard expressions and reducing repetitive communication due to misunderstandings, thus improving instruction accuracy and shortening the time required for a single communication. The two-way error correction mechanism, with the control tower and pilot as input parties, enables end-to-end standardization verification, further ensuring the accuracy and efficiency of air-to-ground communication. In addition, presenting instructions to the receiver in the target language overcomes language barriers.
[0119] In summary, this application, supported by multimodal technology, integrates key technologies such as speech enhancement processing algorithms and dual-channel feedback mechanisms. It also combines comprehensive processing capabilities including speech enhancement, multilingual translation, and language output standardization. This effectively reduces communication costs between air traffic controllers and pilots, significantly improving communication efficiency. Furthermore, it accurately resolves issues such as non-standard expressions and pronunciation differences across regions, achieving unified and standardized instruction output through standardized processing. This ensures pilots receive accurate flight instructions promptly even in complex linguistic environments, strengthening flight safety at the language interaction level and comprehensively guaranteeing flight safety and operational reliability.
[0120] Example 2:
[0121] Please see Figure 6 The embodiments of this application provide an aviation speech transcription device based on multimodal integration, including a signal module 10, a text module 20, a correction module 30 and a communication module 40;
[0122] Among them, the signal module 10 is used to acquire the initial voice signal of the input party;
[0123] Text module 20 is used to convert the initial speech signal into target language text;
[0124] The correction module 30 is used to generate correction suggestions for the target language text based on the multimodal large model and output them to the input. The multimodal large model is obtained by training an initial model based on a comprehensive loss function. The initial model is established by setting task heads including speech recognition, speech synthesis, and multimodal matching tasks based on the multimodal features of speech. The comprehensive loss function is established by quantifying the difference between generated data and real data based on the loss function in each task of the task head.
[0125] The communication module 40 is used to acquire the final speech signal after the input party has corrected its expression according to the correction suggestions, generate instructions in the target language based on the final speech signal and send them to the receiver; wherein the input party and the receiver are the tower and the pilot, respectively, forming a two-way communication mechanism.
[0126] In one embodiment, the signal module 10 specifically comprises:
[0127] The initial voice signal emitted by the input party (such as the control tower or pilot) is captured in real time through the microphone array or voice acquisition module of the aviation intercom system. This signal contains the original voice waveform, background environmental noise and possible transmission interference, forming the original data input source for the multimodal processing flow.
[0128] In one embodiment, text module 20 specifically comprises:
[0129] After the initial speech signal is input into the Conv-TasNet speech enhancement algorithm model, preprocessing operations such as normalization are first performed on the initial speech signal to ensure that it meets the model's input criteria. Then, the preprocessed initial speech signal is fed into a one-dimensional convolutional neural network. By setting a convolutional kernel size of 40 and a convolutional stride of 20, feature extraction and encoding are performed on the signal to obtain the processed audio features. The Conv-TasNet speech enhancement algorithm model is an end-to-end speech separation and enhancement framework based on deep convolutional neural networks, specifically designed to solve the problem of speech signal extraction in complex environments, and is particularly suitable for scenarios with strong noise and multi-source interference.
[0130] The processed audio features are input into a 6-dimensional, 256-layer Transformer model for masking. Combined with a multimodal large model, mask information is generated to accurately distinguish speech signals from noise components. The multimodal large model is used to provide feedback on environmental features and identify noise patterns, thus providing cross-modal references for the Transformer to generate masks and compensating for the information loss of a single audio modality in complex environments. The specific construction process of the multimodal large model will be described in detail in subsequent sections.
[0131] Subsequently, a one-dimensional convolutional layer with a kernel size of 40 and a stride of 20 is used to decode the processed audio features based on the generated mask information. By reconstructing and restoring the effective speech features identified by the mask, the final output is one-dimensional decoded audio data, thus obtaining the enhanced speech signal.
[0132] The ASR engine is used to convert the enhanced speech signal into the corresponding source language text to obtain the first language text.
[0133] Leveraging the multilingual understanding and generation capabilities of a multimodal large-scale model, the first-language text is preprocessed to obtain second-language text conforming to a pre-defined aviation dialogue norm rule set. This preprocessing includes locating and removing non-semantic characters introduced during speech-to-text transcription based on noise features in the first-language text, and mapping colloquial expressions in the first-language text to terminological expressions conforming to the pre-defined aviation dialogue norm rule set. The "pre-defined aviation dialogue norm rule set" is a standardized language interaction framework built upon established dialogue standards in the aviation industry. Its core is to transform industry-standard aviation terminology, command formats, and dialogue processes into a computer-recognizable rule system, ensuring the accuracy, standardization, and professionalism of aviation dialogues through systematic standardization. This rule set achieves standardization from multiple dimensions: clarifying standard expressions for aviation-specific terminology to avoid colloquial or ambiguous expressions; standardizing sentence structure and logical order in dialogues to ensure clear information transmission; defining prohibited content such as vague words and non-standard abbreviations; and providing emergency dialogue templates for special scenarios, forming a comprehensive language interaction norm system covering all scenarios.
[0134] In response to the frequent language conversion needs in aviation scenarios, such as the translation between English and other languages, the target language is pre-defined, and a multi-language conversion model is built or an appropriate language conversion interface is called to achieve accurate and efficient conversion between different languages and meet the standardization requirements of cross-language interaction in aviation communications.
[0135] Relying on a multilingual conversion model or a compatible language conversion interface, second-language text is converted into corresponding text content in the target language, resulting in target language text. This conversion process strictly adheres to the principle of semantic consistency, ensuring that the converted text is highly semantically consistent with the original speech information, fluent and accurate, and precisely conveys the key information in the original speech, ultimately generating target language text that conforms to the expression norms of the target language.
[0136] For examples of this application, please refer to [link / reference]. Figure 2 , Figure 2 This is a flowchart of speech enhancement provided in this application embodiment, which shows the specific processing flow of processing the initial speech signal (original audio) to obtain the enhanced speech signal, including the encoding stage, the masking stage and the decoding stage.
[0137] In this embodiment, the text module 20, through a one-dimensional convolutional neural network, can effectively capture local features and global trends in speech signals, preserving key information of speech better than traditional methods. By clearly distinguishing speech from noise through a mask, it essentially labels audio features, accurately locating noise regions and filtering them during decoding to prevent noise from being misidentified as valid speech.
[0138] Furthermore, by locating and removing non-semantic characters introduced during speech transcription, it is possible to effectively filter out background noise or transcription errors remaining after speech signal enhancement, ensuring the purity of the first-language text. First, standardizing the source language text into a second-language text that conforms to industry rules, and then translating it into the target language, ensures that the target language text accurately conveys the semantics while strictly adhering to aviation dialogue standards, avoiding issues of imprecise expression in cross-language translation.
[0139] In one embodiment, the correction module 30 includes a data unit, a speech unit, a text unit, a fusion unit, a model unit, a recognition unit, a synthesis unit, a matching unit, a comprehensive unit, and a conversion unit;
[0140] The data unit is used to acquire multimodal data in the aviation communication environment, including voice samples, associated text data, and aviation dialogue-related images, which are affected by aircraft engine noise and wind noise in complex aviation environments.
[0141] The speech unit is used to encode the speech data in multimodal data according to the convolutional neural network, and to capture the long-term dependencies in the speech data using the self-attention mechanism to obtain the speech feature sequence.
[0142] The text unit is used to encode the text data in the multimodal data according to the bidirectional long short-term memory network to obtain the text feature sequence;
[0143] The fusion unit first adjusts the feature sequence length through pooling operations to achieve precise alignment of the speech and text feature sequences in the time dimension. Then, it concatenates and integrates the two along the feature dimension. Next, a multi-head self-attention mechanism dynamically weights and fuses the concatenated mixed features to ultimately obtain multimodal features. In this process, the attention mechanism autonomously learns and assigns corresponding weight coefficients based on the actual importance of different modal features in the current task. Specifically, in aviation voice command analysis scenarios, if the speech signal has high clarity, the system will assign greater weight to the speech features; if the text description contains key information such as "altitude instructions" or "runway number," the weight of the text features will be increased accordingly, thereby effectively highlighting core features and suppressing interference from irrelevant information.
[0144] The model unit is used to set task heads, including speech recognition tasks, speech synthesis tasks, and multimodal matching tasks, based on multimodal features, and to build an initial model based on the task heads.
[0145] The speech recognition task converts feature sequences into text sequences using a CTC decoder or attention decoder. The speech synthesis task generates corresponding speech waveforms based on text features using a vocoder. The multimodal matching task outputs the matching probability of speech and text sharing the same semantic meaning through contrastive learning, such as verifying whether a speech command and a text command are consistent, thereby outputting the matching probability. Furthermore, the CTC decoder (Connectionist Temporal Classification Decoder) is an algorithm specifically designed for processing time-series data, playing a crucial role, especially in the field of speech recognition.
[0146] In this embodiment, the data unit, speech unit, text unit, fusion unit, and model unit achieve weighted fusion of speech and text features, enabling the model to simultaneously grasp the correlation information of two modalities and avoid the limitations of a single modality. The speech recognition and synthesis tasks enable bidirectional conversion between speech and text, meeting the requirements for command transmission and repetition in aviation communications; the multimodal matching task ensures semantic consistency between speech and text through contrastive learning, technically avoiding the risk of discrepancies between spoken and written language.
[0147] The recognition unit is used in speech recognition tasks to measure the difference between the text transcription results output by the model and the preset real labels based on the CTC loss function or the attention loss function, and to establish the first loss function. Among them, the CTC loss function (Connectionist Temporal Classification Loss) is a loss function specifically used for temporal sequence modeling, and is widely used in scenarios such as speech recognition and handwriting recognition where the lengths of the input and output sequences are inconsistent.
[0148] The synthesis unit is used in speech synthesis tasks to calculate the difference between the speech signal generated by the model and the preset real speech signal based on the mean square error loss function or the spectral convergence loss function, and to establish a second loss function.
[0149] The matching unit is used in multimodal matching tasks to narrow the distance between speech and text features with the same semantic information based on the contrastive loss function, while widening the distance between features with different semantic information, and establishing a third loss function. It should be noted that this setting is based on the speech and text matching task. In practical applications, it can be adjusted and set according to other specific task types (such as speech and image matching, text and image matching, etc.).
[0150] The synthesis unit is used to establish a comprehensive loss function based on the first loss function, the second loss function, and the third loss function.
[0151] The integrated unit is also used to construct sample pairs based on multimodal data using a contrastive learning method; among them, positive sample pairs consist of speech features and text features carrying the same semantic information, while negative sample pairs consist of a combination of features belonging to different semantic information.
[0152] The synthesis unit also employs the Adam optimization algorithm and its variants, using the comprehensive loss function as the optimization objective, to conduct initial model training in conjunction with multimodal data and constructed sample pairs. During training, a learning rate decay strategy is simultaneously introduced, gradually reducing the learning rate as the number of training epochs increases. This allows the model to more accurately capture subtle features hidden in the data in later iterations, further improving the model's learning accuracy and generalization ability, thus obtaining a large multimodal model. Furthermore, to prevent overfitting and enhance generalization ability, multiple regularization strategies are employed simultaneously: firstly, L2 regularization, which limits the parameter size by adding the L2 norm term of the model parameters to the loss function to avoid excessive model complexity; secondly, Dropout technology, which randomly discards some neuron outputs during training, forcing the model to rely on distributed feature learning and improving robustness against interference; and thirdly, data augmentation techniques, which add noise, adjust speech rate and pitch, etc., to the speech data in the multimodal data, and perform synonym replacement and sentence restructuring operations on the text data, thereby expanding the diversity of the training data and strengthening the model's adaptability to variant samples. Among them, the Adam (Adaptive Moment Estimation) optimization algorithm is a deep learning optimization algorithm that combines momentum and adaptive learning rate characteristics, and is widely used in neural network training.
[0153] The integration unit is also used to conduct systematic evaluation and validation of the model periodically after the multimodal large model is built. Using validation set data, it comprehensively tests the model's performance metrics on core tasks such as speech enhancement, multilingual conversion, and dialogue standard calibration. Based on the evaluation results, iterative optimizations are performed on the model, such as adjusting the inter-layer connection structure and optimizing hyperparameter configurations like the learning rate decay coefficient, thereby continuously improving the model's prediction accuracy and robustness against interference in complex scenarios.
[0154] The integrated unit is also used to embed the trained and optimized multimodal large model into the software architecture of the aviation intercom headset system, ensuring seamless integration and efficient data interaction with hardware devices such as the headset's voice input module and OLED display, as well as software modules such as voice enhancement, multilingual conversion, and dialogue standard calibration. Simultaneously, compatible software interfaces and data transmission protocols are developed to construct a real-time processing and analysis workflow for multimodal data such as speech and text using the multimodal large model. Specifically, after beamforming the original audio, the model sequentially performs speech recognition to text conversion, Conv-TasNet noise reduction, and text recognition analysis and optimization, ultimately converting the standard-compliant text content into speech for output to the receiver, forming a fully automated processing loop.
[0155] In this embodiment, the recognition unit, synthesis unit, matching unit, and synthesis unit use the CTC loss function in the speech recognition task to solve the alignment ambiguity between speech and text, accurately measuring the difference between the transcription result and the real label. In the speech synthesis task, the mean squared error loss function is used. This function is sensitive to continuous signal errors and can effectively measure the acoustic feature difference between the generated speech and the real speech, ensuring the naturalness and accuracy of the synthesized speech. In the multimodal matching task, the contrastive loss function is used. By narrowing the feature distance of similar semantics, it can ensure the semantic consistency between speech and text, avoiding information misalignment. The comprehensive loss combines and optimizes these three tasks, improving the performance of a single task while taking into account other tasks, avoiding getting trapped in local optima. Furthermore, training based on multimodal data from the aviation communication environment can accurately capture the professional features of speech and text in this scenario, enabling the model to adapt to industry needs from the ground up.
[0156] The conversion unit is used to perform structured decomposition of the target language text based on a multimodal large model, and to call the preset aviation dialogue norm rule set to perform multi-dimensional comparison between the structured decomposed target language text and the corresponding entries in the preset aviation dialogue norm rule set.
[0157] The conversion unit is also used to generate corresponding correction suggestions based on the preset aviation dialogue norms if the target language text contains expressions that deviate from the preset aviation dialogue norms rule set, such as the use of non-standard aviation terminology or errors in the instruction structure.
[0158] The conversion unit is also used to convert correction suggestions into voice prompts, which are then played to the input party via an audio channel, or to display the correction suggestions in structured text on the input party's screen, thereby guiding the input party to conduct dialogue in accordance with the specifications. Specifically, it may play voice prompts such as "The terminology you are using is incorrect, it should be XX" in the headset, or it may present specific standard expressions on the OLED display, ensuring that the input party can clearly obtain and adopt the correction information. In the process of converting correction suggestions into voice prompts or text displays, multi-dimensional and multimodal fusion analysis is performed on the correction suggestions themselves, the input party's dialogue history context, and the current flight status, so as to accurately generate voice prompt content or structured text display information that is suitable for the scenario, ensuring that the correction information is not only in line with the real-time communication context, but also matches the urgency and information priority of the flight operation stage, improving the input party's acceptance of correction suggestions and execution efficiency.
[0159] This embodiment's conversion unit decomposes the target language text using structured decomposition, breaking down complex dialogue content into multiple analyzable units. These units are then compared across multiple dimensions with a pre-defined aviation dialogue specification rule set, comprehensively identifying potential problems and avoiding overlooked details. Furthermore, this comparison method ensures that corrective suggestions fully align with professional aviation dialogue standards. It provides two output modes: voice prompts and structured text display, allowing for flexible selection based on the input user's real-time status, enhancing ease of use.
[0160] For the text module 20 and correction module 30 in the embodiment, when the control tower or pilots communicate using a non-common language, the system can achieve rapid and accurate target language conversion, completely eliminating language barriers to ensure barrier-free communication between the two parties. This reduces the risk of communication errors caused by language barriers at the source, and simultaneously improves work efficiency and operational safety. At the same time, the multimodal big data model performs real-time monitoring and intelligent analysis of each round of dialogue. Once it identifies expressions that deviate from the standard, it immediately triggers the correction mechanism, pushing standardized suggestions to the user through voice prompts or text displays, guiding them to complete the dialogue interaction according to standard scripts, thereby further enhancing the accuracy and professional standardization of information transmission.
[0161] In one embodiment, the communication module 40 specifically comprises:
[0162] Obtain the final speech signal after the input party corrects their expression based on the correction suggestions;
[0163] Based on the above processing methods, the final voice signal undergoes further voice processing, language conversion, and standard calibration to ensure it strictly conforms to the expression requirements of the preset aviation dialogue specification rule set. On this basis, instructions presented in the target language are generated from the processed final voice signal. This not only ensures clear and intelligible speech and a speech rate within the standard range for aviation communication, but also that the intonation conforms to industry standards, comprehensively adapting to the communication needs of aviation scenarios.
[0164] Furthermore, the instruction is simultaneously sent to the receiver's visual interaction layer, such as being displayed on a miniature OLED display on the headset; and the display interface adopts a minimalist design logic to highlight key instruction content such as altitude parameters and runway number, so that the receiver can quickly complete visual confirmation.
[0165] It should be noted that in this second embodiment, the input party and the receiver correspond to the control tower and the pilot, respectively. The two form a two-way communication mechanism in which the tower can act as the input party to send instructions and as the receiver to receive feedback from the pilot, and vice versa, thereby realizing a closed loop of information interaction.
[0166] For examples of this application, please refer to [link / reference]. Figure 3-5 ;
[0167] Figure 3 This is a multilingual conversion flowchart provided in this application embodiment, illustrating the voice data conversion process of this embodiment two, including the logical relationships and execution steps of each stage from data input to data output.
[0168] Figure 4 This is a flowchart of the dialogue standard calibration process provided in this application embodiment, which shows the dialogue calibration process of this embodiment two, including the logical relationships and execution steps of each stage from data input to specification output.
[0169] Figure 5 This is a system architecture diagram provided in the embodiments of this application. Specifically, it takes "obtaining tower instructions and outputting instructions to the pilot" as the application scenario, and clearly shows the overall system architecture involved in the process in Embodiment 2.
[0170] Overall, this embodiment has the following beneficial effects:
[0171] This application converts the initial speech signal from the input party into text, avoiding interference from cabin noise and improving readability. The initial model is configured with three tasks: speech recognition, speech synthesis, and multimodal matching, corresponding to speech-to-text, text-to-speech, and speech-to-text alignment, respectively, and can collaboratively address issues such as accent discrepancies and terminology misuse. The comprehensive loss function optimizes the overall model performance by quantifying the differences between the output of each task and real data, ensuring recognition and conversion accuracy in complex scenarios. Furthermore, multi-task training allows the model to simultaneously master multi-dimensional features of speech, text, and context, adapting to complex communication scenarios. Based on this, correction suggestions are generated for the target language text using a large multimodal model, directly standardizing non-standard expressions and reducing repetitive communication due to misunderstandings, thus improving instruction accuracy and shortening the time required for a single communication. The two-way error correction mechanism, with the control tower and pilot as input parties, enables end-to-end standardization verification, further ensuring the accuracy and efficiency of air-to-ground communication. In addition, presenting instructions to the receiver in the target language overcomes language barriers.
[0172] In summary, this application, supported by multimodal technology, integrates key technologies such as speech enhancement processing algorithms and dual-channel feedback mechanisms. It also combines comprehensive processing capabilities including speech enhancement, multilingual translation, and language output standardization. This effectively reduces communication costs between air traffic controllers and pilots, significantly improving communication efficiency. Furthermore, it accurately resolves issues such as non-standard expressions and pronunciation differences across regions, achieving unified and standardized instruction output through standardized processing. This ensures pilots receive accurate flight instructions promptly even in complex linguistic environments, strengthening flight safety at the language interaction level and comprehensively guaranteeing flight safety and operational reliability.
[0173] Example 3:
[0174] This application provides a computer-readable storage medium, which includes a stored computer program, wherein the computer program, when running, controls the device where the computer-readable storage medium is located to execute the aforementioned aviation speech transcription method based on multimodal integration.
[0175] The aforementioned multimodal integrated aviation speech transcription method, when implemented as a software functional unit and used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0176] The above are preferred embodiments of the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for transcribing aviation speech based on multimodal integration, characterized in that, include: Acquire the initial speech signal from the input party; Convert the initial speech signal into target language text; Correction suggestions are generated for the target language text based on a multimodal large model and output to the input. The multimodal large model is obtained by training an initial model using a comprehensive loss function. The initial model is established by setting task heads including speech recognition, speech synthesis, and multimodal matching tasks based on the multimodal features of speech. The comprehensive loss function is established by quantifying the difference between generated data and real data in each task of the task head based on the loss function. The system acquires the final speech signal after the input party has corrected its expression according to the correction suggestions, generates instructions in the target language based on the final speech signal, and sends them to the receiver; wherein the input party and the receiver are the control tower and the pilot, respectively, forming a two-way communication mechanism.
2. The aviation speech transcription method based on multimodal integration as described in claim 1, characterized in that, The initial model is established by setting task heads based on the multimodal features of speech, including speech recognition, speech synthesis, and multimodal matching tasks. Specifically: Acquire multimodal data in an aviation communication environment; The speech data in the multimodal data is feature-encoded using a convolutional neural network to obtain a speech feature sequence; The text data in the multimodal data is feature-encoded using a bidirectional long short-term memory network to obtain a text feature sequence; The speech feature sequence and the text feature sequence are weighted and fused to obtain multimodal features; The task head, which includes a speech recognition task, a speech synthesis task, and a multimodal matching task, is set according to the multimodal features. The initial model is established based on the task head. The speech recognition task converts the feature sequence into a text sequence through a decoder. The speech synthesis task generates the corresponding speech waveform based on the text features through a vocoder. The multimodal matching task outputs the matching probability that the speech and text have the same semantics through contrastive learning.
3. The aviation speech transcription method based on multimodal integration as described in claim 1, characterized in that, The comprehensive loss function is established by quantifying the difference between generated data and real data based on the loss function in each task of the task header, specifically as follows: In the speech recognition task, a first loss function is established based on the CTC loss function to measure the difference between the text transcription results output by the model and the preset real labels; In the speech synthesis task, a second loss function is established by calculating the difference between the speech signal generated by the model and the preset real speech signal based on the mean square error loss function. In the multimodal matching task, a third loss function is established to narrow the distance between speech and text features with the same semantic information based on the contrastive loss function; The comprehensive loss function is established based on the first loss function, the second loss function, and the third loss function.
4. The aviation speech transcription method based on multimodal integration as described in claim 1, characterized in that, Converting the initial speech signal into target language text specifically involves: The initial speech signal is subjected to feature extraction and encoding operations using a one-dimensional convolutional neural network to obtain processed audio features. The processed audio features are then masked to generate mask information; wherein the mask information is used to distinguish between speech and noise. The audio features are decoded based on the mask information to obtain the enhanced speech signal; The enhanced speech signal is converted into text content corresponding to the target language to obtain the target language text.
5. The aviation speech transcription method based on multimodal integration as described in claim 4, characterized in that, The enhanced speech signal is converted into text content corresponding to the target language to obtain the target language text, specifically as follows: The enhanced speech signal is converted into the corresponding source language text to obtain the first language text; The first language text is preprocessed according to the multimodal large model to obtain a second language text that conforms to the preset aviation dialogue standard rule set; wherein, the preprocessing includes locating and removing non-semantic characters introduced during speech transcription based on the noise feature information in the first language text, and mapping the spoken expression text in the first language text to terminology expression text that conforms to the preset aviation dialogue standard rule set. The second language text is converted into the text content corresponding to the target language to obtain the target language text.
6. The aviation speech transcription method based on multimodal integration as described in claim 1, characterized in that, Based on a multimodal large model, correction suggestions are generated for the target language text and output to the input, specifically as follows: The target language text is structurally decomposed based on the multimodal large model, and a preset aviation dialogue norm rule set is invoked to compare the structurally decomposed target language text with the corresponding entries in the preset aviation dialogue norm rule set in multiple dimensions. If the target language text contains expressions that deviate from the preset aviation dialogue specification rule set, then corresponding correction suggestions are generated according to the preset aviation dialogue specification rule set; The correction suggestions can be converted into voice prompts and played to the input device via an audio channel, or the correction suggestions can be displayed on the input device's screen in structured text form.
7. An aviation speech transcription device based on multimodal integration, characterized in that, It includes a signal module, a text module, a correction module, and a communication module; The signal module is used to acquire the initial voice signal from the input party. The text module is used to convert the initial speech signal into target language text; The correction module is used to generate correction suggestions for the target language text based on a multimodal large model and output them to the input. The multimodal large model is obtained by training an initial model using a comprehensive loss function. The initial model is established by setting task heads including speech recognition, speech synthesis, and multimodal matching tasks based on the multimodal features of speech. The comprehensive loss function is established by quantifying the difference between generated data and real data in each task of the task head based on the loss function. The communication module is used to acquire the final speech signal after the input party has corrected its expression according to the correction suggestions, generate instructions presented in the target language based on the final speech signal, and send them to the receiver; wherein the input party and the receiver are the control tower and the pilot, respectively, forming a two-way communication mechanism.
8. The aviation speech-to-text device based on multimodal integration as described in claim 7, characterized in that, The correction module includes a data unit, a speech unit, a text unit, a fusion unit, and a model unit; The data unit is used to acquire multimodal data in an aviation communication environment; The speech unit is used to perform feature encoding on the speech data in the multimodal data according to the convolutional neural network to obtain a speech feature sequence; The text unit is used to encode the text data in the multimodal data according to the bidirectional long short-term memory network to obtain a text feature sequence; The fusion unit is used to perform weighted fusion of the speech feature sequence and the text feature sequence to obtain multimodal features; The model unit is used to set the task head, which includes a speech recognition task, a speech synthesis task, and a multimodal matching task, according to the multimodal features, and to establish the initial model according to the task head; wherein, the speech recognition task converts the feature sequence into a text sequence through a decoder, the speech synthesis task generates the corresponding speech waveform according to the text features through a vocoder, and the multimodal matching task outputs the matching probability that the speech and text have the same semantics through contrastive learning.
9. The aviation speech-to-text device based on multimodal integration as described in claim 8, characterized in that, The correction module includes an identification unit, a synthesis unit, a matching unit, and a comprehensive unit; The recognition unit is used to establish a first loss function in the speech recognition task by measuring the difference between the text transcription result output by the model and the preset real label according to the CTC loss function. The synthesis unit is used to calculate the difference between the speech signal generated by the model and the preset real speech signal according to the mean square error loss function in the speech synthesis task, and establish a second loss function. The matching unit is used to establish a third loss function in the multimodal matching task by narrowing the distance between speech and text features with the same semantic information based on the contrast loss function. The synthesis unit is used to establish the synthesis loss function based on the first loss function, the second loss function, and the third loss function.
10. A storage medium, characterized in that, The storage medium stores a computer program, which is called and executed by a computer to implement an aviation speech transcription method based on multimodal integration as described in any one of claims 1 to 6.