ASR Model Runtime Training via Pilot Edits
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Automatic Speech Recognition (ASR) models used in aircraft systems for transcribing cockpit communications are not 100% accurate due to noisy input from VHF/HF channels, requiring large amounts of manually transcribed data for training, which is labor-intensive and inefficient, especially for domain-specific accuracy in ATC-pilot communications.
Innovation Solution
A system and method that enables training of the ASR model during runtime by displaying speech-to-text samples to users for feedback, allowing them to correct errors and update the model without offline training or sophisticated learning techniques, using a background processor that operates within the transcription system to improve accuracy without interfering with ongoing transcription processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large amounts of manually transcribed data are used for training the ASR model, then the accuracy of speech-to-text conversion is improved, but the labor intensity and time requirements increase significantly
Solution Approach 1:
The system enables the ASR model to self-improve by automatically utilizing transcription corrections made during normal operational use. The background processor collects correction data from pilots who correct transcription errors, and this data is automatically used to retrain and update the ASR model without requiring separate manual transcription efforts or offline training sessions.
Solution Approach 2:
The system implements a feedback mechanism where transcription errors are identified and corrected during runtime, and these corrections are fed back to retrain the ASR model. The background processor continuously monitors transcription accuracy, collects correction data, and uses this feedback to iteratively improve the model's performance in real-world operational conditions.
2Measurement precision
If offline training methods are used to improve ASR model accuracy, then domain-specific accuracy is enhanced, but the system requires significant manual transcription and processing effort
Solution Approach 1:
The system eliminates the need for manual offline transcription and processing by automatically collecting correction data during normal operations. The background processor handles data collection, processing, and model retraining automatically, allowing the system to maintain and improve domain-specific accuracy without requiring manual intervention for data preparation.
Solution Approach 2:
The ASR model training process transitions from periodic offline batch processing to continuous online learning. The background processor continuously collects correction data and retrains the model in real-time during operational periods, ensuring the model continuously improves without interrupting service or requiring separate training phases.
3Measurement precision
If the ASR model is updated during runtime, then the accuracy improves continuously, but there is a risk of interfering with ongoing transcription operations
Solution Approach 1:
The system separates the transcription function from the model training function by implementing a background processor that operates independently. The main transcription operations continue uninterrupted in the foreground, while the background processor collects correction data and retrains the model in parallel, preventing interference between these two critical functions.
Solution Approach 2:
The background processor prepares model updates in advance by collecting and processing correction data before implementing changes to the ASR model. This preliminary preparation allows the system to update the model with minimal disruption to ongoing transcription operations, maintaining system stability while enabling continuous improvement.
Data Source
AI summary
Systems and methods are provided for training of an Automatic Speech Recognition (ASR) model during runtime of a transcription system, the system includes a background processor configured to operate with the transcription system to display a speech-to-text sample of an audio segment of a cockpit communication with an identifier which is converted using an ASR model wherein the background processor receives a response by a user during runtime of the transcription system and display of the speech-to-text sample and causes a change to the identifier to either a positive or negative attribute upon a determination of the correctness of a conversion process of the speech-to-text sample using the ASR model by review of a display of the content of the speech-to-text sample; and to train the ASR model based on information associated with the content of the speech-to-text sample in accordance with the response by the user.


