Cockpit Speech Recognition Acoustic Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Acoustic speech recognition systems in aircraft face challenges in achieving high accuracy due to the difficulty in collecting and labeling voice corpus data, especially with diverse speakers and ambient noise environments.

Innovation Solution

The method involves obtaining voice data articulations, performing multi-level augmentations by enhancing and suppressing acoustic frequency components, and combining them with noise-based audio data to generate a corpus audio data set, which is used to train the ASR model, allowing for accurate speech recognition with limited voice corpus data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a large amount of voice corpus data is collected for training, then speech recognition accuracy is improved, but data collection and labeling difficulty increases due to diverse speakers and ambient noise

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoiddata collection and labeling ease
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The patent applies parameter changes by transforming the voice corpus data through multiple augmentation techniques including frequency domain transformations (enhancing and suppressing acoustic frequency components), time-domain transformations (pitch shifting, time stretching), and adding ambient noise at different levels. These parameter transformations generate diverse training samples from limited original data, improving speech recognition accuracy without requiring extensive manual data collection

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent uses copying by creating synthetic augmented versions of the original voice corpus data. Through multi-level augmentation processes, the system generates multiple copies of training samples with varied characteristics (different noise levels, frequency modifications, pitch variations) that simulate diverse speaking conditions without requiring actual collection of additional real-world voice data

Inventive Principle:
Principle #26Copying

2Measurement precision

If more voice corpus data is collected to improve accuracy, then ASR performance is improved, but the burden on flight crews increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtraining time and crew burden
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-processing and augmenting the voice corpus data before the actual speech recognition training. The multi-level augmentation process (frequency transformations, noise addition, pitch shifting) is performed in advance to create a comprehensive training dataset, eliminating the need for flight crews to participate in extensive data collection and labeling activities during operations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates synthetic copies of training data through automated augmentation processes, replacing the need for flight crews to provide extensive real-world voice samples. The copying process generates diverse training scenarios artificially, reducing the time and effort required from flight crews while maintaining high training data quality

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10997967B2Methods and systems for cockpit speech recognition acoustic model training with multi-level corpus data augmentation
Publication Date: 2021.05.04 HONEYWELL INTERNATIONAL INC
  • US10997967B2 patent drawing
  • US10997967B2 patent drawing
  • US10997967B2 patent drawing

AI summary

A method for initializing a device for performing acoustic speech recognition (ASR) using an ASR model, by a computer system including at least one processor and a system memory element. The method includes obtaining a plurality of voice data articulations of predetermined phrases, by the at least one processor via a user interface. The plurality of voice data articulations includes a first quantity of audio samples of actual articulated voice data, and each of the plurality of voice data articulations includes one of the audio samples including acoustic frequency components. The method further includes performing a plurality of augmentations to the plurality of voice data articulations of predetermined phrases, to generate a corpus audio data set that includes the first quantity of audio samples and a second quantity of audio samples including augmented versions of the first quantity of audio samples.