Method and system for training acoustic model of cockpit speech recognition using multi-level corpus data augmentation

Through the multi-level corpus data augmentation method, the augmented speech corpus data set is generated and used to train the ASR model, which solves the problem of difficulty in collecting speech corpus data in the aviation industry and achieves high-accuracy speech recognition.

CN111833850BActive Publication Date: 2025-05-27HONEYWELL INTERNATIONAL INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010194390.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-18
Filing Date
2020-03-19
Publication Date
2025-05-27
Estimated Expiration
2040-03-19

AI Technical Summary

Technical Problem

In aviation industry applications, it is difficult to collect and label voice corpus data, making it difficult for existing ASR technologies to achieve acceptable speech recognition accuracy.

Method used

By using a multi-level corpus data augmentation method, a corpus audio data set, including multiple augmented audio samples, is generated to train the ASR model. The method includes first and second stage augmentation, respectively, by enhancing and suppressing the audio components, and combining the transformed speech data with the noise data.

Benefits of technology

It achieves a relatively high speech recognition accuracy under the conditions of limited voice corpus data, reducing the time and burden of the crew training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111833850B_ABST
    Figure CN111833850B_ABST
Patent Text Reader

Abstract

The present invention is titled "Method and System for Training Cockpit Speech Recognition Acoustic Model Using Multi-Level Corpus Data Augmentation". The present invention discloses a method for initializing a device that enables a computer system including at least one processor and a system memory element to perform ASR using an acoustic speech recognition (ASR) model. The method includes obtaining, via at least one processor through a user interface, multiple speech data pronunciations of a predetermined phrase. The multiple speech data pronunciations include a first number of audio samples of the actually pronounced speech data, and each of the multiple speech data pronunciations includes one of the audio samples, and the audio samples include audio components. The method further includes performing multiple augmentations on the multiple speech data pronunciations of the predetermined phrase to generate a corpus audio data set, the corpus audio data set including a first number of audio samples and a second number of audio samples, and the second number of audio samples including an augmented version of the first number of audio samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to acoustic speech recognition methods and systems. More specifically, the present disclosure relates to methods and systems for training an acoustic model for cockpit speech recognition using multi-level corpus data augmentation. Background Art

[0002] In modern aircraft, advances in sensor and information processing technologies have resulted in a significant increase in the amount of information available to pilots. While this generally enables pilots to have better situational awareness at any given time, it often requires pilots to scan information from several sources in order to obtain that situational awareness. Additionally, due to the increased complexity of modern aircraft, pilots may be required to control more aircraft systems and subsystems than were present in technologically less complex aircraft in the past.

[0003] During aircraft operation, pilots are required to accurately determine and maintain a continuous perception of various elements of the current aircraft state, such as speed, altitude, position, flight direction, external atmospheric conditions, cabin conditions, fuel state, and rates of change of various parameters, as well as numerous other elements. Additionally, it is particularly important to ensure that the aircraft operates normally within various parameter limits during takeoff and landing, and that external conditions are favorable for takeoff or landing maneuvers. However, generally speaking, given the number of parameters that pilots need to accurately determine and monitor during various phases of aircraft operation, pilots may have only a very limited amount of time to make important decisions regarding aircraft control. Additionally, pilots may often be required to remove one hand from the control instruments and shift their attention from the task at hand to manipulating physical components of the user interface (e.g., keys, dials, buttons, control levers, etc.) in order to change aircraft operation based on information associated with the monitored parameters. Monitoring and controlling an aircraft can sometimes place considerable stress on pilots.

[0004] One method / system developed in recent years to assist pilots in maintaining situational awareness and reducing the manipulation of physical components of the user interface is acoustic speech recognition (ASR). The ASR method / system receives voice input from a pilot or air traffic controller and makes appropriate changes to the aircraft systems, which would otherwise require pilot input. For example, the ASR method / system may be able to receive voice input from an air traffic controller (sent to the aircraft via radio) that indicates a request to change radio frequency, altitude, heading, speed, or some other aircraft operating parameter, and can identify and automatically enter it at the appropriate system of the aircraft, thus relieving the pilot of this burden. In another scenario, the ASR method / system may be able to receive voice input from a pilot that indicates a command to change radio frequency, altitude, heading, speed, or some other aircraft operating parameter, and can identify and automatically enter it at the appropriate system of the aircraft, thus relieving the pilot of this burden.

[0005] One challenge with ASR technology is to achieve an acceptable level of speech recognition accuracy to avoid incorrect input to the aircraft systems. In prior art ASR methods / systems, the acceptable level of speech recognition accuracy is based on model training with a large amount of "speech corpus" data. As used herein, the term "speech corpus" data refers to the body of voice recordings used to train an ASR system. However, in aviation industry applications, it is difficult to collect and label speech corpus data due to the many different speakers (any given aircraft is typically flown by many different crew members) and the surrounding sound environment (engine sounds and other sounds during flight can distort the speech corpus data).

[0006] Based on the above, it is desirable to provide an aircraft cockpit acoustic speech recognition method and system that achieves a relatively high level of accuracy using limited speech corpus data. Additionally, other desirable features and characteristics of the present disclosure will become apparent in light of the following detailed description and the appended claims, taken in conjunction with the accompanying drawings, the summary of the invention, the technical field, and this background art. Summary of the Invention

[0007] Generally speaking, the present disclosure relates to methods and systems for improving acoustic speech recognition. According to one exemplary embodiment, a method for initializing a device for performing acoustic speech recognition (ASR) using an ASR model by a computer system including at least one processor and a system memory element is disclosed. The method includes obtaining, via a user interface by at least one processor, a plurality of speech data pronunciations of a predetermined phrase. The plurality of speech data pronunciations includes a first number of audio samples of the actually pronounced speech data, and each of the plurality of speech data pronunciations includes one of the audio samples, and the audio samples include audio components. The method further includes performing multiple augmentations on the plurality of speech data pronunciations of the predetermined phrase to generate a corpus audio data set, the corpus audio data set including the first number of audio samples and a second number of audio samples, the second number of audio samples including augmented versions of the first number of audio samples. Performing multiple augmentations includes performing a first-level augmentation by processing each of the plurality of speech data pronunciations to enhance a first subset of the audio components and suppress a second subset of the audio components to generate transformed speech data pronunciations including a plurality of speech transformations. Performing multiple augmentations further includes performing a second-level augmentation by processing the transformed speech data pronunciations. Performing the second-level augmentation includes combining, by at least one processor, the transformed speech data pronunciations with noise-based audio data to generate combined speech data pronunciations, and adjusting, by at least one processor, the level of the noise-based audio data for each of the combined speech data pronunciations to generate a corpus audio data set including various noise levels. Each audio sample of the corpus audio data set includes one of the plurality of speech transformations and one of the various noise levels. The method further includes training, by at least one processor, the ASR model using the corpus audio data set to perform ASR.

[0008] According to another exemplary embodiment, a computer system that uses an acoustic speech recognition (ASR) model to perform ASR includes a system memory element, a user interface, and at least one processor. The at least one processor is configured to obtain, via the user interface, a plurality of speech data pronunciations of a predetermined phrase. The plurality of speech data pronunciations includes a first number of audio samples of the actually pronounced speech data, and each of the plurality of speech data pronunciations includes one of the audio samples, and the audio samples include audio components. The at least one processor is further configured to perform multiple augmentations on the plurality of speech data pronunciations of the predetermined phrase to generate a corpus audio data set, the corpus audio data set including the first number of audio samples and a second number of audio samples, the second number of audio samples including augmented versions of the first number of audio samples. The at least one processor performs the multiple augmentations by performing a first-level augmentation and a second-level augmentation, the first-level augmentation including processing each of the plurality of speech data pronunciations to enhance a first subset of the audio components and suppress a second subset of the audio components to generate transformed speech data pronunciations including a plurality of speech transformations. The second-level augmentation includes processing the transformed speech data pronunciations by combining the transformed speech data pronunciations with noise-based audio data to generate combined speech data pronunciations and adjusting the level of the noise-based audio data for each of the combined speech data pronunciations to generate a corpus audio data set including various noise levels. Each audio sample of the corpus audio data set includes one of the plurality of speech transformations and one of the various noise levels. The at least one processor is further configured to use the corpus audio data set to train the ASR model to perform ASR.

[0009] The present invention content is provided to describe selected concepts in a simplified form, and these selected concepts will be further described in the detailed description according to various embodiments covering the concepts described in the invention content. The present invention content is not intended to identify the key or essential features of the subject matter of the present disclosure by reference to the claims or otherwise, nor is the present invention content intended to be used as an aid in determining the full scope of the disclosed subject matter, and the full scope of the disclosed subject matter is appropriately determined by reference to the various embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] A more complete understanding of the present disclosure can be obtained from the accompanying drawings, in which like reference numerals indicate like elements, and in which:

[0011] Figure 1 is a schematic block diagram of an aircraft system having an integrated acoustic speech recognition system according to an exemplary embodiment;

[0012] Figure 2 is a functional block diagram of an acoustic speech recognition system according to an exemplary embodiment;

[0013] Figure 3 is a system diagram showing the design and operation of an augmentation module of an acoustic speech recognition system according to an exemplary embodiment; Figure 2

[0014] Figure 4 is a system diagram showing the design and operation of a voice transformation processing module of an augmentation module according to an exemplary embodiment; Figure 3

[0015] Figure 5 is a system diagram showing the design and operation of a noise injection processing module of an augmentation module according to an exemplary embodiment; and Figure 3

[0016] Figure 6 is a representation of a derived speech corpus dataset generated by an augmentation module according to an exemplary embodiment; Figure 3 DETAILED DESCRIPTION

[0017] The following detailed description is merely exemplary in nature and is not intended to limit the invention or the application and uses of the invention. As used herein, the word "exemplary" means "serving as an example, instance, or illustration". Thus, any embodiment of an acoustic speech recognition system or method described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. All embodiments described herein are exemplary embodiments provided to enable persons skilled in the art to make or use the invention, without limiting the scope of the invention as defined by the claims.

[0018] Although embodiments of the subject matter of the present invention implemented in an aircraft-based speech recognition system will be described in detail below, it should be understood that the embodiments may be implemented in various other types of speech recognition systems and / or devices. These various types of systems and / or devices include, but are not limited to, systems implemented in aircraft, spacecraft, motor vehicles, watercraft, other types of vehicles and vessels, air traffic control (ATC) systems, electronic recording systems, robotic systems, hardware system control and monitoring systems (e.g., for manufacturing, energy production, mining, construction, etc.), computer system control and monitoring systems, and network control and monitoring systems, among others. Thus, the reference herein to an aircraft-based speech recognition system is not intended to be limiting. Persons skilled in the art will understand from the description herein that other embodiments may be implemented in various other types of systems.

[0019] ​​​​The subject matter of the present invention can generally be used in a variety of diverse applications that can benefit from speech recognition, particularly speech-activated operation control based on speech recognition technology. For example, the subject matter of the present invention can be used in the context of speech-activated vehicle operations (e.g., airplanes, helicopters, automobiles, or ships), speech-activated air traffic control operations, and speech-activated electronic document and / or record access processes, among others. In the following description, an exemplary application of speech-activated aircraft operations will be described in more detail. Those skilled in the art will understand based on the description herein that other embodiments can be implemented to perform other types of operations.

[0020] In the context of speech-activated aircraft operations, a speech processing method and apparatus in accordance with the subject matter of the present invention can be used to assist cockpit personnel (e.g., pilots, co-pilots, and navigators) in performing checklist-related actions, data input actions, data retrieval actions, and system control actions, among others. For example, the speech processing method of the present invention can be used to perform checklist-related actions by helping to ensure that all checklist items associated with parameter checks and tasks during takeoff and landing have been properly completed. Data input actions can include hands-free selection of radio frequencies / channels, warning level settings, navigation information specifications, etc. Data retrieval actions can include hands-free retrieval of data (e.g., navigation, operation, and task-related data). System control actions can include hands-free control of various aircraft systems and modules, as will be described in more detail subsequently.

[0021] In the presently described embodiment, a speech recognition system is operatively coupled to a host system, and speech commands and information recognized by the speech recognition system can be transmitted to the host system to control the operation of the host system, input data into the host system, and / or retrieve data from the host system. For example, and without limitation, the host system coupled with the speech recognition system can be any system selected from the following: vehicle control systems, aircraft control systems, spacecraft control systems, motor vehicle control systems, watercraft control systems, air traffic control (ATC) systems, electronic record systems, robotic systems, hardware system control and monitoring systems, computer system control and monitoring systems, network control and monitoring systems, portable systems for emergency search and rescue (e.g., for first responders at the scene), industrial monitoring and control systems (e.g., used in the context of power plants, refineries, offshore oil drilling stations, etc.), and various other types of systems.

[0022] In the presently described embodiments, the ASR system is a speaker-dependent speech recognition system. More specifically, the speaker-dependent speech recognition system implements a training process for the system user, during which user utterances (e.g., speech) are input into the system, digitized, and analyzed to create a voice profile that can be used during future interactive sessions to improve the accuracy of speech recognition. Such a speech recognition system stores the voice profile within the system itself. In the context of an aircraft, for example, a pilot in the cockpit of a first aircraft can interact with a cockpit-based speech recognition system to perform the training process, and the voice profile generated during the training process is stored in the speech recognition system and can be used during future operations of that aircraft. However, when the pilot enters the cockpit of a second aircraft, the pilot must use prior art to perform another training process to generate a voice profile for storage in the speech recognition system of the second aircraft. To this end, the embodiments of the present disclosure operate using a relatively small amount of speech corpus data, such that the training process is relatively compact and does not significantly impact other aircraft operations in terms of the time required by the flight crew. Now, various embodiments will be described in more detail in conjunction with Figures 1 to 6 More detailed descriptions of the various embodiments.

[0023] Figure 1 FIG. 6 is a schematic block diagram of an aircraft system 100 having an integrated acoustic speech recognition system according to an exemplary embodiment. The illustrated embodiment of aircraft system 100 includes, but is not limited to: at least one processing system 102; a suitable amount of data storage 104; a graphics and display system 106; a user interface 108; a control surface actuation module 110; other subsystem control modules 112; a voice input / output (I / O) interface 116; and a radio communication module 120. These elements of aircraft system 100 may be coupled together by a suitable interconnection architecture 130 that accommodates the transmission of data communication, control, or command signals within aircraft system 100 and / or the delivery of operating power. It should be understood that Figure 1 FIG. 6 is a simplified representation of aircraft system 100 that will be used for purposes of explanation and convenience of description, and Figure 1 is not intended to limit the application or scope of the subject matter in any way. In practice, as will be understood in the art, aircraft system 100 and the host aircraft will include other devices and components for providing additional functionality and features. Additionally, although Figure 1 aircraft system 100 is depicted as a single unit, any number of physically distinct hardware or device artifacts may be used to implement the individual elements and components of aircraft system 100 in a distributed manner.

[0024] The processing system 102 may be implemented or realized using one or more general-purpose processors, content addressable memories, digital signal processors, application specific integrated circuits, field programmable gate arrays, any suitable programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The processor device may be implemented as a microprocessor, controller, microcontroller, or state machine. Additionally, the processor device may be implemented as a combination of computing devices (e.g., a combination of a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a digital signal processor core, or any other such configuration). As described in more detail below, the processing system 102 may implement a speech recognition algorithm and, when operating in this context, may be considered a speech recognition system. Additionally, the processing system 102 may generate commands and may transmit the commands to various other system components via the interconnect architecture 130. Such commands may cause various system components to change their operation, provide information to the processing system 102, or perform other actions (non-limiting examples of which will be provided below).

[0025] The data store 104 may be implemented as RAM memory, flash memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. In this regard, the data store 104 may be coupled to the processing system 102 such that the processing system 102 can read information from and write information to the data store 104. In an alternative, the data store 104 may be integrated into the processing system 102. As an example, the processing system 102 and the data store 104 may reside in an ASIC. In practice, the functions or logic modules / components of the aircraft system 100 may be implemented using program code held in the data store 104. For example, the processing system 102, the graphics and display system 106, the control modules 110, 112, the voice I / O interface 116, and / or the radio communication module 120 may have associated software program components stored in the data store 104. Additionally, the data store 104 may be used to store data for supporting the implementation of speech recognition and the operation of the aircraft system 100 (e.g., speech training data), as will become apparent from the following description.

[0026] The graphics and display system 106 includes one or more displays and associated graphics processors. The processing system 102 and the graphics and display system 106 cooperate to display, present, or otherwise convey one or more graphical representations, composite displays, graphical icons, visual symbols, or images associated with the operation of the host aircraft. One embodiment of the aircraft system 100 may use existing graphics processing skills and techniques in conjunction with the graphics and display system 106. For example, the graphics and display system 106 may be suitably configured to support well-known graphics technologies such as, but not limited to, VGA, SVGA, UVGA, etc.

[0027] The user interface 108 is suitably configured to receive input from a user (e.g., a pilot) and, in response to the user input, provide appropriate command signals to the processing system 102. The user interface 108 may include any one or any combination of various known user interface devices or technologies, including but not limited to: cursor control devices such as a mouse, trackball, or joystick; a keyboard; buttons; switches; knobs; control levers; or dials. Additionally, the user interface 108 may cooperate with the graphics and display system 106 to provide a graphical user interface. Thus, a user may manipulate the user interface 108 by moving a cursor symbol presented on the display, and the user may use the keyboard to input text data and so on. For example, a user may manipulate the user interface 108 to initiate or affect the execution of a speech recognition application by the processing system 102 and the like.

[0028] In one exemplary embodiment, the radio communication module 120 is suitably configured to support data communication between the host aircraft and one or more remote systems. For example, the radio communication module 120 may be designed and configured to enable the host aircraft to communicate with an air traffic control (ATC) system 150. In this regard, the radio communication module 120 may include or support a data link subsystem that may be used to provide ATC data to the host aircraft and / or transmit information from the host aircraft to the ATC system 150, preferably in accordance with known standards and specifications. In certain specific implementations, the radio communication module 120 is also used to communicate with other aircraft in the vicinity of the host aircraft. For example, the radio communication module 120 may be configured to be compatible with Automatic Dependent Surveillance - Broadcast (ADS - B) technology, Traffic Collision Avoidance System (TCAS) technology, and / or similar technologies.

[0029] The control surface actuation module 110 includes electrical and mechanical systems configured to control the orientation of various flight control surfaces (e.g., ailerons, flaps, rudders, etc.). The processing system 102 and the control surface actuation module 110 cooperate to adjust the orientation of the flight control surfaces to affect the attitude and flight characteristics of the host aircraft. The processing system 102 may also cooperate with other subsystem control modules 112 to affect various aspects of aircraft operation. For example, without limitation, the other subsystem control modules 112 may include, but are not limited to, a landing gear actuation module, a cabin environmental control system, a throttle control system, a propulsion system, a radar system, and a data input system.

[0030] The voice I / O interface 116 is suitably configured to couple the headset 160 to the system 100, enabling the system 100 to communicate with a system user (e.g., a pilot) by voice. For example, when a system user utters words that are captured by the microphone 162 as an analog signal, the voice I / O interface 116 digitizes the analog voice signal and provides the digital voice signal to the processing system 102 for analysis by a speech recognition algorithm. Additionally, the processing system 102 and other system components (e.g., the radio communication module 120) may provide a digital voice signal to the voice I / O interface 116, which may generate an analog voice signal from the digital voice signal and provide the analog voice signal to one or more speakers 164 of the headset 160.

[0031] Figure 2 is a functional block diagram of an acoustic speech recognition system 200 according to one exemplary embodiment. The illustrated embodiment of the speech recognition system 200 includes, but is not limited to: a speech processing module 202; a speech input module 204; a command processing module 206; a speech training data update module 208; a speech training data cache 210; and an augmentation module 216. The processing system of the aircraft system (e.g., Figure 1 the processing system 102) may implement the speech processing module 202, the speech input module 204 (in combination with Figure 1 the voice I / O interface 116 of Figure 1 ), the command processing module 206, and the speech training data update module 208 by executing program code associated with these various functions. Additionally, the speech training data cache 210 may be implemented as a data store that is closely coupled to the processing system, or as a separate data store of the system (e.g.,

[0032] First, during the training process, speech corpus data 270 is provided to the speech training data cache of system 200. The training process generates and stores initial speech training data (e.g., one or more initial speech profiles) for an individual (e.g., a pilot or other system user). The initial speech training data can be generated by a speech recognition system within the cockpit or by a separate training system configured to generate the initial speech training data. According to one embodiment, the training system can be configured to generate different speech profiles for different ambient noise conditions (e.g., a quiet environment, a low engine noise environment, and / or a high engine noise environment). Thus, system 200 uses this training process to obtain multiple speech data pronunciations of a predetermined phrase. These audio samples include the actually pronounced speech data (i.e., the data that only defines the spoken words) and other audio components present in the surrounding environment, as initially pointed out above.

[0033] The speech processing module 202 executes a speech recognition algorithm that uses the speech corpus data stored in the speech training data cache to identify one or more recognized terms from the digital speech signal 280. The digital speech signal 280 is generated in response to a spoken utterance by the system user (e.g., by Figure 1 microphone 162 and the speech I / O interface 116), and the speech input module 204 is configured to receive the speech signal 280 and transmit the speech signal to the speech processing module 202. In one embodiment, the speech input module 204 can cache the received speech data within the speech signal 280 for later use by the speech processing module 202.

[0034] The speech processing module 202 can execute any type of speech recognition algorithm that uses the speech corpus data (i.e., the training data) in conjunction with speech recognition. For example, and without limitation, the speech processing module 202 can execute a speech recognition algorithm that uses a Hidden Markov Model (HMM) to model sequential speech data and identify patterns therein, and uses the speech training data (e.g., speech profile data) to train the system. An HMM is a statistical model that outputs a sequence or number of symbols. For example, an HMM can periodically output a sequence of n-dimensional real-valued vectors (e.g., cepstral coefficients). Each word (or phoneme) has a different output distribution, and an HMM for a sequence of words or phonemes can be established by concatenating independently trained HMMs for individual words and phonemes. When a new utterance is provided to the speech recognition algorithm, speech decoding can use the Viterbi algorithm to find the best path. Alternatively, the speech recognition algorithm can use dynamic time warping, artificial neural network techniques, and Bayesian networks, or other speech recognition techniques. The present disclosure should not be considered limited to any particular speech recognition algorithm.

[0035] According to one embodiment, a speech recognition algorithm generates recognized terms from a known vocabulary term set that is application-specific (e.g., for controlling an aircraft) and is also known to the system user. The known vocabulary terms can include terms associated with typical commands that the system user can issue (e.g., "Change the radio frequency to 127.7 megahertz", "Lower the flaps to 15 degrees", etc.). The speech processing module 202 transmits the recognized vocabulary terms to the command processing module 206, which is configured to determine a system response based on the command formed by the recognized terms and generate a control signal 290 to cause the aircraft system to implement the system response. In various embodiments, the command processing module 206 is configured to generate a control signal 290 to affect the operation of one or more aircraft subsystems, the one or more aircraft subsystems being selected from a group consisting of: a radio communication module (e.g., Figure 1 module 120), a graphics and display system (e.g., Figure 1 system 106), a control surface actuation module (e.g., Figure 1 module 110), a landing gear actuation module, a cabin environmental control system, a throttle control system, a propulsion system, a radar system, a data input system, and other types of aircraft subsystems.

[0036] According to one embodiment, the command processing module 206 implements an appropriate system response (i.e., generates an appropriate control signal 290) by executing applications associated with various known commands. More specifically, the command processing module 206 can map the recognized speech commands received from the speech processing module 202 to specific application actions (e.g., actions such as storing data, retrieving data, or controlling a component of the aircraft, etc.) using command / application mapping data 222. The command processing module 206 can also transmit information about the action to an application programming interface (API) 224 associated with the host system component, and the API is configured to execute the action.

[0037] For example, for an identified command related to cockpit operation, the command processing module 206 may map the command to a cockpit operation application action and may transmit information about the action to the API 224, and then the API may initiate an appropriate cockpit operation application, which is configured to initiate one to N different types of actions associated with cockpit operation. The identified commands related to cockpit operation may include, for example but without limitation: i) checklist-related commands (e.g., commands associated with ensuring that all parameter checks and tasks associated with the takeoff / landing checklist have been completed); ii) data input-related commands (e.g., commands for setting radio frequencies, selecting channels, setting warning levels (e.g., low fuel level, etc.), and other commands); iii) commands associated with controlling a multifunctional display (e.g., a radar display, a jammer display, etc.); and iv) data retrieval-related commands (e.g., retrieving data associated with tasks, speed, altitude, attitude, position, flight direction, landing approach angle, external atmospheric conditions, cabin conditions, fuel status, and rates of change of various parameters). Some cockpit operation application actions mapped to the command (and executed) may include providing human-perceivable information (e.g., data) via the display system and / or the audio system. A response generator associated with the cockpit operation application may generate a response signal and provide the response signal to the display and / or audio components of the user interface (e.g., Figure 1 the graphics of and the display system 106 and / or the voice I / O interface 116). The user interface may interpret the response signal and appropriately control the display and / or audio components to generate human-perceivable information (e.g., displayed information or audio output information) corresponding to the interpreted response signal.

[0038] As another example, for an identified command of "change the radio frequency to 127.7 megahertz", the command processing module 206 may map the command to a radio control application action and may transmit information about the action to the API 224, and then the API may initiate an appropriate radio control application. The execution of the radio control application may cause the generation of a response signal to the radio communication module (e.g., Figure 1 the radio communication module 120), thereby causing the radio communication module to switch the frequency to 127.7 megahertz. Similarly, the command processing module 206 may initiate an application that causes the graphics and display system (e.g., Figure 1 the system 106) to change the displayed information, an application that causes the control surface actuation module (e.g., Figure 1 the module 110) to change the configuration of one or more flight control surfaces, and / or an action that causes other subsystem control modules to change their operations (e.g., lower / retract the landing gear, change the cabin pressure or temperature, contact the flight attendants, etc.).

[0039] According to one embodiment, the speech training data update module 208 is configured to generate updated speech training data and metadata based on the digital speech signal 280 and provide the updated speech training data and metadata to the speech training data cache 210. In one embodiment, the updated speech training data is generated and provided when the association between the current speech sample (e.g., a speech sample obtained from the user after system startup) and the previously acquired speech profile is insufficient. More specifically, the speech training data update module 208 may enter an update mode and generate a new speech profile that reflects the new speech training data. Then the speech training data update module 208 automatically updates the speech training information in the speech training data cache 210. Thus, the speech training data update module 208 has the ability to train speech profiles and add the speech profiles to the set of existing profiles of the user. The speech training data update module 208 may use any of a variety of standard learning machines or learning system concepts (e.g., a neural network-based training process) to provide the ability to generate the updated speech training data.

[0040] According to one embodiment, the acoustic speech recognition system 200 of the present disclosure further includes an augmentation module 216. The augmentation module 216 interfaces with the speech training data cache and augments the speech corpus data included therein. Thus, the augmentation module provides a system function related to allowing the speech processing module to accurately recognize the spoken words (signal 280) using a relatively small amount of speech corpus data.

[0041] Figure 3 FIG. is a system diagram that provides more details regarding the design and operation of the augmentation module 216. As previously noted, the initial speech training data stored in the speech training data cache 210 includes one or more audio samples, and the one or more audio samples include the actually pronounced speech data (i.e., the data that only defines the spoken words; the speech corpus data 304) and other audio components present in the surrounding environment (the cockpit noise profile data 302). The augmentation module 216 processes each of these data 302, 304 in different ways (at modules 312 and 314 respectively, as will be described in more detail below), and then recombines them in various combinations to generate the augmented (derived) speech corpus set data 316.

[0042] Within the augmentation module 216 ( Figure 2 and Figure 3 ), and further referring to Figure 4, initial training speech corpus data 304 is provided to a speech transformation processing module 312, which executes a "speech random" transformation algorithm 402. Algorithm 402 is called a "speech random" algorithm because, for each instance of a given frequency component (i.e., range) in a speech data utterance, a random determination is made as to whether such frequency component will be transformed (via enhancement or suppression). Thus, for a frequency component to be enhanced, some instances of that frequency component in the speech utterance will be randomly enhanced and some instances will not be enhanced. Similarly, for a frequency component to be enhanced, some instances of that frequency component in the speech utterance will be randomly suppressed and some instances will not be suppressed. Whether a frequency component is a frequency component to be enhanced or suppressed is a predetermined aspect of the algorithm and may vary from implementation to implementation. When a speech utterance has a speech random transformation applied to it by module 312, the speech utterance sounds like different speakers (i.e., it does not sound like the speaker who produced the utterance). Thus, the output of module 312 is speech corpus data that includes speech data utterances that appear to be spoken by multiple different speakers, in Figure 4 304 as a corpus (410-A, 410-B, 410-C, ... 410-X) having transformed speech A to X. Thus, the speech transformation processing module 312 performs a first level of augmentation by processing each speech data utterance in the data 304 to enhance a first subset of acoustic components and suppress a second subset of acoustic components, thereby generating a transformed speech data utterance including a plurality of speech transformations.

[0043] Further in the augmentation module 216 ( Figure 2 and Figure 3 ) and further reference Figure 5, the voice transformations 410-A through 410-X are provided to the noise injection processing module 314. The noise injection processing module 314 also receives the previously described cockpit noise profile data 302 from the cache 210, as well as noise level control data 318. The noise level control data 318 is a predetermined factor by which the noise level (e.g., decibel level) of the cockpit noise profile data 302 is to be adjusted (such as reduced), and can vary depending on the implementation. Thus, at module 314, the cockpit noise profile data 302 is received and adjusted according to various factors determined by the noise level control data 318, and then each such adjusted cockpit noise profile data is combined with each of the voice transformations 410-A through 410-X to generate corpus data (510-1) with "Level 1" noise, corpus data (510-2) with "Level 2" noise, corpus data (510-3) with "Level 3" noise, up to corpus data (510-N) with "Level N" noise. Thus, the noise injection processing module 314 performs a second level of augmentation by processing the transformed voice data pronunciations ( Figure 4 ; 410-A through 410-X), specifically by combining the transformed voice data pronunciations 410-A through 410-X with noise-based audio data to generate combined voice data pronunciations (corpus data 510-1 through 510-N), where the level of the noise-based cockpit profile audio data has been adjusted for each of the combined voice data pronunciations to generate a corpus audio data set including various cockpit noise levels.

[0044] At the processing module 314, each voice transformation 410-A through 410-X is combined with each cockpit noise level (from control data 318), and thus an augmented (derived) voice corpus data set 316 is generated using X×N augmented voice pronunciations. This data set 316 is represented in Figure 6 , where pronunciations 602 and 603 represent "Level 1" noise with "voice" A and B respectively, pronunciations 604 and 605 represent "Level 2" noise with "voice" A and B respectively, and pronunciation 606 represents "Level X" noise with "voice" N. Thus, the corpus data set 316 can be provided to and stored in the cache 210, as Figure 3As shown above, the speech processing module 202 can utilize the augmented data set 316 when performing its speech recognition function. The augmented data set includes X×N additional speech pronunciation entries for each initial speech pronunciation. When adopting all these additional (system-generated) reference points of the predefined training pronunciations, it is expected that the accuracy of the speech recognition performed at the module 202 will be relatively greater than the case where the module 202 only uses the initial (speaker-generated) pronunciations as references. Therefore, the amount of time required for the crew to generate additional pronunciations is avoided while still maintaining an acceptable level of speech recognition accuracy.

[0045] Although at least one exemplary embodiment has been presented in the foregoing detailed description, it should be understood that there are numerous variations. It should also be understood that one exemplary embodiment or a plurality of exemplary embodiments are merely examples and are not intended to limit the scope, applicability, or configuration of the present disclosure in any way. On the contrary, the foregoing detailed description will provide those skilled in the art with a convenient roadmap for implementing one or more exemplary embodiments. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the scope of the present disclosure as set forth in the appended claims and their legal equivalents.

Claims

1. A method for initializing a device for a computer system including at least one processor and a system memory element to perform acoustic speech recognition (ASR) using an ASR model, the method comprises: obtaining, by the at least one processor via a user interface, a plurality of speech data pronunciations of a predetermined phrase, wherein the plurality of speech data pronunciations includes a first number of audio samples of actually pronounced speech data, and wherein each of the plurality of speech data pronunciations includes one of the audio samples, the audio samples including an audio component; performing augmentation a plurality of times on the plurality of speech data pronunciations of the predetermined phrase to generate a corpus audio data set, the corpus audio data set including the first number of audio samples and a second number of audio samples, the second number of audio samples including augmented versions of the first number of audio samples, in the following manner: performing a first level of augmentation by processing each of the plurality of speech data pronunciations to enhance a first subset of the audio components and suppress a second subset of the audio components to generate transformed speech data pronunciations including a plurality of speech transformations; and performing a second level of augmentation by processing the transformed speech data pronunciations, in the following manner: combining, by the at least one processor, the transformed speech data pronunciations with noise-based audio data to generate combined speech data pronunciations; and adjusting, by the at least one processor, the level of the noise-based audio data for each of the combined speech data pronunciations to generate the corpus audio data set including various noise levels, wherein each audio sample of the corpus audio data set includes one of the plurality of speech transformations and one of the various noise levels; and training, by the at least one processor, the ASR model using the corpus audio data set to perform ASR.

2. The method according to claim 1, wherein the device is implemented in an aircraft, and wherein obtaining the plurality of speech pronunciations is performed using a headset including a microphone and a speaker communicatively coupled to the aircraft.

3. The method according to claim 1, wherein performing the first level of augmentation includes utilizing a speech random transformation algorithm that randomly selects the first subset and the second subset.

4. The method according to claim 1, wherein the first subset of the audio components includes frequency components in the same frequency range.

5. The method according to claim 1, wherein the second subset of the audio components includes frequency components in the same frequency range.

6. The method according to claim 1, wherein the device is implemented in an aircraft, and wherein adjusting the noise-based audio data includes adjusting cockpit noise profile data.

7. The method according to claim 1, wherein the device is implemented in an aircraft, and wherein the method further includes receiving crew or air traffic control voice communications and using the ASR model to automatically identify words spoken in the voice communications.

8. The method according to claim 7, further comprising automatically performing an aircraft function based on the identified spoken word.

9. The method according to claim 1, further comprising generating an updated ASR model using further multiple speech data pronunciations of the predetermined phrase subsequently received via the user interface.

10. A computer system for performing ASR using an acoustic speech recognition (ASR) model, the computer system comprising: a system memory element; a user interface; and at least one processor, wherein the at least one processor is configured to: obtain, via the user interface, multiple speech data pronunciations of a predetermined phrase, wherein the multiple speech data pronunciations include a first number of audio samples of actually pronounced speech data, and wherein each of the multiple speech data pronunciations includes one of the audio samples, the audio samples including an audio component; perform multiple augmentations on the multiple speech data pronunciations of the predetermined phrase to generate a corpus audio data set, the corpus audio data set including the first number of audio samples and a second number of audio samples, the second number of audio samples including an augmented version of the first number of audio samples, in the following manner: perform a first-level augmentation by: processing each of the multiple speech data pronunciations to enhance a first subset of the audio components and suppress a second subset of the audio components to generate transformed speech data pronunciations including multiple speech transformations; and perform a second-level augmentation by processing the transformed speech data pronunciations, in the following manner: combining the transformed speech data pronunciations with noise-based audio data to generate combined speech data pronunciations; and adjusting the level of the noise-based audio data for each of the combined speech data pronunciations to generate the corpus audio data set including various noise levels, wherein each audio sample of the corpus audio data set includes one of the multiple speech transformations and one of the various noise levels; and using the corpus audio data set to train the ASR model to perform ASR.

Citation Information

Patent Citations

  • Method and apparatus for training a voice recognition model database

    CN105580071A

  • Poem recitation evaluation method, system, terminal and storage medium

    CN107316638A