Data augmentation for each generation of training acoustic models

By enhancing the training data in the frequency band energy domain during the training cycle and using multiple sets of enhancement parameters, the overfitting problem of the speech analysis system during training in noisy and echo environments is solved, thereby improving the real-world adaptability and robustness of the speech analysis system.

CN114175144BActive Publication Date: 2026-03-17DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-30
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing speech analysis systems suffer from performance degradation during training due to mismatches between real-world conditions such as noise, reverberation, and echo and training conditions. Traditional multi-style training methods suffer from overfitting due to limited enhancement parameters, which affects real-world performance.

Method used

During the training cycle, training data is augmented using the frequency band energy domain, multiple sets of augmentation parameters are used to avoid overfitting, and the robustness of the system is improved by coordinating multiple microphones to estimate the user's position.

Benefits of technology

This study achieves efficient training of the speech analysis system in noisy and echo environments, avoids overfitting, and improves the system's performance and robustness in the real world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114175144B_ABST
    Figure CN114175144B_ABST
Patent Text Reader

Abstract

In some embodiments, methods and systems for training an acoustic model, where the training includes a training loop (including at least one epoch) following a data preparation phase. During the training loop, the training data is augmented to generate augmented training data. At least some of the augmented training data is used to train the model during each epoch of the training loop. The augmented training data used during each epoch can be generated by augmenting at least some of the training data differently (e.g., using different sets of augmentation parameters to augment). In some embodiments, the augmentations are performed in the frequency domain, where the training data is organized into frequency bands. The acoustic model can be of a type used (trained for) to perform speech analysis (e.g., wake-word detection, voice activity detection, speech recognition, or speaker recognition) and / or noise suppression.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 880,117, filed July 30, 2019, and U.S. Patent Application No. 16 / 936,673, filed July 23, 2020, which are incorporated herein by reference. Some embodiments relate to the subject matter of U.S. Patent Application Publication No. 2019 / 0362711, filed May 2, 2019. Technical Field

[0003] This invention relates to systems and methods for implementing speech analysis (e.g., wake word detection, speech activity detection, speech recognition, or speaker identification) and / or noise suppression using training. Some embodiments relate to systems and methods for training acoustic models (e.g., implemented by intelligent audio devices). Background Technology

[0004] Here, we use the term "smart audio device" to refer to either a single-purpose audio device or a smart device that includes a virtual assistant (e.g., a connected virtual assistant). A single-purpose audio device is a device (e.g., a television or mobile phone) that includes or is coupled to at least one microphone (and optionally includes or is coupled to at least one speaker) and / or at least one speaker (and optionally includes or is coupled to at least one microphone), and is largely or primarily designed to perform a single purpose. While televisions can generally play (and are considered capable of playing) audio from program material, in most cases, modern televisions run some operating system on which applications, including applications for watching television, run natively. Similarly, audio input and output in mobile phones can do many things, but these are served by applications running on the phone. In this sense, a single-purpose audio device with one or more speakers and one or more microphones is often configured to run native applications and / or services to directly use one or more speakers and one or more microphones. Some single-purpose audio devices can be configured to be grouped together to enable audio playback over a range of regions or user-configured areas.

[0005] A virtual assistant (e.g., a connected virtual assistant) is a device (e.g., a smart speaker or voice assistant integrated device) that includes or is coupled to at least one microphone (and optionally includes or is coupled to at least one speaker), and that provides the capability to use multiple devices (different from the virtual assistant) for applications that, in some sense, enable the cloud or are not implemented in or on the virtual assistant itself. Virtual assistants can sometimes work together, for example, in a very discrete and conditionally defined manner. For example, in some sense, two or more virtual assistants work together, with one (i.e., the one most certain it heard the wake word) responding to that word. Connected devices can form a cluster that can be managed by a master application, which can be (or include or implement) the virtual assistant.

[0006] Here, "wakeword" is used broadly to refer to any sound (e.g., a word uttered by a person, or some other sound) in which the smart audio device is configured to wake up in response to the detection ("hearing") of a sound (using at least one microphone included in or coupled to the smart audio device, or at least one other microphone). In this context, "wake-up" means that the device enters a state of waiting (i.e., listening) for a sound command.

[0007] Here, the term "wake word detector" refers to a device configured (or software, e.g., a lightweight code snippet for configuring the device) to search (e.g., continuously) for alignment between real-time sound (e.g., speech) features and a pre-trained model. Typically, a wake word event is triggered whenever the wake word detector determines that the wake word probability (the probability of detecting a wake word) exceeds a predetermined threshold. For example, the threshold may be a predetermined threshold that is adjusted to provide a good trade-off between the false acceptance rate and the false rejection rate. After the wake word event, the device may enter a state (i.e., a "wake-up" state or an "attentional" state) in which it listens for commands and passes the received commands to a larger, more computationally intensive device (e.g., a recognizer) or system.

[0008] Orchestrated systems involving multiple smart audio devices require knowledge of the user's location in order to at least: (a) select the optimal microphone for voice pickup; and (b) emit audio from a perceived location. Existing technologies include sound source localization algorithms that select a single microphone (which captures audio indicating high wake-word confidence) and use multiple synchronized microphones to estimate the user's coordinates relative to the device.

[0009] More generally, when training audio machine learning systems (e.g., wake word detectors, speech activity detectors, speech recognition systems, speaker recognizers or other speech analysis systems and / or noise suppressors), especially those based on deep learning, it is often important to augment a clean training dataset by adding reverberation, noise, and other conditions that the system will encounter when running in the real world.

[0010] Speech analysis systems (e.g., noise suppression systems, wake word detectors, speech recognizers, and speaker (voice) recognizers) are often trained from training example corpora. For example, a speech recognizer can be trained from a large number of records of people speaking single words or phrases, as well as transcriptions or labels of what they say.

[0011] In such training systems, it is often desirable to record clean speech (e.g., in low-noise and low-reverberation environments, such as recording studios or soundproof rooms, using microphones close to the speaker's mouth), because such clean speech corpora can be collected efficiently on a large scale. However, once trained, such speech analysis systems struggle to perform well under real-world conditions that do not closely match the conditions used to collect the training set. For example, speech from a person speaking into a microphone a few meters away in a typical home or office room is usually contaminated with noise and reverberation.

[0012] In this scenario, it is common for one or more devices (e.g., smart speakers) to play music (or other sounds, such as podcasts, radio broadcasts, or telephone content) while a person is speaking. This music (or other sound) can be considered an echo and can be canceled, suppressed, or managed by an echo management system that runs before the speech analysis system. However, such an echo management system cannot perfectly remove echoes from the recorded microphone signal, and echo remnants may remain in the signal presented to the speech analysis system.

[0013] Furthermore, speech analysis systems often need to operate without a complete understanding of the microphone's frequency response and sensitivity parameters. These parameters can also change over time as the microphone ages and as the speaker moves their position within the acoustic environment.

[0014] This can lead to a situation where there is a significant mismatch between the examples presented to the speech analysis system during training and the actual audio presented to the system in the real world. These mismatches in noise, reverberation, echo, level, equalization, and other aspects of the audio signal often degrade the performance of speech analysis systems trained on clean speech. Therefore, it is often desirable to enhance clean speech training data during training by adding noise, reverberation, and / or echo, and by changing the level and / or equalization of the training data. This is commonly referred to as "multi-style training" in speech technology.

[0015] Traditional multi-style training methods often involve augmenting the PCM data during the data preparation phase before training begins to create new PCM datasets. Because augmented data must be saved to disk, storage, etc., the diversity of augmentations that can be applied before training is limited. For example, a 100GB training set augmented with 10 different sets of augmentation parameters (e.g., 10 different room acoustics) would occupy 1000GB. This limits the number of different augmentation parameters that can be chosen and often leads to overfitting of the acoustic model to a specific set of selected augmentation parameters, resulting in suboptimal performance in the real world.

[0016] Traditional multistyle training is typically accomplished by augmenting the data in the time domain (e.g., by convolution with impulse responses) before the main training loop, and often suffers from severe overfitting because the number of augmented versions that can actually be created for each training vector is limited. Summary of the Invention

[0017] In some embodiments, a method for training an acoustic model includes a data preparation phase and a training cycle following the data preparation phase, wherein the training cycle includes at least one generation. In such an embodiment, the method includes the steps of: in the data preparation phase, providing (e.g., receiving or generating) training data, wherein the training data is or includes at least one example (e.g., multiple examples) of audio data (e.g., each example of audio data is a frame sequence of audio data, and the audio data indicates at least one utterance of a user); during the training cycle, augmenting the training data to generate augmented training data; and during each generation of the training cycle, training the model using at least some of the augmented training data. For example, augmented training data used during each generation (during the training cycle) can be generated by differently augmenting (e.g., using different sets of augmentation parameters) at least some of the training data. For example, augmented training data can be generated for each generation of the training cycle (including by applying different random augmentations to a set of training data for each generation), and the model can be trained during said each generation using the augmented training data generated for each generation. In some embodiments, the enhancement is performed in the frequency band energy domain (i.e., in the frequency domain, the training data is organized into frequency bands). For example, the training data may be acoustic features (organized in frequency bands) derived from the output of one or more microphones (e.g., extracted from audio data indicating the output).

[0018] Acoustic models can be of the type used (e.g., models that have been trained to) perform speech analysis (e.g., wake word detection, speech activity detection, speech recognition, or speaker recognition) and / or noise suppression.

[0019] In some embodiments, performing augmentation during training cycles (e.g., in the frequency band energy domain) rather than during the data preparation phase allows for the more efficient use of a larger number of different augmentation parameters (e.g., extracted from multiple probability distributions) than in conventional training, and can prevent the acoustic model from overfitting to a particular set of selected augmentation parameters. Typical embodiments can be efficiently implemented in GPU-based deep learning training schemes (e.g., using GPU hardware commonly used for training speech analysis systems built on neural network models, and / or GPU hardware used in common deep learning software frameworks. Examples of such software frameworks include, but are not limited to, PyTorch, Tensorflow, or Julia), and allow for very fast training times and elimination (or at least substantial elimination) of the overfitting problem. Typically, it is not necessary to save the augmented data to disk or other storage prior to training. Some embodiments avoid the overfitting problem by allowing the selection of different sets of augmentation parameters to augment the training data used for training during each generation of training (and / or augment different subsets of the training data used for training during each generation of training).

[0020] Some embodiments of the present invention envision a system for coordinating (or orchestrating) intelligent audio devices, wherein at least one (e.g., all or some) of the devices is (or includes) a speech analysis system (e.g., a wake word detector, a speech activity detector, a speech recognition system, or a speaker identifier) ​​and / or a noise suppression system. For example, in a system (including orchestrated intelligent audio devices) that requires indication of when a wake word (issued by the user) has been heard and attention (i.e., listening) to commands from the user, training according to embodiments of the present invention can be performed to train at least one element of the system to recognize the wake word. In a system including orchestrated intelligent audio devices, multiple microphones (e.g., asynchronous microphones) may be available, each microphone being included in or coupled to at least one intelligent audio device. For example, at least some microphones may be discrete microphones (e.g., in a home appliance) that are not included in any intelligent audio device but are coupled to at least one intelligent audio device (such that their output can be captured by it). In some embodiments, each wake word detector (or each smart audio device including a wake word detector) or another subsystem of the system (e.g., a classifier) ​​is configured to estimate the user's location (e.g., which of several different regions the user is located in) by applying a classifier driven by multiple acoustic features derived from at least some microphones (e.g., asynchronous microphones). The goal may not be to estimate the user's exact location, but rather to form a robust estimate of discrete regions (e.g., in the presence of strong noise and residual echoes).

[0021] It is conceivable that the user, the smart audio device, and the microphone are in an environment where sound can propagate from the user to the microphone (e.g., the user's residence or business premises), and this environment includes predetermined areas. For example, the environment may include at least the following areas: a food preparation area; a dining area; an open area of ​​a living space; a television area of ​​a living space (including a television sofa), etc. During system operation, it is assumed that the user is physically located in one of these areas ("user area") at all times, and the user area may change from time to time.

[0022] Microphones can be asynchronous (i.e., digitally sampled using different sampling clocks) and randomly located. A user's region can be estimated using a data-driven approach that is driven by multiple high-level features, at least partially derived from at least one of a set of wake word detectors. These features (e.g., wake word confidence and reception level) typically consume very little bandwidth and can be transmitted asynchronously to a central classifier with very little network load.

[0023] Some aspects of the embodiments relate to implementing smart audio devices, and / or coordinating smart audio devices.

[0024] Aspects of the invention include systems configured (e.g., programmed) to perform any embodiment of the method or steps of the invention, and tangible, non-transitory computer-readable media (e.g., disks or other tangible storage media) that implement non-transitory storage of data, storing code for performing (e.g., executable to perform) any embodiment of the method or steps of the invention. For example, embodiments of the system of the invention may be or include a programmable general-purpose processor, digital signal processor, GPU, or microprocessor, any of which is programmed with software or firmware and / or otherwise configured to perform various operations on data, including embodiments of the method or steps of the invention. Such a general-purpose processor may be or include a computer system including input devices, memory, and a processing subsystem programmed (and / or otherwise configured) to perform embodiments of the method (or steps of the invention) in response to data asserted thereto. Some embodiments of the system of the invention may be (or) implemented as a cloud service (e.g., components of the system are located in different locations, and data is transferred between these locations, for example, via the Internet).

[0025] Naming and Terminology

[0026] Throughout this disclosure, including in the claims, the terms "loudspeaker" and "amplifier" are used synonymously to refer to any sound-generating transducer (or group of transducers) driven by a single loudspeaker feed. A typical headset includes two loudspeakers. A loudspeaker can be implemented to include multiple transducers (e.g., a woofer and a tweeter), all of which are driven by a single common loudspeaker feed (the loudspeaker feed may undergo different processing in different circuit branches coupled to the different transducers).

[0027] Throughout the disclosure, including in the claims, the expression “on” a signal or data (e.g., filtering, scaling, transforming, or applying gain to a signal or data) is used broadly to mean performing an operation directly on the signal or data, or on a processed version of the signal or data (e.g., a version of the signal that has undergone preliminary filtering or preprocessing before the operation is performed on it).

[0028] Throughout this disclosure, including in the claims, the term "system" is used broadly to refer to a device, system, or subsystem. For example, a subsystem implementing a decoder can be called a decoder system, and a system including such a subsystem (e.g., a system that generates an X output signal in response to multiple inputs, wherein the subsystem generates M inputs and receives additional XM inputs from an external source) can also be called a decoder system.

[0029] Throughout this disclosure, including in the claims, the term "processor" is used broadly to refer to a system or device that is programmable or otherwise configurable (e.g., with software or firmware) to perform operations on data (e.g., audio data). Examples of processors include field-programmable gate arrays (or other configurable integrated circuits or chipsets), digital signal processors programmed and / or otherwise configured to perform pipelined processing of audio data, graphics processing units (GPUs) configured to perform processing of audio data, programmable general-purpose processors or computers, and programmable microprocessor chips or chipsets.

[0030] Throughout the disclosure, including the claims, the terms "coupled" or "coupled" are used to indicate a direct or indirect connection. Therefore, if a first device is referred to as being coupled to a second device, the connection can be either a direct connection or an indirect connection via other devices and connections.

[0031] Throughout the disclosure, including in the claims, “audio data” means data indicating sound (e.g., speech) captured by at least one microphone, or data generated (e.g., synthesized) such that said data can be rendered as sound (e.g., speech) playback (by at least one speaker), or can be used to train a speech analysis system (e.g., a speech analysis system operating only in the frequency band energy domain). For example, audio data can be generated to serve as an alternative to data indicating sound (e.g., speech) captured by at least one microphone. Here, the expression “training data” means audio data that can be used (or intended for) training an acoustic model.

[0032] Throughout the disclosure, including the claims, the term “addition” (e.g., the step of “adding” enhancements to training data) is used broadly to refer to addition (e.g., mixing or otherwise combining) and approximate implementations of addition. Attached Figure Description

[0033] Figure 1 It is an environment diagram that includes a system containing a set of smart audio devices.

[0034] Figure 1 A is a flowchart of the traditional multi-style training process for acoustic models.

[0035] Figure 1B This is a flowchart of the multi-style training process of the acoustic model according to an embodiment of the present invention.

[0036] Figure 2 This is a schematic diagram of another environment that includes users and a system comprising a set of smart audio devices.

[0037] Figure 3 This is a block diagram of the components of a system that can be implemented according to embodiments of the present invention.

[0038] Figure 3A This is a block diagram of the components of a system that can be implemented according to embodiments of the present invention.

[0039] Figure 4 These are a set of graphs illustrating an example of fixed-spectrum smooth noise addition (enhancement) according to an embodiment of the present invention.

[0040] Figure 5 This is a graph illustrating an example of an embodiment of the invention including microphone equalization enhancement.

[0041] Figure 6 This is a flowchart of the training process according to an embodiment of the present invention, wherein the enhancement includes the addition of variable spectral semi-stationary noise.

[0042] Figure 7This is a flowchart of the training process according to an embodiment of the present invention, wherein the enhancement includes the addition of non-stationary noise.

[0043] Figure 8 This is a flowchart of the training process according to an embodiment of the present invention, wherein the enhancement implements a simplified reverberation model.

[0044] Figure 9 This is a flowchart of a method for enhancing input features (128B) and generating class label data (311-314) according to an embodiment of the present invention for training a model. The model classifies the time-frequency blocks of the enhanced features into speech, stationary noise, non-stationary noise, and reverberation categories, and can be used to train a model for noise suppression (including suppression of non-speech sounds).

[0045] Figure 10 It is to augment the training data (e.g., based on...) Figure 9 The four example graphs of the data (310) generated by the method are each generated by augmenting the same set of training data (training vectors) for use during different generations of model training. Detailed Implementation

[0046] Many embodiments of the present invention are technically possible. Those skilled in the art will understand from this disclosure how they can be implemented. Referring to the accompanying drawings, examples of embodiments of the systems and methods of the present invention will now be described.

[0047] Figure 1 This is a schematic diagram of an environment (living space) including a system comprising a set of intelligent audio devices (devices 1.1) for audio interaction, a speaker (1.3) for audio output, and controllable lights (1.2). In one example, each device 1.1 includes (and / or is coupled to) at least one microphone, such that the environment also includes the microphone, and the microphone provides the device 1.1 with a sense of where (e.g., in which area of ​​the living space) a user (1.4) who issues a wake-up command (device 1.1 is configured to recognize the sound as a wake-up sound in a specific context) is. The system (e.g., one or more of its devices 1.1) can be configured to implement embodiments of the invention. Various methods can be used to obtain information from... Figure 1 The devices together acquire information and use it to provide a location estimate of the user who issues (e.g., speaks) the wake word.

[0048] In living spaces (e.g., Figure 1 In a living space, there exists a set of natural activity zones where people will perform tasks or activities, or cross thresholds. These activity zones (areas) are places where it may be difficult to estimate the user's position (e.g., determining an uncertain location) or context. Figure 1 In the example, the key area is

[0049] 1. Kitchen sink and food preparation area (in the upper left area of ​​the living space);

[0050] 2. Refrigerator door (right side of the sink and food preparation area);

[0051] 3. Dining area (located in the lower left area of ​​the living space);

[0052] 4. Open areas of the living space (to the right of the sink and food preparation area and the dining area);

[0053] 5. TV sofa (on the right side of the open area);

[0054] 6. The television itself;

[0055] 7. Table; and

[0056] 8. Doorway or entrance passage (in the upper right area of ​​the living space).

[0057] According to some embodiments of the invention, a system that estimates where a signal (e.g., a wake word or other attention-grabbing signal) originates or is derived (e.g., an uncertain estimate of its origin) can have several determined confidence levels (or multiple assumptions) on that estimate. For example, if a user happens to be near a boundary between areas of the system environment, the uncertain estimate of the user's location can include a determined confidence level of the user's location in each area. In some conventional implementations of voice interfaces (e.g., Alexa), it is required that the voice assistant's voice be emitted from only one location at a time, which forces the system to consider the individual locations (e.g., Figure 1 The assistant's voice source can be selected individually from one of the eight speaker positions (1.1 and 1.3). However, based on simple hypothetical role-playing, it is clear (in this conventional implementation) that the selected position of the assistant's voice source (i.e., the position included in or coupled to the assistant's speaker) is likely to be a natural return response or focus for expressing attention.

[0058] Figure 2 This is a schematic diagram of another environment (109), which is an acoustic space including a user (101) uttering direct speech 102. This environment also includes a system comprising a set of intelligent audio devices (103 and 105), speakers for audio output, and a microphone. The system can be configured according to embodiments of the invention. The speech uttered by user 101 (sometimes referred to herein as the speaker) can be recognized as a wake word by one or more elements of the system.

[0059] More specifically, Figure 2 The system components include:

[0060] 102: Direct local voice (sent by user 101);

[0061] 103: Voice assistant device (coupled to multiple speakers). Device 103 is closer to user 101 than device 105, therefore device 103 is sometimes referred to as the "proximal" device, while device 105 is referred to as the "distal" device;

[0062] 104: Multiple microphones in (or coupled to) the near-side device 103;

[0063] 105: Voice assistant device (coupled to multiple speakers);

[0064] 106: Multiple microphones in (or coupled to) the remote device 105;

[0065] 107: Household appliances (such as lamps); and

[0066] 108: Multiple microphones in (or coupled to) household appliance 107. Each microphone 107 is also coupled to at least one of devices 103 or 105.

[0067] Figure 2 The system may also include at least one speech analysis subsystem (e.g., the one described below that includes classifier 207). Figure 3 The system is configured to perform speech analysis on the system's microphone output (e.g., including classifying features derived from the microphone output) (e.g., indicating the probability of a user in each of multiple regions of environment 109). For example, device 103 (or device 105) may include a speech analysis subsystem, or the speech analysis subsystem may be implemented separately from devices 103 and 105 (but coupled to devices 103 and 105).

[0068] Figure 3 This is a block diagram of the elements of a system that can be implemented according to embodiments of the present invention (e.g., wake word detection or other speech analysis processing can be implemented by training according to embodiments of the present invention). Figure 3 The system (including a region classifier) ​​is implemented in an environment with regions and includes:

[0069] 204: Distributed throughout the listening environment (e.g., Figure 2 Multiple speakers in the environment;

[0070] 201: Multichannel loudspeaker renderer, whose output is used as both loudspeaker drive signal (i.e., loudspeaker feed for driving loudspeaker 204) and echo reference;

[0071] 202: Multiple speaker reference channels (i.e., speaker feed signals output from renderer 202, which are provided to echo management subsystem 203);

[0072] 203: Multiple echo management subsystems. The reference input for subsystem 203 is all (or a subset thereof) of the speaker feeds output from renderer 202;

[0073] 203A: Multiple echo management outputs, each echo management output being output from one of the subsystems 203, and each echo management output having a decayed echo (relative to the input of the associated subsystem 203);

[0074] 205: Distributed throughout the listening environment (e.g., Figure 2 Multiple microphones in the listening environment. The microphones may include both array microphones in multiple devices and point microphones distributed throughout the listening environment. The output of microphone 205 is provided to echo management subsystem 203 (i.e., each echo management subsystem 203 captures the output of a different subset of microphone 205 (e.g., one or more microphones));

[0075] 206: Multiple wake word detectors, each taking audio output from one of the subsystems 203 as input and outputting multiple features 206A. The features 206A output from each subsystem 203 may include (but are not limited to): wake word confidence, wake word duration, and a measure of reception level. Each detector 206 may implement a model trained according to embodiments of the present invention;

[0076] 206A: Multiple features derived from (and from the output of) all wake word detectors 206;

[0077] 207: A region classifier that acquires features 206A (as input) from the wake-word detector 206 for all microphones 205 in the acoustic space. The classifier 207 can implement a model trained according to an embodiment of the invention; and

[0078] 208: Output of region classifier 207 (e.g., indicating the posterior probability of multiple regions).

[0079] The following description Figure 3 Example implementation of region classifier 207.

[0080] Let x i (n) is the i-th microphone signal in discrete time n, where i = {1...N} (i.e., microphone signal x). i (n) represents the outputs of N microphones 205. The N signals x in the echo management subsystem 203... i The processing of (n) produces a "clean" microphone signal e. i (n), where i = {1...N}, and each signal is generated in discrete time n. Clean signal e i (n), in Figure 3 The vector w, referred to as 203A, is fed into the wake word detector 206. Each wake word detector 206 generates a feature vector w. i (j), in Figure 3 This is referred to as 206A, where j = {1...J} is the index corresponding to the j-th wake word utterance. Classifier 207 will aggregate the feature set... As input.

[0081] For example, specify a set of area labels C k k = {1...K}, corresponding to areas (K distinct areas) within an environment (e.g., a room). These areas could include a sofa area, a kitchen area, a reading chair area, etc.

[0082] In some implementations, classifier 207 estimates the posterior probability p(C) of the feature set W(j) by using, for example, a Bayesian classifier. k |W(j))(and output an indicator signal). Probability p(C k |W(j)) instructs the user to perform each region C k The probability of (for the j-th utterance and the k-th region, for each region C) k (and each utterance), and is an example of the output 208 of classifier 207.

[0083] Typically, training data is collected by having users utter a wake word near a desired area (e.g., the center and outermost edge of a sofa) for each area. The word may be repeated several times. Then, the user moves to the next area and continues until all areas are covered.

[0084] An automated prompting system can be used to collect this training data. For example, during this process, the user might see the following prompts on the screen, or hear these notifications:

[0085] • Move to the sofa

[0086] • Say the wake-up phrase ten times while shaking your head.

[0087] • Move to a position between the sofa and the reading chair, and say the wake-up phrase ten times.

[0088] • "Stand in the kitchen, as if you're cooking, and then say the wake-up phrase ten times."

[0089] The model trained by classifier 207 (or another model trained according to an embodiment of the invention) may or may not be labeled. In the case of labeling, each training utterance is paired with a hard label.

[0090]

[0091] And the model is fitted to best fit the labeled training data. Without loss of generality, appropriate classification methods may include:

[0092] • Bayesian classifiers, for example, have each class distribution described by a multivariate normal distribution, a full covariance Gaussian mixture model, or a diagonal covariance Gaussian mixture model;

[0093] • Vector quantization;

[0094] • Nearest neighbor method (k-means);

[0095] • A neural network with a softmax output layer, where each output corresponds to a class;

[0096] • Support Vector Machine (SVM); and / or

[0097] • Improving technologies, such as gradient boosters (GBM)

[0098] In the unlabeled case, training the model implemented by classifier 207 (or training another model according to an embodiment of the invention) includes automatically segmenting the data into K clusters, where K may also be unknown. For example, unlabeled automatic segmentation can be performed using classic clustering techniques, such as the k-means algorithm or Gaussian mixture modeling.

[0099] To improve robustness, regularization can be applied to model training (which can be performed according to an embodiment of the method of the present invention), and the model parameters can be updated over time as new utterances are made.

[0100] The following describes other aspects of the example, in which embodiments implementing the method of the invention are used to train a model (e.g., by...). Figure 3 (Model implemented by component 207 of the system).

[0101] Example feature set (e.g., Figure 3 Feature 206A (derived from the microphone output in the region of the environment) includes features indicating the likelihood of wake-up word confidence, the average reception level over the estimated duration of the most confident wake-up word, and the maximum reception level over the duration of the most confident wake-up word. For each wake-up word utterance, the features can be normalized relative to their maximum values. Training data can be labeled, and a fully covariant Gaussian mixture model (GMM) can be trained to maximize the expected value of the training labels. The estimated region is the class with the highest posterior probability.

[0102] The above description relates to learning an acoustic region model from a set of training data collected during a collection process (e.g., a cue collection process). In this model, training time (operating in configuration mode) and runtime (operating in normal mode) can be considered as two different modes in which the system's microphone can operate. An extension of this scheme is online learning, where some or all of the acoustic region model is learned or adjusted online (i.e., during operation in normal mode).

[0103] Online learning models may include the following steps:

[0104] 1. Whenever a user says a wake word, predict which region the user is in based on the prior region mapping model (e.g., offline learning during the setup phase or online learning during the previous learning epoch).

[0105] 2. Obtain implicit or explicit feedback regarding whether the prediction is correct; and

[0106] 3. Update the region mapping model based on feedback.

[0107] Explicit techniques for obtaining feedback include:

[0108] • Use a voice user interface (UI) to ask the user if the prediction is correct. For example, you could provide the user with a voice indicating the following: "I think you are on the sofa. Please say 'yes' or 'no'."

[0109] • Use the voice UI at any time to notify the user to correct incorrect predictions. (For example, you could provide the user with a voice indicating the following: "I am now able to predict where you are when you speak to me. If I predict incorrectly, I will say, 'Amanda, I am not on the sofa. I am sitting in the reading chair.'")

[0110] • Use the voice UI to notify the user at any time that a correct prediction can be rewarded. (For example, you could provide the user with a voice indicating the following: "I am now able to predict where you are when you speak to me. If my prediction is correct, you can say something like 'Amanda, yes. I am on the couch' to help further improve my prediction.")

[0111] • This includes physical buttons or other UI elements that users can interact with to provide feedback (e.g., thumb up and / or thumb down buttons on a physical device or in a smartphone app).

[0112] The goal of predicting the acoustic region (where the user is located) could be to inform microphone selection or adaptive beamforming schemes that attempt to pick up sound more effectively from the user's acoustic region, for example, to better recognize commands following a wake word. In this case, implicit techniques for obtaining feedback on the quality of the region prediction could include:

[0113] • Penalties may result in incorrect prediction of commands following a wake word. Agents that might instruct incorrect recognition could include the user shortening the voice assistant's response to the command, for example, by issuing the opposite command, such as, "Amanda, stop!";

[0114] • Penalties result in the speech recognizer successfully recognizing a low-confidence prediction of a command. Many automatic speech recognition systems have the ability to return to the confidence level, and their results can be used for this purpose;

[0115] • Penalties prevent second-pass wake word detectors from retrospectively detecting wake word predictions with high confidence; and / or

[0116] • Enhance predictions that lead to high-confidence recognition of wake words and / or correct recognition of user commands.

[0117] Techniques for posteriorly updating region mapping models after one or more wake words have been spoken include:

[0118] • Adaptive maximum a posteriori (MAP) model of Gaussian mixture model or nearest neighbor model; and / or

[0119] • Reinforcement learning, such as reinforcement learning of neural networks, involves determining new network weights by associating appropriate “one-hot” (in the case of correct prediction) or “one-cold” (in the case of incorrect prediction) ground truth labels with the softmax output and applying online backpropagation.

[0120] Figure 3A This is a block diagram illustrating examples of components of device (5), which may be configured to perform at least some of the methods disclosed herein. In some examples, device 5 may be or may include a personal computer, a desktop computer, a graphics processing unit (GPU), or another local device configured to provide audio processing. In some examples, device 5 may be or may include a server. According to some examples, device 5 may be a client device configured to communicate with a server via a network interface. Components of device 5 may be implemented by hardware, software stored on a non-transitory medium, firmware, and / or combinations thereof. Figure 3A The types and numbers of components shown, as well as other figures disclosed herein, are illustrated by way of example only. Alternative implementations may include more, fewer, and / or different components.

[0121] Figure 3AThe device 5 includes an interface system 10 and a control system 15. The device 5 may be referred to as a system, and its components 10 and 15 may be referred to as subsystems of such a system. The interface system 10 may include one or more network interfaces, one or more interfaces between the control system 15 and the memory system, and / or one or more external device interfaces (e.g., one or more Universal Serial Bus (USB) interfaces). In some embodiments, the interface system 10 may include a user interface system. The user interface system may be configured to receive input from a user. In some embodiments, the user interface system may be configured to provide feedback to the user. For example, the user interface system may include one or more displays with corresponding touch and / or gesture detection systems. In some examples, the user interface system may include one or more microphones and / or speakers. According to some examples, the user interface system may include means for providing haptic feedback, such as motors, vibrators, etc. The control system 15 may, for example, include a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, and / or discrete hardware components. The control system 15 may also include one or more devices implementing non-transitory memory.

[0122] In some examples, device 5 can be implemented in a single device. However, in some implementations, device 5 can be implemented in more than one device. In some such implementations, the functionality of control system 15 can be included in more than one device. In some examples, device 5 can be a component of another device.

[0123] In some embodiments, device 5 is or implements a system for training an acoustic model, wherein training includes a data preparation phase and a training cycle following the data preparation phase, and wherein the training cycle includes at least one generation. In some such embodiments,

[0124] Interface system 10 is or implements a data preparation subsystem, which is coupled and configured to implement a data preparation phase, including receiving or generating training data, wherein the training data is or includes at least one example of audio data, and

[0125] The control system 15 is or implements a training subsystem coupled to a data preparation subsystem and is configured to augment training data during training cycles to generate augmented training data, and to train a model using at least some (e.g., different subsets) of the augmented training data during each generation of the training cycle (e.g., generating different subsets of the augmented training data during training cycles for use during different generations of the training cycle, each said subset being generated by differently augmenting at least some of the training data).

[0126] According to some examples, embodiments of the present invention are implemented... Figure 3 System components 201, 203, 206, and 207 can be transmitted via one or more systems (e.g., Figure 3A This is achieved through a control system 15). Similarly, elements of other embodiments of the invention (e.g., configured to implement, as referred here) Figure 1B The elements of the described method can be transmitted via one or more systems (e.g., Figure 3A The control system 15) is used to achieve this.

[0127] refer to Figure 1 A. Next, we will describe an example of a traditional multi-style training method. Figure 1 A is a flowchart (100A) of a conventional multi-style training method for training acoustic model 114A. This method can be implemented by a programmed processor or other system (e.g., Figure 3A The control system 15) is implemented, and the steps (or stages) of the method can be implemented by one or more subsystems of the system. Here, such a subsystem is sometimes referred to as a "unit", and the steps (or stages) implemented thereunder are sometimes referred to as functions.

[0128] Figure 1 Method A includes a data preparation phase 130A and a training loop (training phase) 131A executed after the data preparation phase 130A. In this method, augmentation functions (units) 103A augment the audio training data (110) during the data preparation phase 130A.

[0129] Figure 1 The elements of A include the following:

[0130] • 101A: Indication of separation between the data preparation phase 130A and the training cycle (training phase) 131A. Phases 130A and 131A can generally be considered two distinct phases of the complete training process. Each pass through training cycle 131A can be referred to as a generation;

[0131] 102A: Start of program flow (i.e., execution) Figure 1 (The start of execution of method A);

[0132] • 103A: Enhancement Function / Unit. This function (or unit) acquires training data (training set) 110 and applies enhancements to it (e.g., adding reverberation, adding stationary noise, adding non-stationary noise, and / or adding simulated echo residuals) to generate enhanced training data (enhanced training set) 112A;

[0133] • 104A: Feature Extraction Function / Unit. This function (or unit) takes enhanced training data 112A (e.g., time-domain PCM audio data) as input and extracts features 113A (e.g., Mel frequency cepstral coefficients (MFCC), "logmelspec" coefficients (logarithmic coefficients indicating the power of frequency bands spaced apart to occupy equal or substantially equal portions of the Mel spectrum, coefficients indicating the power of frequency bands spaced apart to occupy at least substantially equal portions of the logarithmic spectrum, and / or perceptual linear predictor (PLP) coefficients) for training model 114A;

[0134] • 105A: The prediction phase (or unit) of training phase 131A. In phase 105A, training features (113A) pass through the model (114A). For example, if training phase 131A is or implements the Expectation Maximization (EM) algorithm, then phase 105A (sometimes called feature 105A) can be referred to as the "E-step". When model 114A is an HMM-GMM acoustic model, an applicable variant of the EM algorithm is the Baum Welch algorithm. On the other hand, if model 114A is a neural network model, then prediction feature 105A corresponds to the forward-running network;

[0135] • 106A: Update phase of training phase 131A. In this training phase, the predicted output from phase 105A is compared with the ground truth label (e.g., 121C) using some loss criterion, and the determined loss is used to update model 114A. If training phase 131A is the EM algorithm, then phase 106A can be referred to as the "M-step". On the other hand, if training phase 131A is a neural network training process, then phase 106A can correspond to calculating the gradient of the CTC (Connectionist Temporal Classification) loss function and backpropagation;

[0136] • 107A: Apply convergence (stopping) criteria to determine whether to stop training phase 131A. Typically, training phase 131A will require running multiple iterations until it is determined that the convergence criterion is met. Examples of stopping criteria include (but are not limited to):

[0137] Run a fixed number of epochs / passes (through the loop of phase 131A); and / or

[0138] Wait until the change in training loss from one generation to another is less than the threshold.

[0139] • 108A: Stop. In this method (which can be implemented as a computer program running on a processor), once control reaches this point, training of model 114A is complete;

[0140] • 110: Training set: Training data (e.g., multiple example audio utterances) used to train the acoustic model 114A. Each audio utterance may include or contain PCM speech data 120A and some kind of label or transcription 121A (e.g., such a label could be "cat");

[0141] • 112A: Enhanced Training Set. Enhanced training set 112A is an enhanced version of training set 110 (e.g., it may include multiple enhanced versions of audio utterances from training set 110). In one example, the enhanced PCM utterance 122A (of set 112A) is an enhanced version of PCM speech data 120A, and set 112A includes (i.e., retains) the label “cat” (121B) copied from input label 121A. However, as... Figure 1 As shown in Figure A, enhancement unit 103A has generated enhanced PCM data 122A to include the following additional features:

[0142] ο123A, 123B: Examples of non-stationary noise; and

[0143] ο124A: Reverberation;

[0144] • 113A: Enhanced features corresponding to the conventional enhanced training set 112A (determined by function 104A). In this example, feature set 113A includes the Mel spectrum (127A) corresponding to PCM utterance 122A. Enhanced feature set 113A contains the following enhancements corresponding to features 123A, 123B, and 124A:

[0145] ο125A-125D: Examples of non-stationary noise (time-limited and frequency-limited); and

[0146] ο126A: Reverberation;

[0147] ·120A: PCM speech data of a sample utterance in the training set (110);

[0148] • 121A: A label (e.g., transcription) of an example utterance (corresponding to PCM speech data 120A) in the training set (110);

[0149] • 127A: Enhance the features of a utterance in feature set 113A (e.g., spectrogram or logmelspec features);

[0150] • 130A: Data Preparation Phase. This occurs only once and therefore may not be highly optimized. In typical deep learning training, the computations for this phase are performed on the central processing unit; and

[0151] 131A: Main training phase (loop). Since the phases in this loop (105A, 106A, 107A) run on the main pass / generation, these operations are typically highly optimized and run on multiple GPUs.

[0152] Next, refer to Figure 1B This describes an example of a multi-style training method according to embodiments of the present invention. The method can be implemented by a programmed processor or other system (e.g., Figure 3A The control system 15) is implemented, and the steps (or stages) of the method can be implemented by one or more subsystems of the system. Here, such a subsystem is sometimes referred to as a "unit", and the steps (or stages) implemented thereunder are sometimes referred to as functions.

[0153] Figure 1B This is a flowchart (100B) of a multi-style training method used to train acoustic model 114B. Figure 1B The method includes a data preparation phase 130B and a training loop (training phase) 131B executed after the data preparation phase 130B. Figure 1B In the method, enhancement feature 103B enhances the audio training data (features 111B generated from training data 110) during training loop 131B to generate enhanced feature 113B.

[0154] and Figure 1 In contrast to A's traditional approach, it enhances (through...) Figure 1B The function / unit 103B is performed in the feature domain and during the training cycle (phase 131B), not during the data preparation phase. Figure 1 Phase 130A of A is performed directly on the input audio training data (110). Figure 1B The elements include the following:

[0155] 102B: Program flow begins (i.e., execution begins) Figure 1B (The start of the program execution of the method);

[0156] • 101B: Indication of separation between the data preparation phase 130B and the training loop (training phase) 131B. Phases 130B and 131B can generally be considered as two distinct phases of the entire training process. During training loop 131B, training of model 114B typically occurs over a series of training iterations, and each such iteration is sometimes referred to here as a "pass" or "minibatch" through the training loop.

[0157] ·110: Training dataset. Figure 1B The training data 110 can be compared with Figure 1 The training data for A's traditional method is the same as 110;

[0158] • 111B: Unenhanced training set features (generated by feature extraction function 104A, which can be compared with...) Figure 1 Method A has the same functionality as 104A.

[0159] • 103B: Data augmentation function (or unit). This function (unit) takes feature 111B (determined in data preparation phase 130B) and applies enhancements to it, thereby generating enhanced feature 113B. Examples of enhancements will be described here, including (but not limited to) the addition of reverberation, the addition of stationary noise, the addition of non-stationary noise, and / or the addition of simulated echo residuals. Figure 1 In contrast to the traditional data augmentation unit 103A, unit 103B operates as follows:

[0160] ο In the functional domain. Therefore, in a typical implementation, it can be implemented quickly and efficiently on a GPU as part of the deep learning training process; and

[0161] Within training loop 131B (i.e., during each pass through training loop 131B or during each generation of training loop 131B). Therefore, during each generation of training, different enhancement conditions (e.g., different room / reverberation models, different noise levels, different noise spectra, different non-stationary noise or music residual patterns) can be selected for each training example in training set 110.

[0162] • 104A: Feature Extraction Function / Unit. This function (unit) takes input training data 110 (e.g., input time-domain PCM audio data) and extracts features 111B from it, which are used to enhance function (unit) 103B and to train model 114B. Examples of features include (but are not limited to) Mel frequency cepstral coefficients (MFCC), "logmelspec" (logarithmic coefficients indicating the power of frequency bands spaced apart to occupy equal or substantially equal portions of the Mel spectrum), coefficients indicating the power of frequency bands spaced apart to occupy at least substantially equal portions of the logarithmic spectrum, and / or perceptual linear predictor (PLP) coefficients;

[0163] • 105B: The prediction phase of training loop 131B, where augmented training data 113B is passed through the model being trained (114B). Phase 105B can be synchronized with phase 105A ( Figure 1A) This is performed in the same manner, but typically using augmented training data 113B (augmented features generated by unit 103B) that can be updated during each generation of the training cycle, instead of using augmented training data generated in the data preparation phase before the training cycle is executed. In some implementations, stage 105B may use (e.g., in one or more generations of the training cycle) unaugmented training data (feature 111B) instead of augmented training data 113B. Therefore, data flow path 115B (or a path similar to or analogous to the example data flow path 115B) can be used to provide augmented training data 113B for stage 105B;

[0164] • 106B: Update phase (unit) of training loop 131B. This can be compared with ( Figure 1 Phase (unit) 106A is the same as A, but it typically operates on augmented training data 113B (generated during training phase 131B) rather than on the data preparation phase performed before the training phase (such as...). Figure 1 A) The augmented training data generated during the process is used for manipulation. In some implementations, due to... Figure 1B The novel design of the training process now facilitates the activation of optional data flow path 115B to allow stage 106B to access and use unenhanced training data 111B (e.g., ground truth labels 121E), rather than just enhanced training data 113B.

[0165] • 107B: Apply the convergence (stopping) criterion to determine whether to stop the training phase (loop) 131B. Training loop 131B typically requires running multiple iterations (i.e., multiple passes or generations of loop 131B) until the convergence criterion is satisfied, at which point it stops. Step 107B can be combined with… Figure 1 Step A is the same as step 107A;

[0166] • 108B: Stop. In this method (which can be implemented as a computer program running on a processor), once control reaches this point, training of model 114B is complete;

[0167] • 113B: Enhanced training set functionality. (Compared to...) Figure 1 Compared to the traditionally generated augmented feature 113A, augmented feature 113B is temporary intermediate data that is only needed during training in one generation (e.g., mini-batch or pass) of loop 131B. Therefore, feature 113B can be efficiently hosted in GPU memory;

[0168] • 114B: The model being trained. Model 114B can be compared with... Figure 1 Model A is the same as or similar to model 114A, but is trained based on data (in the feature domain) augmented from the training loop by augmentation unit 103B;

[0169] • 115B: Allows the update stage / unit 106B (and / or the prediction stage / unit 105B) to access an optional data flow path of the unenhanced feature 111B. This is in the present invention Figure 1B In the embodiments rather than in Figure 1 In traditional methods, this is convenient and memory-efficient, and allows (in...) Figure 1B In the examples, at least some of the following types of models are effectively trained:

[0170] This refers to a model implemented by a noise suppression system, where augmented data is the network's input and unaugmented data is the network's desired output. Such networks are typically trained using the mean squared error (MSE) loss criterion.

[0171] This is a model implemented by a noise suppression system, where augmented data is the network's input, and the gain from obtaining unaugmented data is the network's desired output. Such networks are typically trained using the mean squared error (MSE) loss criterion.

[0172] This is a model implemented by a noise suppression system, where augmented data is the network's input, and the network's desired output is the probability that each frequency band in each frame contains desired speech (e.g., the opposite of unwanted noise). Some such systems can distinguish between multiple unwanted artifacts (e.g., stationary noise, non-stationary noise, reverberation). Based on this output, the suppression gain can be determined. Such networks are typically trained using a cross-entropy loss criterion.

[0173] A model implemented by a speech analysis system (e.g., a wake-word detector, an automatic speech recognizer) in which estimates of signal-to-noise ratio (SNR), signal-to-echo ratio (SER), and / or directly to reverberation ratio (DRR) are used to weight the network input or as additional input to the network to obtain better results in the presence of noise, echo, and / or reverberation. At runtime, the SNR, SER, and / or DRR can be estimated by some signal processing component, such as a noise estimator, echo predictor, echo canceller, echo suppressor, or reverberation modeler. Here, during training, path 115B allows the deriving of ground truth SNR, SER, and / or DRR by subtracting unenhanced feature 111B from enhanced feature 113B;

[0174] ·120A-121A: with Figure 1 The corresponding (same number) element in the traditional example of A is the same;

[0175] • 121E: Label of one of the unreinforced training features 111B (copied from training set 110 by feature extraction unit 104A);

[0176] • 121F: Label of one of the augmented training features 113B (copied from training set 110 by augmentation unit 103B);

[0177] • 125A-125D: Examples of feature elements corresponding to feature 113B of non-stationary noise added by enhancement unit 103B;

[0178] • 126B: An example of a characteristic element corresponding to the reverberation characteristic 113B added by the enhancement unit 103B;

[0179] • 128B: Features of a utterance in the unenhanced training set features 111B (e.g., spectrogram or logmelspec features);

[0180] • 130B: Data preparation phase for training. It should be understood that in... Figure 1B In this embodiment, if the enhancement parameters of enhancement unit 103B change, it is not necessary to rerun the data preparation stage 130B; and

[0181] ·131B: Figure 1B The main training phase in the embodiment. Figure 1B In this embodiment, the data increment occurs during the main training loop 131B.

[0182] Embodiments of the present invention may be applied (e.g., by...) Figure 1B Examples of augmentation types (applied to training data features by augmentation feature 103B) include (but are not limited to) the following:

[0183] • Fixed-spectrum stationary noise: For example, for each utterance in the training set (e.g., each utterance in training set 110 or each utterance indicated by training set 110, and therefore each utterance in feature set 111B or each utterance indicated by feature set 111B), a random signal-to-noise ratio (SNR) is selected from a distribution of SNR values ​​(e.g., a normal distribution with a mean of 45 dB and a standard deviation of 10 dB), and stationary noise with a fixed spectrum (e.g., white noise, pink noise, or Hoechst noise) is applied at a selected level below the incoming speech signal (determined by the selected SNR value). When the input features (which will be enhanced by applying noise) are in dB band power, the added noise corresponds to taking the maximum value of the sub-band of noise power and signal power. (Refer to...) Figure 4 Describe an example of stationary noise enhancement with a fixed spectrum;

[0184] • Variable-spectrum semi-stationary noise: For example, a random SNR is selected (for example, a fixed-spectrum stationary noise), and the random stationary noise spectrum is also obtained from a distribution (e.g., a linear slope value distribution in dB / octave, or a distribution on the DCT values ​​of a log-Mel spectrum (cepstrum). The noise is then applied with a selected shape at a selected level (determined by the selected SNR value). In some embodiments, the shape of the noise changes slowly over time, for example, by selecting the rate of change per cepstrum value per second and using that rate of change to modulate the shape of the applied noise (e.g., during a generation, or between consecutive generations). (See reference...) Figure 6 Examples describing variable-spectral semi-stationary noise enhancement;

[0185] • Non-stationary noise: Noise added at random locations in the time and / or frequency spectrum of each training data feature to be augmented. For example, for each training utterance, ten rectangles are obtained, each with a random start and end time, a random start and end frequency band, and a random SNR. Noise is added within each rectangle with the corresponding SNR. (Refer to...) Figure 7 Describe an example of non-stationary noise enhancement;

[0186] • Reverberation: Applying a reverberation model (e.g., random RT60, mean free path, and distance from source to microphone) to the training data (features) to be enhanced generates enhanced training data. Typically, different reverberations are applied during the training loop (e.g., ...). Figure 1B The augmented training data used during each generation of the cycle 131B. The term "RT60" refers to the time required for the pressure level of the sound (emitted from the sound source) to decrease by 60 dB after the sound source is suddenly turned off. The reverberation model used to generate the augmented training data (features) can be of the type described in the aforementioned U.S. Patent Application Publication No. 2019 / 0362711. References will be made below. Figure 8 Describe an example of enhancement using reverberation (applying a simplified reverberation model);

[0187] • Simulated Echo Residue: To simulate the leakage of music through an echo canceller or echo suppressor (i.e., the model to be trained is such that it operates after an echo cancellation or echo suppression model and in the presence of an echo), the augmented example adds music-like noise (or other simulated echo residuals) to the feature to be augmented. Such augmented training data may be useful for training echo cancellation or echo suppression models. Some devices (e.g., some smart speakers and other smart audio devices) must routinely recognize speech incident on their microphones when music is played from their speakers and typically use echo cancellers or echo suppressors (which can be trained according to some embodiments of the invention) to partially remove echoes. A well-known limitation of echo cancellation and echo suppression algorithms is their performance degradation in “two-way conversation” scenarios, where the microphone simultaneously picks up speech and other auditory events as echo signals. For example, even when playing music or other audio, the microphone on a smart device frequently picks up the echo from the device's speaker as well as the speech of a nearby user. Under such “two-talk” conditions, echo cancellation or suppression using an adaptive filtering process may experience erroneous convergence, and the increased number of echoes may “leak.” In some cases, it may be desirable to simulate this behavior in different generations of the training loop. For example, in some embodiments, the magnitude of the added simulated echo residual (during augmentation of the training data) is at least partially based on the utterance energy (indicated by the unaugmented training data). Thus, augmentation is performed in a manner determined in part by the training data. Some embodiments gradually increase the magnitude of the added simulated echo residual over the duration during which the utterance appears in the unaugmented training vector. Examples of simulated echo residual augmentation are described below with reference to the Julia code listings (“List 1” and “List 1B”);

[0188] • Microphone Equalization: Speech recognition systems often need to operate without complete knowledge of the microphone hardware equalization characteristics. Therefore, it can be beneficial to apply a range of microphone equalization characteristics to the training data during different generations of the training cycle. For example, (during each generation of the training cycle) a random microphone tilt in dB / octave (e.g., relative to a normal distribution with a mean of 0 dB / octave and a standard deviation of 1 dB / octave) can be selected, and (during relevant generations) a filter with a corresponding linear amplitude response can be applied to the training data. When the feature domain is logarithmic (e.g., dB) band power, this can correspond to adding an octave-scale offset to each band, proportional to the distance from a reference band. Figure 5 An example describing microphone equalization enhancement;

[0189] • Microphone Shutdown: Another microphone frequency response characteristic that is not necessarily known in advance is low-frequency cutoff. For example, one microphone may pick up signals as low as 200Hz, while another microphone may pick up frequency components of signals as low as 50Hz (e.g., speech). Therefore, enhancing the features of training data by applying a random low-frequency cutoff (high-pass) filter can improve performance on a range of microphones; and / or

[0190] • Level: Another microphone parameter is the level or volume gain, which can vary depending on the microphone and acoustic environment. For example, some microphones may be more sensitive than others, and some speakers may sit closer to the microphone than others. Furthermore, some speakers may speak louder than others. Therefore, speech recognition systems must process speech within a certain range of input levels. Thus, it can be beneficial to vary the level of the input features during training. When the features are band power in dB, this can be achieved by taking a random level offset from a distribution (e.g., a uniform distribution over [-20, +20] dB) and adding that offset to all band power.

[0191] Reference Figure 4 Embodiments of the present invention are described, including fixed-spectrum stationary noise enhancement.

[0192] exist Figure 4 In this example, according to this embodiment, by feeding training data (e.g., Figure 1B Feature 111B) adds fixed-spectral stationary noise to enhance the training data (e.g., by...). Figure 1B (Function / Unit 103B execution). Figure 4 The elements include the following:

[0193] • 210: Noise spectrum example (noise power graph at dB comparison frequency);

[0194] • 211A: The flat portion of example spectrum 210, which is the spectrum 210 at the reference frequency f peak (exist Figure 4 The portion marked as having a frequency below 212 (f) peak One example value is 200Hz;

[0195] ·211B: Example spectrum 210 is higher than frequency f peak Part of the spectrum 210. Part 211B rolls off at a constant slope in dB / octave. According to Hoth's experiment (see Hoth, Daniel, The Journal of the Acoustical Society of America 12, 499 (1941); https: / / doi.org / 10.1121 / 1.1916129This represents a typical average noise roll-off of 5 dB / octave in a real room.

[0196] ·212: Reference frequency (f peak Below this frequency, the average spectrum is modeled as flat;

[0197] • 213: Plots of spectra 214, 215, and 216 (in power at dB contrast frequencies). Spectrum 214 is an example of the average speech power spectrum (e.g., the training data to be augmented), and spectrum 215 is an example of the equivalent noise spectrum 215. Spectrum 216 is an example of the noise to be added to the training data (e.g., the noise to be added by...). Figure 1B (Function / unit 103B is added to the training vector in the feature domain);

[0198] ·214: Example average speech spectrum of a training utterance (training vector);

[0199] • 215: Equivalent noise spectrum. It is formed by shifting the noise spectrum 210 by this equivalent noise power, such that the average power across all frequency bands of the equivalent noise spectrum 215 equals the average power across all frequency bands of the average speech spectrum 214. The equivalent noise power can be calculated using the following formula:

[0200] in,

[0201] x i It is the average speech spectrum in frequency band i, measured in decibels (dB).

[0202] n i It is the prototype noise spectrum in frequency band i, in decibels (dB), and

[0203] There are N frequency bands;

[0204] • 216: Added noise spectrum. This is the noise spectrum to be added to the training vector (in the feature domain). It is formed by shifting the equivalent noise spectrum 215 down according to the signal-to-noise ratio, which is derived from the signal-to-noise ratio distribution 217. Once created, the noise spectrum 216 is added to all frames of the training vector in the feature domain by taking the maximum value of the noise spectrum 216 and the signal band power 214 in each time-frequency block (i.e., included in the enhanced training vector); and

[0205] ·217: Signal-to-noise ratio (SNR) distribution. In the training cycle (e.g., Figure 1B In each generation / pass of the loop 131B, the signal-to-noise ratio (SNR) is extracted from distribution 217 (e.g., by...). Figure 1BFunction / unit 103B) is used to determine the noise to be applied in this generation / pass to augment each training vector (e.g., via function / unit 103B). Figure 4 In the example shown, the signal-to-noise ratio distribution 217 is a normal distribution with a mean of 45 dB and a standard deviation of 10 dB.

[0206] Reference Figure 5 Another embodiment of the invention, including microphone equalization enhancement, is described. Figure 5 In the example, the training data (e.g., Figure 1B Feature 111B) is enhanced by applying a filter with a randomly chosen linear magnitude response to it during each generation of the training cycle (e.g., using a different filter for each different generation). Figure 1B Function / Unit 103B). The characteristics of the filter (for each generation) are determined from a randomly selected microphone tilt (e.g., a tilt in dB / octave selected from a normal distribution of microphone tilts). Figure 5 The elements include the following:

[0207] • 220: Example microphone equalization spectrum. Spectrum (curve) 220 is a plot of the gain (dB) versus frequency (octaves) of all frames of a training vector to be added to a generation / pass in the training loop. In this example, curve 220 is linear in dB / octaves;

[0208] ·221: At the reference frequency f (i.e., including) ref (For example, f) ref The reference point (curve 220) in the frequency band of 1 kHz. Figure 5 In the middle, the equalized spectrum 220 at the reference frequency f ref The power is 0dB; and

[0209] ·222: Point 222 on curve 220 in a frequency band with arbitrary frequency f. At point 222, the equalization curve 220 has a gain “g” dB, where for a randomly chosen slope T in dB / octave, g = T log2(f – f ref For example, the value of T can be randomly obtained (for each generation / each pass).

[0210] refer to Figure 6 The invention described below (e.g., by means of) Figure 1BAnother embodiment of unit 103B is applied to augmentation of training data, wherein the augmentation includes the addition of variable-spectral semi-stationary noise. In this embodiment, for each generation (pass) of the training cycle (or once, for multiple consecutive generations of the training cycle), the SNR is extracted from the signal-to-noise ratio (SNR) distribution (as in the embodiment employing fixed-spectral stationary noise addition). Furthermore, for each generation of the training cycle (or once, for multiple consecutive generations of the training cycle), a random stationary noise spectrum is selected from the distribution of the noise spectral shape (e.g., a distribution of linear slope values ​​in dB / octave, or a distribution on the DCT values ​​of the log-Mel spectrum (cepstrum). For each generation, augmentation (i.e., noise, whose power as a function of frequency is determined by the selected SNR and the selected shape) is applied to each set of training data (e.g., each training vector). In some implementations, for example, the noise shape is modulated slowly over time (e.g., within a generation) by selecting a rate of change for each cepstrum value per second and using that rate of change.

[0211] Figure 6 The elements include the following:

[0212] • 128B: A set of input (i.e., unenhanced) training data features (e.g., temporally, the “logmelspec” band power in dB for multiple Mel-spaced bands). Feature 128B (which may be referred to as training vectors) can be or include Figure 1B One or more features of a utterance in the unenhanced training set 111B (i.e., indicated by it), and assumed to be speech data in the following description;

[0213] • 121E: Metadata associated with training vector 128B (e.g., transcription of spoken words);

[0214] •231: Speech power calculation. This can be done during the preparation time before training begins (e.g., in...). Figure 1B During the preparation phase 130B, the step of calculating the speech power of the training vector 128B is performed;

[0215] ·232: The randomly selected signal-to-noise ratio (e.g., in dB) is randomly selected from a distribution (e.g., a normal distribution with a mean of 20 dB and a standard deviation (between training vectors) of 30 dB) (e.g., in each generation);

[0216] ·233: For noise to be added to the training vector, the initial spectrum or cepstrum is randomly selected (randomly selected in each generation, or randomly selected once before the first generation);

[0217] •234: Select a random rate of change (e.g., dB / s) for the initial spectrum or cepstrum. Changes occurring at this rate may occur on different frames of the training vector (within one generation) or across different generations;

[0218] • 236: Based on the parameters selected in steps 233 and 234, calculate the effective noise spectrum or cepstral of the noise to be applied to each frame of the training vector. Noise with the same effective noise spectrum or cepstral can be applied to all frames of a training vector in one generation of training, or noise with different effective noise spectra (or cepstrals) can be applied to different frames of the training vector in one generation. To generate the noise spectrum or cepstral, zero or more (e.g., one or more) random, stationary narrowband tones can be included in (or added to) it;

[0219] • 235: An optional step to convert the effective noise cepstral spectrum into a spectral representation. If cepstral representation is used for 233, 234, and 236, step 235 converts the effective noise cepstral spectrum 236 into a spectral representation;

[0220] • 237A: By attenuating the noise spectrum generated during step 235 (or step 236, if step 235 is omitted) using the SNR value 240, a noise spectrum is generated for all (or some) frames to be applied to a training vector in one generation of training. In step 237A, the noise spectrum is attenuated (e.g., amplified) such that it is lower than the speech power determined in step 231 at a selected SNR value 232 (or higher than the speech power determined in step 231 if the selected SNR value 232 is negative);

[0221] ·237: The complete semi-stationary noise spectrum generated in step 237A;

[0222] • 238: Combine the clean (unenhanced) input features 128B with the semi-steady-state noise spectrum 237. If working in the logarithmic (e.g., dB) domain, the sum of the noise band power and the corresponding speech power can be approximated by taking (i.e., included in the enhanced training vector 239A) the maximum value of each element of the speech power and the corresponding noise band power;

[0223] • 239A: Augmented training vectors (generated during step 238) to be presented to the model (e.g., a DNN model) for training. For example, augmented training vector 239A can be generated by ( Figure 1B (of) Function 103B generated for Figure 1B An example of augmented training data 113B used in one generation of training loop 131B;

[0224] • 239B: Metadata associated with the training vector 239A (possibly required for training) (e.g., transcription of spoken words); and

[0225] • 239C: Indicates that metadata (e.g., transcription) 239B can be directly copied from the input (metadata 121E of training data 128B) to the output (metadata 239B of augmented training data 239A) data flow path because the metadata is not affected by the augmentation process.

[0226] refer to Figure 7 The invention described below (e.g., by) Figure 1B Another embodiment of the unit 103B is applied to the augmentation of training data, wherein the augmentation includes the addition of non-stationary noise. Figure 7 The elements include the following:

[0227] • 128B: A set of input (i.e., unenhanced) training data features (e.g., temporally, the “logmelspec” band power in dB for multiple Mel-spaced bands). Feature 128B (which may be referred to as training vectors) can be or include Figure 1B One or more features of a utterance in the unenhanced training set 111B (i.e., indicated by it), and assumed to be speech data in the following description;

[0228] •231: Speech power calculation. This can be done during the preparation time before training begins (e.g., in...). Figure 1B During the preparation phase 130B, the step of calculating the speech power of the training vector 128B is performed;

[0229] ·232: The randomly selected signal-to-noise ratio (e.g., in dB) is randomly selected from a distribution (e.g., a normal distribution with a mean of 20 dB and a standard deviation (between training vectors) of 30 dB) (e.g., during each generation);

[0230] •240: The time for randomly selecting an event to be inserted. The step of selecting a time (at which the event will be inserted) can be performed by drawing a random number of frames from a uniform distribution (e.g., the number of frames corresponding to the number of training vectors 128B, for example, in the range of 0-300ms), and then drawing random inter-event periods from a similar uniform distribution until the end of the training vector is reached;

[0231] ·241: For each event, a randomly selected cepstrum or spectrum (e.g., selected by drawing from a normal distribution);

[0232] ·242: For each event, the start rate and release rate are randomly selected (e.g., selected by drawing from a normal distribution);

[0233] ·243: The step of calculating the cepstral or spectrum of the resulting events for each frame of the training vector from the selected parameters 240, 241, and 242;

[0234] • 235: An optional step to convert the cepstral representation of each event into a spectral representation. If step 243 is performed in the cepstral domain using the cepstral representation of 241, then step 235 converts each cepstral representation computed in step 243 into a spectral representation;

[0235] • 237A: By attenuating (or amplifying) each noise spectrum generated during step 235 (or step 243, if step 235 is omitted) using the SNR value 232, a sequence of noise spectra to be applied to a training vector in one generation of training is generated. The noise spectra to be attenuated (or amplified) in step 237A can be considered as non-stationary noise events. In step 237A, each noise spectrum is attenuated (e.g., amplified) such that it has a selected SNR value 232 lower than the speech power determined in step 231 (or higher than the speech power determined in step 231 if the selected SNR value 232 is negative);

[0236] • 244: The complete non-stationary noise spectrum, which is the sequence of noise spectra generated in step 237A. The non-stationary noise spectrum 244 can be considered as a sequence of individual noise spectra, each noise spectrum corresponding to a discrete synthesized noise event (including...). Figure 7 The different sequences of synthetic noise events 245A, 245B, 245C and 245D shown are described, wherein individual spectra in the sequence will be applied to individual frames of training vector 128B;

[0237] • 238: Combine the clean (unenhanced) input features 128B with the semi-steady-state noise spectrum 244. If working in the logarithmic (e.g., dB) domain, the sum of the noise band power and the corresponding speech power can be approximated by taking (i.e., included in the enhanced training vector 239A) the maximum value of each element of the speech power and the corresponding noise band power;

[0238] ·246A: Augmented training vectors to be presented to the model (e.g., a DNN model) for training (in... Figure 7 (Generated during step 238). For example, the enhanced training vector 246A can be generated by ( Figure 1B The function 103B generates the following: Figure 1B An example of the augmented training data 113B used in one generation (i.e., the current pass) of training loop 131B; and

[0239] ·245A-D: Synthetic noise events in the noise spectrum 244.

[0240] The invention described below (for example, by) Figure 1B Another embodiment of the augmentation of unit 103B for training data. In this embodiment, the augmentation implements and applies a simplified reverberation model. This model is an improved (simplified) version of the energy domain reverberation algorithm described in the aforementioned U.S. Patent Application Publication No. 2019 / 0362711. The simplified reverberation model has only two parameters: RT60 and direct reverberation ratio (DRR). The mean free path and source distance are summarized as DRR parameters.

[0241] Figure 8 The elements include the following:

[0242] • 128B: A set of input (i.e., unenhanced) training data features (e.g., temporally, the “logmelspec” band power in dB for multiple Mel-spaced bands). Feature 128B (which may be referred to as training vectors) can be or include Figure 1B One or more features of a utterance in the unenhanced training set 111B (i.e., indicated by it), and assumed to be speech data in the following description;

[0243] ·250: A specific frequency band "i" of the 128B training vector to be reverberated. During execution Figure 8 In this method, reverberation can be added sequentially to each frequency band of vector 128B, or added in parallel to all frequency bands. Figure 8 The following description pertains to the enhancement of a specific frequency band (250) of vector 128B;

[0244] ·251:x[i,t], a value in band 250 at time “t”, which is the input power of band “i” of data 128B at time “t”, in dB;

[0245] ·252: The step of subtracting parameter 263 (DRR) from the input band power 251 to determine x[i,t]–DRR;

[0246] ·253: The steps to determine the maximum value of state[i, t-1] + α[i] and x[i, t] – DRR, where “x[i, t] – DRR” is the output of step 252 and “state[i, t-1] + α[i]” is the output of step 255;

[0247] ·254: The state variable state[i, t] is updated for each frame of vector 128B. For each frame t, step 253 uses state[i, t-1], and then the result of step 253 is written back to state[i, t].

[0248] ·255: The steps to calculate the value “state[i, t-1] + α[I]”;

[0249] ·256: The steps that generate noise (e.g., Gaussian noise with a mean of 0 dB and a standard deviation of 3 dB);

[0250] ·257: The step of canceling the reverberation tail (output of step 255) with noise (generated in step 256);

[0251] • 258: The step of determining the maximum value of the reverberation energy (output of step 257) and the direct energy (251). This is the step of combining the reverberation energy and the direct energy, and is an approximation (approximate implementation) of the step of adding the reverberation energy value to the corresponding direct energy value;

[0252] ·259: The output power y[i, t] at time “t” for frequency band “i”, as determined in step 258;

[0253] •260: Reverberant output features used to train the model (i.e., augmented training data) (Feature 260 is...) Figure 1B Example of augmented training data 113B, which is used to train model 114b);

[0254] ·260A: Output power 259 for all times “t”, which is a frequency band (the “i”th frequency band) of output feature 260 and is generated in response to frequency band 250 of training vector 128B;

[0255] ·261: Execution Figure 8 Time indication for the step (dashed line). Figure 8 All elements above the midline 261 (i.e., 262, 263, 262A, 263A, 264, 264A, 265, 266, 266A, and 266B) are generated or executed once per generation for each training vector. Figure 8 All elements below the midline 261 are generated or executed once per frame for each training vector in each generation;

[0256] ·262: Indicates for each generation (e.g., Figure 1B The training loop (131B, each generation) uses a randomly selected reverberation time RT60 (e.g., in milliseconds) as a parameter for each training vector (128B). Here, "RT60" represents the time required for the pressure level of the sound (emitted from the sound source) to decrease by 60 dB after the sound source is suddenly turned off. For example, the RT60 parameter 262 can be extracted from a normal distribution with a mean of 400 milliseconds and a standard deviation of 100 milliseconds.

[0257] ·262A: Display parameter 262 (in Figure 8The data stream path marked "RT60" was used to execute step 264;

[0258] • 263: This parameter represents the direct reverberation ratio (DRR) value (e.g., in dB) randomly selected for each training vector in each generation. For example, the DRR parameter 263 can be extracted from a normal distribution with a mean of 8 dB and a standard deviation of 3 dB.

[0259] ·263A: Shows DRR parameter 263 (in Figure 8 The data stream path marked as "DRR(dB)" is used once per frame (to perform step 252);

[0260] ·264: Steps to derate broadband parameter 262 (RT60) in terms of frequency to address the phenomenon that most rooms have more reverberant energy at high frequencies than at low frequencies;

[0261] • 264A: The derating RT60 parameters (labeled "RT60i") generated for frequency band "i" in step 263. Each derating parameter (RT60... i ) is used to enhance data 251("x[i,t]") in the same frequency band "i";

[0262] •265: Parameter Δt indicates the frame duration. For example, parameter Δt can be expressed in milliseconds;

[0263] ·266: The steps for calculating the coefficient “α[I]”, where the index “i” represents the “i”th frequency band, are as follows: α[i]=-60(Δt) / RT60 i Among them, "RT60" i " is parameter 264A, Δt is parameter 265;

[0264] ·266B: The coefficient “α[i]” generated for frequency band “i” in step 266; and

[0265] ·266A: Shows the data flow path where each coefficient 266B (“α[i]”) is used once per frame (execute step 255).

[0266] refer to Figure 9 The following describes a time-frequency block classifier training pipeline 300 implemented in some embodiments of the present invention. The training pipeline 300 (e.g., in...) Figure 1B In some embodiments of training loop 131B, the training data (input feature 128B) is augmented in the training loop of the multi-style training method, and class label data (311, 312, 313, and 314) is also generated in the training loop. The augmented training data (310) and class labels (311-314) can be used (e.g., in the training loop) to train the model (e.g., Figure 1B Model 114B), such that the trained model (e.g., by Figure 3 A classifier 207 or another classifier implementation can be used to classify time-frequency slices of input features as speech, stationary noise, non-stationary noise, or reverberation. A model trained in this way can be used for noise suppression (e.g., including classifying time-frequency blocks of input features as speech, stationary noise, non-stationary noise, or reverberation, and suppressing unwanted non-speech sounds). Figure 9 Steps 303, 315, and 307 are performed to augment the input training data (e.g., in the "logmelspec" band energy domain on the GPU) during training.

[0267] • 300: Time-frequency block classifier training pipeline;

[0268] • 128B: Input features are acoustic features derived from the output of a set of microphones (e.g., at least some of the microphones in a system that includes orchestrated smart devices). Features 128B (sometimes called vectors) are organized as time-frequency blocks of data;

[0269] • 301: Speech mask, describing the prior probability of speech primarily contained in each time-frequency block (in a clean input vector of 128B). The speech mask comprises data values, each corresponding to a probability (e.g., within a range including high and low probabilities). For example, such a speech mask can be generated using Gaussian mixture modeling at the level of each frequency band in the vector of 128B. For instance, a diagonal covariance Gaussian mixture model containing two Gaussians can be used.

[0270] ·302: Data preparation phase for multi-style training methods (e.g., Figure 1B Phase 130A) and training cycles (e.g., Figure 1B The line separates the training loop 131A from the data preparation phase. Everything to the left of this line (i.e., the generation of features 128B and mask 301) occurs during the data preparation phase. Everything to the right of this line occurs for each vector in each generation and can be implemented on a GPU.

[0271] • 304: Synthesized stationary (or semi-stationary) noise. For example, noise 304 could be like... Figure 6 An example of synthesized semi-stationary noise generated in the same way as element 237. To generate noise 304, one or more random stationary narrowband tones can be included in its spectrum (e.g., as shown in the image). Figure 6 (as indicated in the description of element 236);

[0272] ·305: Synthesized non-stationary noise. In Figure 7 In the example, element 244 is a synthesized non-stationary noise;

[0273] • 303: Steps (or units) to enhance the clean feature 128B by combining it with stationary (or semi-stationary) noise 304 and / or non-stationary noise 305. If operating in the logarithmic power domain (e.g., dB), the feature and noise can be approximated by taking the maximum value in terms of elements;

[0274] • 306: Dirty features (enhanced features) created during step 303 by combining clean feature 128B with stationary (or semi-stationary) and / or non-stationary noise;

[0275] ·315: The step (or unit) of enhancing dirty feature 306 by applying reverberation (e.g., synthesized reverberation) to dirty feature 306. This enhancement can be achieved, for example, by... Figure 8 This is achieved by executing steps within the training loop (using...) Figure 8 (Values ​​generated during the data preparation phase);

[0276] • 308: Enhanced features generated by step 315 (with added reverberation, such as synthetic reverberation);

[0277] • 307: A step (or unit) of applying leveling, equalization, and / or microphone cutoff filtering to feature 308. This processing is an enhancement of feature 308, which is one or more of the types described above, such as leveling, microphone equalization, and microphone cutoff enhancement;

[0278] • 310: The final enhanced feature generated by step 307. Feature 310 (which can be presented to a system implementing the model to be trained, for example, to a network of such a system) contains at least some of the following: synthesized stationary (or semi-stationary) noise, non-stationary noise, reverberation, level, microphone equalization, and microphone cutoff enhancement;

[0279] • 309: Steps (or units) for class tags. In Figure 9 The step (or unit) identified as "class label logic" tracks the dominant type of enhancements applied throughout the illustrated process to generate enhancement feature 310 for each time-frequency block (if any of them are dominant). For example: in each time-frequency block where clean speech is still the dominant contributor (or for each time-frequency block) (i.e., if enhancements 303, 315, and 307 are not considered dominant), step / unit 309 in its P speech A 1 is recorded in output (311), while a 0 is recorded for all other outputs (312, 313, and 314); in each time-frequency block where reverberation is the dominant contributor (or for each time-frequency block), step / unit 309 will record a 1 in its P... reverb The output (314) is recorded as 1, while all other outputs (311, 312, and 313) are recorded as 0; and so on;

[0280] • 311, 312, 313, and 314: Training class labels P speech (Label 311 indicates no enhancement is dominant), P stationary (Label 312 indicates predominantly stable or semi-stationary noise enhancement), P nonstationary (Label 313 indicates that non-stationary noise enhancement is dominant) and P reverb (Label 314 indicates that reverberation enhancement is dominant).

[0281] Class labels 311-314 can be compared with the model output (the output of the trained model) to calculate the loss gradient during backpropagation during training. Already (or currently) being used. Figure 9 The classifier trained by the scheme (e.g., implementation) Figure 1B The classifier of model 114B can, for example, in its output (e.g., Figure 1B The output of the prediction step 105B of the training loop includes an element-wise softmax, which indicates the probability of speech, stationary noise, non-stationary noise, and reverberation in each time-frequency block. These predicted probabilities can be compared with class labels 311-314 using, for example, cross-entropy loss and backpropagation gradients, to (e.g., in...) Figure 1B In step 106 of the training loop, the model parameters are updated.

[0282] Figure 10 Examples of four augmented training vectors (400, 401, 402, and 403) are shown, each augmented by adjusting the same training vector (e.g., ...). Figure 9 The input features (128B) are generated by applying different augmentations so that they can be used during different training iterations in the training cycle. Each of the augmented training vectors (400-403) is ( Figure 9 An example of the enhanced feature 310, which has been generated to implement Figure 9 The method is used during different training iterations of the training cycle. Figure 10 middle:

[0283] • The augmented training vector 400 is an instance of the augmented feature 310 from the first generation of training;

[0284] • The augmented training vector 401 is an instance of the augmented feature 310 in the second-generation training;

[0285] • Augmented training vector 402 is an instance of augmented feature 310 on the third-generation training; and

[0286] • Enhanced training vector 403: An instance of enhanced feature 310 on fourth-generation training.

[0287] Each of vectors 400-403 includes a banded frequency component (in bands) for each frame in the sequence, with frequency indicated on the vertical axis and time indicated on the horizontal axis (in frames). Figure 10 In the diagram, scale 405 indicates how the shadows of vectors 400-403 (i.e., different brightness in different areas) correspond to power in dB.

[0288] Next, referring to the Julia 1.1 code listing below (“Listing 1”), an example of simulating echo residual enhancement is described. This is done when a processor (e.g., programmed to implement...) Figure 1B When the Julia 1.1 code in Listing 1 is executed (using the 103B processor), it generates the training data to be added to the training data to be augmented (e.g., ...). Figure 1B The simulated echo residuals (similar to musical noise, determined using data values ​​indicating melody, rhythm, and pitch, as shown in the code) of feature 111B can then be added to frames of features (the training data to be enhanced) to generate enhanced features for use in training an acoustic model (e.g., an echo cancellation or echo suppression model) in one generation of the training loop. More generally, simulated music (or other simulated sound) residuals can be combined with (e.g., added to) the training data to generate enhanced training data for use in one generation of the training loop (e.g., Figure 1B The acoustic model is trained in the training loop of generation 131B.

[0289] List 1:

[0290] The generation will be a batch of synthesized music residuals that are combined with a batch of input speech by taking the maximum value in terms of elements, where

[0291] nband: The number of frequency bands.

[0292] nframe: The number of time frames for which residuals are to be generated.

[0293] nvector: The number of vectors to be generated in the batch.

[0294] dt_ms: Frame size in milliseconds.

[0295] meandifflog_fband: This describes how the frequency bands are spaced. For any array of frequency band center frequencies fband, it is calculated using mean(diff(log.(fband))).

[0296] The function below generates a three-dimensional array of residual band energy in dB along the dimension (nband, nframe, nvector).

[0297]

[0298] The following Julia 1.1 code listing (“Listing 1B”) describes another example of simulating echo residual enhancement. This is done when a processor (e.g., programmed to implement...) Figure 1B When the processor of Function 103B executes, the code of List 1B generates the data to be added to the training data to be augmented (e.g., Figure 1B The simulated echo residual (synthesized musical noise) is a feature 111B. The magnitude (amplitude) of the simulated echo residual varies depending on the location of the utterance in the training data (training vectors).

[0299] List 1B:

[0300] """

[0301] The generation will be a batch of synthesized music residuals that are combined with a batch of input speech by taking the maximum value in terms of elements, where

[0302] nband: The number of frequency bands.

[0303] nframe: The number of time frames for which residuals are to be generated.

[0304] nvector: The number of vectors to be generated in the batch.

[0305] dt_ms: Frame size in milliseconds.

[0306] meandifflog_fband: This describes how the frequency bands are spaced. For any array of frequency band center frequencies fband, it is calculated using mean(diff(log.(fband))).

[0307] utterance_spectrum: A three-dimensional array in dimensions (nband, nframe, nvector) representing the frequency band energy of the input speech, expressed in dB.

[0308] The following Julia 1.4 code listing (“List 2”) describes an example implementation of augmenting training data by adding variable spectral stationary noise to the training data (e.g., as referenced above). Figure 6 As described above). When by a processor (e.g., programmed to implement...) Figure 1B When executed by the 103B processor, the code in Listing 2 generates stationary noise (with a variable spectrum) to be combined with the unenhanced training data (e.g., in...). Figure 6 In step 238), thus generating enhanced training data for training the acoustic model.

[0309] In the list ("List 2"):

[0310] The training data being augmented (e.g., Figure 6 The data (128 bytes) is provided to the function `batch_generate_statistic_noise` in the independent variable `x`. In this example, it is a three-dimensional array. The first dimension is the frequency band. The second dimension is time. The third dimension is the number of vectors within the batch (typical deep learning systems divide the training set into batches or "mini-batches" and update the model after running the prediction step 105 bytes for each batch); and

[0311] Voice power (e.g., in) Figure 6 Those generated in step 231 are passed to the batch_generate_statistic_noise function in the nep argument (where "nep" represents the noise equivalent power and can be used). Figure 4 The process shown is used to calculate (Nep is an array because each training vector in the batch has a speech power).

[0312] List 2

[0313]

[0314]

[0315] The following Julia 1.4 code listing (“List 3”) describes an example implementation of augmenting training data by adding non-stationary noise to the training data (e.g., as shown in the above reference). Figure 7 As described above). When by a processor (e.g., programmed to implement...) Figure 1B When the function 103B is executed on a GPU or other processor, the code in Listing 3 generates non-stationary noise, which will be combined with the unenhanced training data (e.g., in...). Figure 7 In step 238), enhanced training data is generated for training the acoustic model.

[0316] In the list ("List 3"):

[0317] The incoming training data (e.g., Figure 7 The data (128B) is presented to the `batch_generate_nonstability_noise` function in the `x` parameter. As shown in Listing 2, it is a three-dimensional array;

[0318] Voice power (e.g., in) Figure 7Those generated in step 231 are presented in the nep argument to the batch_generate_nonstability_noise function;

[0319] The “cepstrum_dB_mean” data describes the data used to generate the cepstrum of random events. Figure 7 The cepstral mean of element 241 in dB; and

[0320] The “cepstrum_dB_stddev” data is used to obtain the cepstrum of random events. Figure 7 The standard deviation of the element 241). In this example, we obtained a 6-dimensional cepstral, so each of these vectors has 6 elements;

[0321] The “attack_cepstrum_dB_per_s_mean” and “attack_cepstrum_dB_per_s_stddev” data describe the distribution from which to obtain the random startup rate. Figure 7 element 242); and

[0322] The “release_cepstrum_dB_per_s_mean” and “release_cepstrum_dB_per_s_stddev” data describe the distribution from which the random release rate is to be obtained. Figure 7 (element 242).

[0323] List 3:

[0324]

[0325]

[0326] "Convert parameters that are independent of banding and time step into coefficients to be used in batch_generate_nonstationary_noise."

[0327]

[0328]

[0329] The helper function calls `batch_generate_nonstationary_noise()` with `Base.randn()` as the random number generator.

[0330]

[0331] "The auxiliary function generates a cepstrum for an event (243)."

[0332]

[0333] """

[0334] Generate non-stationary noise (bandwidth energy in dB) of the same magnitude as the input batch x.

[0335] x is a three-dimensional array describing the bandwidth energy in dB of a batch of training vectors, where:

[0336] - Dimension 1 is the frequency band

[0337] - Dimension 2 is the time frame

[0338] - Dimension 3 indexing on vectors in a batch

[0339] nep is the speech power (e.g., noise equivalent power) of each vector in the batch.

[0340] xrandn is a function that obtains an array of counts from a standard normal distribution. For example, use Base.randn() when running on a CPU, but use CuArrays.randn() when running on a GPU.

[0341]

[0342]

[0343] The following Julia 1.4 code listing (“List 4”) describes an example implementation of augmenting training data (input features) to generate reverberant training data (as described above). Figure 8 As described above). When by a processor (e.g., programmed to implement...) Figure 1B When the function (the graphics processor of the 103B or other processor) is executed, the code in Listing 4 generates a reverberation energy value to combine with the unenhanced training data, and combines that value with the training data (i.e., implements...). Figure 8 Step 258) generates enhanced training data for training the acoustic model.

[0344] List 4:

[0345]

[0346]

[0347] Obtain random reverb parameters from the distribution and return the "x" reverb version.

[0348] x is a three-dimensional array describing the bandwidth energy in dB of a batch of training vectors, where:

[0349] - Dimension 1 is the frequency band

[0350] - Dimension 2 is the time frame

[0351] - Dimension 3 indexing on vectors in a batch

[0352] This function returns a tuple (y, mask), where y is a three-dimensional array of reverberant band energies of the same size as x. Mask is a three-dimensional array of the same size as y, whose:

[0353] For each time-frequency block where reverberation energy has been added, it is 1.

[0354] Otherwise, it is 0.

[0355]

[0356]

[0357] Some embodiments of the present invention include one or more of the following listed aspects:

[0358] EA 1. A method for training an acoustic model, wherein training includes a data preparation phase and a training cycle following the data preparation phase, wherein the training cycle includes at least one generation, the method comprising:

[0359] During the data preparation phase, training data is provided, wherein the training data is or includes at least one example of audio data;

[0360] During the training loop, the training data is augmented to generate augmented training data; and

[0361] During each generation of the training cycle, the model is trained using at least some of the augmented training data.

[0362] EA 2. The method according to EA1, wherein different subsets of augmented training data are generated during training cycles for different generations of training cycles by augmenting at least some of the training data with different sets of augmentation parameters obtained from multiple probability distributions.

[0363] EA 3. The method according to any one of EA1-2, wherein the training data indicates multiple utterances of the user.

[0364] EA 4. The method according to any one of EA1-3, wherein the training data indicates features extracted from temporal input audio data, and the enhancement occurs in at least one feature domain.

[0365] EA 5. The method according to EA4, wherein the characteristic domain is the Mel frequency cepstral coefficient (MFCC) domain, or the logarithm of the band power of multiple frequency bands.

[0366] EA 6. The method according to any one of EA1-5, wherein the acoustic model is a speech analysis model or a noise suppression model.

[0367] EA 7. The method according to any one of EA1-6, wherein the training is or includes training a deep neural network (DNN), or a convolutional neural network (CNN), or a recurrent neural network (RNN), or an HMM-GMM acoustic model.

[0368] EA 8. The method according to any one of EA1-7, wherein the enhancement comprises at least one of the following: adding fixed-spectrum stationary noise, adding variable-spectrum stationary noise, adding noise comprising one or more random stationary narrowband tones, adding reverberation, adding non-stationary noise, adding analog echo residual, analog microphone equalization, analog microphone off, or changing the bandwidth level.

[0369] EA 9. The method according to any one of EA1-8, wherein the enhancement is implemented in or on one or more graphics processing units (GPUs).

[0370] EA 10. The method according to any one of EA1-9, wherein the training data indicates features including frequency bands, the features being extracted from time-domain input audio data, and the enhancement occurring in the frequency domain.

[0371] EA 11. The method according to EA10, wherein each frequency band occupies a constant proportion of the Mel spectrum, or is distributed at equal intervals at logarithmic frequencies, or is distributed at equal intervals at logarithmic frequencies, wherein the logarithm is scaled such that the feature is expressed in decibels (dB) for band power.

[0372] EA 12. The method according to any one of EA1-11, wherein the training is implemented by a control system, the control system comprising one or more processors and one or more devices implementing non-transitory memory, the training comprising providing the training data to the control system, and the training producing a trained acoustic model, wherein the method comprises:

[0373] The parameters of the trained acoustic model are stored in one or more of the devices.

[0374] EA 13. The method according to any one of EA1-11, wherein the enhancement is performed in a manner determined in part based on the training data.

[0375] EA 14. An apparatus comprising an interface system and a control system, the control system including one or more processors and one or more devices implementing non-transitory memory, wherein the control system is configured to perform the method according to any one of EA1-13.

[0376] EA 15. A system configured for training an acoustic model, wherein training includes a data preparation phase and a training cycle following the data preparation phase, wherein the training cycle includes at least one generation, the system comprising:

[0377] A data preparation subsystem, coupled and configured to implement a data preparation phase, includes receiving or generating training data, wherein the training data is or includes at least one example of audio data; and

[0378] The training subsystem, coupled to the data preparation subsystem, is configured to augment the training data during the training cycle to generate augmented training data, and to use at least some of the augmented training data to train the model during each generation of the training cycle.

[0379] EA 16. The system according to EA15, wherein the training subsystem is configured to include augmenting at least some of the training data by using different sets of augmentation parameters obtained from multiple probability distributions, generating different subsets of the augmented training data during the training cycle for different generations of the training cycle.

[0380] EA 17. The system according to EA15 or 16, wherein the training data indicates multiple utterances of the user.

[0381] EA 18. The system according to any one of EA15-17, wherein the training data indicates features extracted from temporal domain input audio data, and the training subsystem is configured to augment the training data in at least one feature domain.

[0382] EA 19. The system according to EA18, wherein the characteristic domain is the Mel frequency cepstral coefficient (MFCC) domain, or the logarithm of the band power of multiple frequency bands.

[0383] EA 20. The system according to any one of EA15-19, wherein the acoustic model is a speech analysis model or a noise suppression model.

[0384] EA 21. The system according to any one of EA15-20, wherein the training subsystem is configured to train a model, including training a deep neural network (DNN), or a convolutional neural network (CNN), or a recurrent neural network (RNN), or an HMM-GMM acoustic model.

[0385] EA 22. The system according to any one of EA15-21, wherein the training subsystem is configured to enhance the training data by performing at least one of the following: adding fixed-spectrum stationary noise, adding variable-spectrum stationary noise, adding noise comprising one or more random stationary narrowband tones, adding reverberation, adding non-stationary noise, adding simulated echo residuals, simulating microphone equalization, simulating microphone shutdown, or changing the bandwidth level.

[0386] EA 23. The system according to any one of EA15-22, wherein the training subsystem is implemented in or on one or more graphics processing units (GPUs).

[0387] EA 24. The system according to any one of EA15-23, wherein the training data indicates features including frequency bands, the data preparation subsystem is configured to extract features from time-domain input audio data, and the training subsystem is configured to augment the training data in the frequency domain.

[0388] EA 25. The system according to EA24, wherein each frequency band occupies a constant proportion of the Mel spectrum, or is equally spaced at logarithmic frequencies, or equally spaced at logarithmic frequencies, wherein the logarithm is scaled such that the characteristic is expressed in decibels (dB) for band power.

[0389] EA 26. The system according to any one of EA15-25, wherein the training subsystem includes one or more processors and one or more devices implementing non-transitory memory, and the training subsystem is configured to generate a trained acoustic model and store parameters of the trained acoustic model in one or more of the devices.

[0390] EA 27. The system according to any one of EA15-26, wherein the training subsystem is configured to augment the training data in a manner determined in part based on the training data.

[0391] Some embodiments of region mapping (e.g., in the context of wake word detection or other speech analysis processing), and some embodiments of the present invention (e.g., for training an acoustic model used in speech analysis processing including region mapping), include one or more of the following:

[0392] Example 1. A method for estimating a user's location in an environment (e.g., as a region label), wherein the environment includes a plurality of predetermined regions and a plurality of microphones (e.g., each microphone is included in or coupled to at least one smart audio device in the environment), the method comprising the steps of:

[0393] Determine (e.g., based at least in part on the microphone's output signal) an estimate of which area the user is located in;

[0394] Example 2. The method of Example 1, where the microphone is asynchronous (e.g., asynchronous and randomly distributed);

[0395] Example 3. The method of Example 1, wherein the model is trained on features derived from multiple wake word detectors about multiple wake word utterances at multiple locations;

[0396] Example 4. The method of Example 1, where the user region is estimated as the class with the highest posterior probability;

[0397] Example 5. The method of Example 1, where the model is trained using training data labeled with a reference region;

[0398] Example 6. The method of Example 1, where the model is trained using unlabeled training data;

[0399] Example 7. The method of Example 1, wherein a Gaussian mixture model is trained on normalized wake-word confidence, normalized average reception level, and maximum reception level;

[0400] Example 8. The method of any of the preceding examples, wherein the adaptation of the acoustic region model is performed online;

[0401] Example 9. The method of Example 8, wherein the adaptation is based on explicit feedback from the user;

[0402] Example 10. The method of Example 8, wherein the adaptation is based on implicit feedback of successful microphone selection or beamforming based on the predicted acoustic region;

[0403] Example 11. The method of Example 10, wherein the implicit feedback includes a response to the user prematurely terminating the voice assistant;

[0404] Example 12. The method of Example 10, wherein the implicit feedback includes the command recognizer returning a low-confidence result; and

[0405] Example 13. The method of Example 10, wherein the implicit feedback includes a two-pass retrospective wake word detector that returns a low confidence that the wake word was spoken.

[0406] Aspects of the invention include systems or devices configured (e.g., programmed) to perform any embodiment of the methods of the invention, and tangible computer-readable media (e.g., disks) storing code for implementing any embodiment or steps of the methods of the invention. For example, the systems of the invention may be or include programmable general-purpose processors, digital signal processors, or microprocessors, programmed with software or firmware and / or otherwise configured to perform various operations on data, including embodiments of the methods of the invention or steps thereof. Such general-purpose processors may be or include computer systems including input devices, memory, and processing subsystems programmed (and / or otherwise configured) to perform embodiments of the methods of the invention (or steps thereof) in response to data asserted thereto.

[0407] Some embodiments of the system of the present invention are implemented as configurable (e.g., programmable) digital signal processors (DSPs) or graphics processing units (GPUs), configured (e.g., programmed and otherwise configured) to perform desired processing on audio signals, including performing embodiments of the methods of the present invention or steps thereof. Optionally, embodiments of the system (or elements thereof) of the present invention are implemented as general-purpose processors (e.g., personal computers or other computer systems or microprocessors that may include input devices and memory), programmed with software or firmware and / or otherwise configured to perform any of a variety of operations including embodiments of the methods of the present invention. Optionally, elements of some embodiments of the system of the present invention are implemented as general-purpose processors, graphics processors, or digital signal processors configured (e.g., programmed) to perform embodiments of the methods of the present invention, and the system also includes other elements (e.g., one or more speakers and / or one or more microphones). The general-purpose processor configured to perform embodiments of the methods of the present invention is typically coupled to input devices (e.g., a mouse and / or keyboard), memory, and a display device.

[0408] Another aspect of the invention is a computer-readable medium (e.g., a disk or other tangible storage medium) that stores code (e.g., an encoder executable to carry out any embodiment of the method or steps of the invention) for implementing the method or steps of the invention.

[0409] While specific embodiments and applications of the invention have been described herein, it will be apparent to those skilled in the art that many variations of the described embodiments and applications are possible without departing from the scope of the invention as described and claimed herein. It should be understood that while certain forms of the invention have been shown and described, the invention is not limited to the specific embodiments described and shown or the specific methods described.

Claims

1. A method of training an acoustic model, wherein the training comprises a data preparation phase and a training loop following the data preparation phase, wherein the training loop comprises a plurality of successive generations, the method comprising: providing training data in the data preparation phase, wherein the training data comprises at least one example of audio data, wherein the training data comprises features extracted from time-domain input audio data and managed in frequency bands; during each generation of the training loop, different augmentations are performed on the same training data, thereby generating different augmented training data, which is temporarily and only used during one generation of training, wherein the augmentations are performed in the frequency domain; and during each generation of the training loop, the model is trained using at least some of the augmented training data, wherein the augmentations comprise extracting a signal-to-noise ratio (SNR) from an SNR distribution, selecting a random stationary noise spectrum from a distribution of noise spectrum shapes, and applying noise to the training data, wherein the power of the noise as a function of frequency is determined according to the extracted SNR and the selected random stationary noise spectrum.

2. The method of claim 1, wherein, Different subsets of augmented training data are generated during the training loop for different generations of the training loop by augmenting at least some of the training data using different sets of augmentation parameters drawn from a plurality of probability distributions.

3. The method of claim 1, wherein the training data is indicative of a plurality of utterances of a user.

4. The method of claim 1, wherein the augmentations occur in at least one feature domain.

5. The method of claim 4, wherein the feature domain is a Mel Frequency Cepstral Coefficient (MFCC) domain, or a log of band powers of a plurality of frequency bands.

6. The method of claim 1, wherein the acoustic model is a speech analysis model or a noise suppression model.

7. The method of claim 1, wherein the training comprises training a deep neural network (DNN), or a convolutional neural network (CNN), or a recurrent neural network (RNN), or an HMM-GMM acoustic model.

8. The method of claim 1, wherein the augmentations comprise at least one of: adding fixed spectral stationary noise, adding variable spectral stationary noise, adding noise comprising one or more random stationary narrowband tones, adding reverberation, adding non-stationary noise, adding simulated echo residual, simulating microphone equalization, simulating microphone turn-off, or changing wideband level.

9. The method of claim 1, wherein the augmentations are implemented in one or more graphics processing units (GPUs).

10. The method of claim 1, wherein each frequency band occupies a constant proportion of a Mel-spectrum, or is equally spaced in log frequency, or is equally spaced in log frequency scaled such that the features represent band powers in decibels (dB).

11. The method of claim 1, wherein the augmentations are performed in a manner that is partially determined from the training data.

12. The method of claim 1, wherein the training is implemented by a control system comprising one or more processors and one or more devices implementing non-transitory memory, the training comprises providing the training data to the control system, and the training produces a trained acoustic model, wherein the method comprises: storing parameters of the trained acoustic model in one or more of the devices.

13. An apparatus comprising an interface system and a control system comprising one or more processors and one or more devices implementing non-transitory memory, wherein the control system is configured to perform the method of any one of claims 1-12.

14. A system configured for training an acoustic model, wherein the training comprises a data preparation phase and a training loop following the data preparation phase, wherein the training loop comprises a plurality of successive generations, the system comprising: a data preparation subsystem coupled and configured to implement the data preparation phase, comprising receiving or generating training data, wherein the training data comprises at least one example of audio data, wherein the training data comprises features extracted from the time domain input audio data and managed in frequency bands; and a training subsystem coupled to the data preparation subsystem and configured to perform different augmentations on the same training data during each generation of the training loop, thereby generating different augmented training data, which is used temporarily and only during one generation of training, and to use at least some of the augmented training data during each generation of the training loop to train the model, wherein the augmentations are performed in the frequency domain, wherein, the augmentations comprise extracting an SNR from a signal-to-noise ratio (SNR) distribution, selecting a random stationary noise spectrum from a distribution of noise spectral shapes, and applying noise to the training data, wherein the power of the noise as a function of frequency is determined according to the extracted SNR and the selected random stationary noise spectrum.

15. The system of claim 14, wherein the training subsystem is configured to augment at least some of the training data by using different sets of augmentation parameters drawn from a plurality of probability distributions, to generate different subsets of the augmented training data during the training loop for different generations of the training loop.

16. The system of claim 14, wherein the training data is indicative of a plurality of utterances of a user.

17. The system of claim 14, wherein the training subsystem is configured to augment the training data in at least one feature domain.

18. The system of claim 17, wherein the feature domain is a Mel Frequency Cepstral Coefficient (MFCC) domain, or a log of band powers of a plurality of frequency bands.

19. The system of claim 14, wherein the acoustic model is a speech analysis model or a noise suppression model.

20. The system of claim 14, wherein the training subsystem is configured to train the model comprising training a deep neural network (DNN), or a convolutional neural network (CNN), or a recurrent neural network (RNN), or an HMM-GMM acoustic model.

21. The system of claim 14, wherein, The training subsystem is configured to augment the training data, including performing at least one of: adding fixed spectral flat noise, adding variable spectral flat noise, adding noise including one or more random flat narrowband tones, adding reverberation, adding non-flat noise, adding simulated echo residual, simulating microphone equalization, simulating microphone cutoff, or changing a wideband level.

22. The system of claim 14, wherein the training subsystem is implemented in or on one or more graphics processing units (GPUs).

23. The system of claim 14, wherein each frequency band occupies a constant proportion of the mel-spectrum, or is equally spaced in log frequency, or is equally spaced in log frequency scaled such that the features represent frequency band power in decibels (dB).

24. The system of claim 14, wherein the training subsystem is configured to augment the training data in a manner determined in part from the training data.

25. The system of claim 14, wherein the training subsystem includes one or more processors and one or more devices implementing non-transitory memory, and the training subsystem is configured to produce a trained acoustic model and store parameters of the trained acoustic model in one or more of the devices.

26. A non-transitory storage medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any of claims 1-12.

27. A device for training an acoustic model, comprising: one or more processors; and a non-transitory storage medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any of claims 1-12.

28. A computer program product comprising instructions that, when executed by one or more processors, cause performance of the method of any of claims 1-12. ​

Citation Information

Patent Citations

  • Training of acoustic models for far-field vocalization processing systems

    US20190362711A1

  • Method and system for training far-field speech acoustic model

    CN107680586A

  • Method for training a statistical classifier with reduced tendency for overfitting

    US5903884A