Speech recognition using natural language understanding related knowledge via deep feedforward neural networks
Through the deep feedforward neural network framework combined with NLU knowledge, ranking the assumptions generated by multiple speech recognition engines is solved, and the problem that speech recognition systems in the prior art are difficult to combine high-level language information, improving the accuracy and efficiency of speech recognition, and are suitable for in-vehicle infotainment and home automation systems.
Patent Information
- Application Number
- CN202010391548.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-10
- Filing Date
- 2020-05-11
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-05-11
AI Technical Summary
When the existing speech recognition system processes the output of multiple speech recognition engines, it is difficult to effectively combine high-level language information, resulting in insufficient accuracy of speech recognition, especially when users need to focus on performing tasks, traditional input devices are inconvenient.
The deep feedforward neural network framework is adopted, combined with natural language understanding (NLU) knowledge, and the assumptions generated by multiple speech recognition engines are ranked. Through the shared projection matrix and BLSTM feature representation, slot filling and intent detection features are extracted to generate the final speech recognition results.
It improves the accuracy and efficiency of speech recognition, reduces dependence on traditional input devices, and is suitable for scenarios such as in-vehicle infotainment systems and home automation systems.
Smart Images

Figure CN111916070B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Provisional Application Serial No. 62 / 846,340, filed on May 10, 2019, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] This disclosure generally relates to the field of automatic speech recognition, and more particularly, to systems and methods for improving the operation of a speech recognition system that utilizes one or more speech recognition engines. Background Art
[0004] Automatic speech recognition is an important technology for implementing a human - machine interface (HMI) in a wide range of applications. In particular, in situations where a human user needs to focus on performing a task, speech recognition is useful, as using traditional input devices such as a mouse and keyboard is inconvenient or impractical. For example, many uses of in - vehicle “infotainment” systems, home automation systems, and small electronic mobile devices such as smartphones, tablets, and wearable computers can employ speech recognition to receive voice - based commands and other inputs from a user. Summary of the Invention
[0005] The framework ranks multiple hypotheses generated by one or more ASR engines for each input speech utterance. The framework jointly achieves ASR improvement and NLU. It utilizes NLU - related knowledge to facilitate the ranking of competing hypotheses and outputs the highest - ranked hypothesis as an improved ASR result along with the NLU result of the speech utterance. The NLU result includes an intent detection result and a slot - filling result.
[0006] The framework includes a deep feed - forward neural network that can extract features from the hypotheses. The features are fed into the input layer, and the same - type features from different hypotheses are concatenated together to facilitate learning. At least two projection layers are applied to each feature type via a shared projection matrix to project the features from each hypothesis into a smaller space, and then a second conventional projection layer projects the smaller space from all hypotheses into a compressed representation. If the number of features for each hypothesis of a feature type is less than a threshold, the projection layer can be bypassed, and the corresponding features extracted from all hypotheses are directly fed into an internal layer.
[0007] When a speech utterance is decoded by different types of speech recognition engines, joint modeling of speech recognition ranking and intent detection is used.
[0008] The framework extracts NLU-related features. Trigger features are extracted based on the slot filling results for each hypothesis. The BLSTM feature represents an intent-sensitive utterance embedding, which is obtained by concatenating the final states of the forward and backward LSTM RNNs in the decoder of the NLU module during the processing of each hypothesis.
[0009] The framework predicts the ranking of competing hypotheses and also generates the NLU results for a given speech utterance. The NLU results include intent detection results and slot filling results. During feature extraction, each input hypothesis is processed by the NLU module to obtain the NLU results. Then the framework predicts the hypothesis with the highest ranking and outputs the intent detection and slot filling results associated with that hypothesis as the NLU results for the input speech utterance. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a block diagram of a speech-based system;
[0011] Figure 2 is an illustration of a framework for ranking hypotheses generated by an ASR engine for a speech utterance;
[0012] Figure 3 is a block diagram of a stand-alone encoder / decoder natural language understanding (NLU) module;
[0013] Figure 4 is a flowchart of a process for performing speech recognition to operate a computerized system;
[0014] Figure 5 is a block diagram of a system configured to perform speech recognition;
[0015] Figure 6 is a flowchart of training speech recognition. DETAILED DESCRIPTION
[0016] As required, detailed embodiments of the present invention are disclosed herein; however, it should be understood that the disclosed embodiments are merely illustrative of the present invention, and the present invention may be embodied in various and alternative forms. The figures are not necessarily drawn to scale; some features may be enlarged or reduced to show details of particular components. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but merely as a representative basis for teaching those skilled in the art to employ the present invention in different ways.
[0017] The term "substantially" may be used in this disclosure to describe disclosed or claimed embodiments. The term "substantially" may modify a value or relative characteristic disclosed in this disclosure. In such cases, "substantially" may indicate that the modified value or relative characteristic is within ±0%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5% or 10% of the value or relative characteristic.
[0018] To facilitate an understanding of the principles of the embodiments disclosed herein, reference is now made to the accompanying drawings and description in the following written specification. The reference is not intended to limit the scope of the subject matter. This disclosure also includes any changes and modifications to the illustrated embodiments, and includes further applications of the principles of the disclosed embodiments that would typically occur to those skilled in the art associated with this disclosure.
[0019] A speech recognition system may use a trained speech recognition engine to convert recorded spoken input from a user into digital data suitable for processing in a computerized system. The speech engine may perform natural language understanding techniques to identify the words spoken by the user and extract semantic meaning from the words to control the operation of the computerized system.
[0020] In some cases, a single speech recognition engine may not be optimal for recognizing speech from a user when the user performs different tasks. Some solutions attempt to combine multiple speech recognition systems to improve the accuracy of speech recognition, including selecting low-level outputs from acoustic models, different speech recognition models, or selecting an entire set of outputs from different speech recognition engines based on a predetermined ranking process. However, the low-level combination of outputs from multiple speech recognition systems does not preserve high-level language information. In other embodiments, multiple speech recognition engines generate complete speech recognition results, but the determination process of which speech recognition result to select from the outputs of multiple speech recognition engines is also a challenging problem. Therefore, it would be beneficial to improve a speech recognition system to increase the accuracy of selecting a speech recognition result from a set of candidate speech recognition results from multiple speech recognition engines.
[0021] As used herein, the term "speech recognition engine" refers to a data model and executable program code that enables a computerized system to identify spoken words from an operator based on recorded audio input data of spoken words received via a microphone or other audio input device. A speech recognition system typically includes a lower-level acoustic model that identifies individual sounds of human speech in a sound recording and a higher-level language model that identifies words and sentences based on sequences of sounds from a predefined language. Speech recognition engines known in the art typically implement one or more statistical models, such as hidden Markov models (HMMs), support vector machines (SVMs), trained neural networks, or other statistical models that use multiple trained parameters applied to feature vectors of input data corresponding to human speech to generate statistical predictions for the recorded human speech. The speech recognition engine uses various signal processing techniques known in the art, for example, to generate feature vectors that extract attributes ("features") of the recorded speech signal and organize the features into one-dimensional or multi-dimensional vectors that can be processed using the statistical model to identify various speech segments, including individual words and sentences. The speech recognition engine can produce results for speech inputs corresponding to individual spoken phonemes and more complex sound patterns, including spoken words and sentences containing sequences of related words.
[0022] As used herein, the term "speech recognition result" refers to the machine-readable output generated by a speech recognition engine for a given input. The result can be, for example, text encoded in a machine-readable format or a collection of additional encoded data used as input to control the operation of an automated system. Due to the statistical nature of the speech recognition engine, in some configurations, the speech engine generates multiple potential speech recognition results for a single input. The speech engine also generates a "confidence score" for each speech recognition result, where the confidence score is a statistical estimate of the likelihood that each speech recognition result is accurate based on the trained statistical model of the speech recognition engine. As described in more detail below, a hybrid speech recognition system uses speech recognition results produced by multiple speech recognition engines to generate additional hybrid speech recognition results and ultimately produces at least one output speech recognition result based on the multiple previously produced speech recognition results. As used herein, the term "candidate speech recognition result" or more simply "candidate result" refers to a speech recognition result that is a candidate for the final speech recognition result from a hybrid speech recognition system that produces multiple candidate results and selects only a subset (typically one) of the results as the final speech recognition result. In various embodiments, candidate speech recognition results include both speech recognition results from general and domain-specific speech recognition engines and hybrid speech recognition results generated by the system using words from multiple candidate speech recognition results.
[0023] As used herein, the term "general speech recognition engine" refers to a speech recognition engine that is trained to recognize a wide range of speech from natural human languages such as English, Chinese, Spanish, Hindi, etc. The general speech recognition engine generates speech recognition results based on a large vocabulary of words and a language model that is trained to broadly cover language patterns in natural languages. As used herein, the term "domain-specific speech recognition engine" refers to a speech recognition engine that is trained to recognize speech inputs in a specific usage area or "domain", where the speech inputs typically include a different vocabulary and potentially different expected grammatical structures from the broader natural language. The domain-specific vocabulary typically includes some terms from the broader natural language, but may include a more narrow overall vocabulary, and in some cases includes specialized terms that are well-known in the specific domain but not officially recognized as part of the official vocabulary in the natural language. For example, in a navigation application, domain-specific speech recognition may recognize terms for roads, towns, or other geographical designations that are not typically recognized as proper names in the more general language. In other configurations, the specific domain uses a set of jargon that is useful for the specific domain but may not be well recognized in the broader language. For example, pilots officially use English as the language for communication, but also use a large number of domain-specific jargon words and other abbreviations that are not part of standard English.
[0024] As used herein, the term "trigger pair" refers to two words, each of which can be a word (e.g., "play") or a predetermined category (e.g., <song name>) that represents a sequence of words (e.g., "Poker Face") that falls within a predetermined category, such as the correct name of a song, a person, a location name, etc. When the words in the trigger pair appear within the words in the utterance text content of the speech recognition result in a specific order, there is a high correlation between the appearance of the earlier word A in the trigger pair A → B observed in the audio input data and the appearance of the later word B. As described in more detail below, after identifying a set of trigger pairs via a training process, the appearance of trigger word pairs in the text of candidate speech recognition results forms a part of the feature vector for each candidate result, and the ranking process uses the feature vector to rank different candidate speech recognition results.
[0025] This disclosure provides an improvement over our prior work, as identified by U.S. Patent No. 10,170,110, which used a deep feed-forward neural network to re-rank multiple hypotheses generated by an automatic speech recognition (ASR) engine for a speech utterance, the entire disclosure of which is incorporated herein by reference. In the prior work, the ranking framework utilized a relatively simple neural network structure that directly extracted features from the ASR results. Here, we refine the neural network structure of the ranking framework and enhance the ranking framework with NLU information in at least two aspects. In one aspect, NLU information (e.g., slot / intent information) is included to facilitate the ranking of ASR hypotheses. The NLU information fed into the neural network includes not only those computed from the ASR information, but also involves NLU-related features such as slot-based trigger features and semantic features representing utterance embeddings sensitive to slots / intents. The framework also jointly trains the ranking task with intent detection, aiming to use the intent information to help distinguish hypotheses. In another aspect, the framework not only outputs the top-ranked hypothesis as the new ASR result, but also outputs NLU results (i.e., slot filling results and intent detection results). Such that in a spoken dialogue system (SDS), the dialogue management component can directly perform subsequent processing based on the output of the proposed ranking framework, leading to the convenient application of the framework in the SDS. Experimental data was collected from an in-vehicle infotainment system, which ranked competing hypotheses generated by three different ASR engines. The experimental results were encouraging, where the system demonstrated the effectiveness of the proposed ranking framework. The experiments also showed that both incorporating NLU-related features and jointly training with intent detection increased the accuracy of the ranking of ASR hypotheses.
[0026] Improvements to ASR can be made in different directions, such as refining the acoustic / language model and adopting an end-to-end mode. Among these directions, post-processing the hypotheses generated by an (multiple) ASR engine has become a popular choice, mainly because it is much more convenient to apply linguistic knowledge to ASR hypotheses than to the decoding search space. Some post-processing methods construct certain confusion networks from ASR hypotheses and then distinguish competing words with the aid of acoustic / linguistic knowledge. Many prior works have used various advanced language models or discriminative models to re-score and rank ASR hypotheses. Ranking methods based on pairwise classification have also been proposed using classifiers based on support vector machines or neural network encoders. From the aspect of knowledge utilization, prior ASR methods only utilized limited linguistic knowledge, mainly modeling word sequences or directly extracting features from word sequences. Here, NLU information such as slots and intents can be shown to improve ASR.
[0027] The present disclosure illustrates a new neural network framework for ranking multiple hypotheses for an utterance. The framework uses all competing hypotheses as inputs and predicts their rankings simultaneously, rather than scoring each hypothesis one by one before ranking or comparing two hypotheses at a time. The framework leverages NLU knowledge to facilitate ranking by modeling using slot / intent-related features and by jointly training using intent detection.
[0028] Figure 1 An in-vehicle information system 100 is depicted, which includes a display such as a head-up display (HUD) 120, or one or more console LCD panels 124, one or more input microphones 128, and one or more output speakers 132. The LCD display 124 and the HCD 120 generate visual output responses from the system 100 based at least in part on voice input commands received by the system 100 from an operator or other passengers of the vehicle. A controller 148 is operatively connected to each of the components in the in-vehicle information system 100. In some embodiments, the controller 148 is connected to or incorporates additional components, such as a global positioning system (GPS) receiver 152 and a wireless network device 154 (such as a modem), to provide navigation and communication with external data networks and computing devices.
[0029] In some operating modes, the vehicle infotainment system 100 operates independently, while in other operating modes, the vehicle infotainment system 100 interacts with a mobile electronic device 170 such as a smartphone, a tablet computer, a laptop computer, or other electronic devices. The vehicle infotainment system communicates with the smartphone 170 using a wired interface (such as USB) or a wireless interface (such as Bluetooth). The vehicle infotainment system 100 provides a voice recognition user interface that enables an operator to control the smartphone 170 or another mobile electronic communication device using voice commands, which reduces distraction when operating the vehicle. For example, the vehicle infotainment system 100 provides a voice interface that enables a passenger in the vehicle (such as the vehicle operator) to make a phone call or send a text message using the smartphone 170 without requiring the operator / passenger to hold or look at the electronic device, the smartphone 170. In some embodiments, the vehicle system 100 provides a voice interface to the electronic device 170 such that the electronic device can launch an application on the smartphone 170 and then navigate the application and input data into the application based on the voice interface. In other embodiments, the vehicle system 100 provides a voice interface to the vehicle such that the operation of the vehicle can be adjusted based on the voice interface. For example, the voice interface can adjust the driving level (powertrain operation, transmission operation, and chassis / suspension operation) such that the vehicle switches from a comfort operation mode to a sport operation mode. In other embodiments, the smartphone 170 includes various devices such as GPS and wireless network devices that supplement or replace the functions of the devices housed in the vehicle.
[0030] The microphone 128 generates audio data based on a voice input received from the vehicle operator or another vehicle passenger. The controller 148 includes: hardware such as a microprocessor, a microcontroller, a digital signal processor (DSP), a single instruction multiple data (SIMD) processor, an application specific integrated circuit (ASIC), or other computing systems that process audio data; and software components that convert an input signal from the microphone 128 into audio input data. As will be explained below, the controller 148 uses at least one general-purpose voice recognition engine and at least one domain-specific voice recognition engine to generate candidate voice recognition results based on the audio input data, and the controller 148 further uses a ranker and a natural language understanding module to improve the accuracy of the final voice recognition result output. Additionally, the controller 148 includes hardware and software components that enable the generation of synthetic voice or other audio output through the speaker 132.
[0031] The vehicle information system 100 uses an LCD panel 124, an HCD 120 projected onto the windshield 102, and gauges, indicator lights, or additional LCD panels located in the instrument panel 108 to provide visual feedback to the vehicle operator. When the vehicle is in motion, the controller 148 optionally deactivates the LCD panel 124 or displays only a simplified output through the LCD panel 124 to reduce distraction to the vehicle operator. The controller 148 uses the HUD 120 to display visual feedback so that the operator can observe the environment around the vehicle while receiving the visual feedback. The controller 148 typically displays simplified data in an area of the HUD 120 corresponding to the peripheral vision of the vehicle operator to ensure an unobstructed view of the road and environment around the vehicle for the vehicle operator.
[0032] As described above, the HUD 120 displays visual information on a portion of the windshield 102. As used herein, the term "HUD" generally refers to various different head-up display devices, including but not limited to a combined head-up display (CHUD) including a separate combiner element, etc. In some embodiments, the HUD 120 displays monochromatic text and graphics, while other HUD embodiments include a multi-color display. Although the HUD 120 is depicted as being displayed on the windshield 102, in alternative embodiments, the head-up unit is integrated with glasses, a helmet visor, or markings worn by the operator during operation.
[0033] The controller 148 includes one or more integrated circuits configured as one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processor (DSP), or any other suitable digital logic device. The controller 148 also includes a memory, such as solid state memory (e.g., random access memory (RAM), read only memory (ROM), etc.), a magnetic data storage device, or other structures for storing programming instructions for the operation of the vehicle information system 100.
[0034] During operation, the vehicle information system 100 receives input requests from a plurality of input devices, including voice input commands received through the microphone 128. In particular, the controller 148 receives audio input data corresponding to the voice from the user via the microphone 128.
[0035] The controller 148 includes one or more integrated circuits configured as a central processing unit (CPU), microcontroller, field programmable gate array (FPGA), application specific integrated circuit (ASIC), digital signal processor (DSP), or any other suitable digital logic device. The controller 148 is also operatively connected to a memory 160 that includes a non-volatile solid state or magnetic data storage device or a volatile data storage device such as random access memory (RAM), and the memory 160 stores programming instructions for the operation of the in-vehicle information system 100. The memory 160 stores model data and executable program instruction codes to implement a plurality of speech recognition engines 162, a feature extractor 164, a natural language understanding unit 166, and a deep neural network ranker 168. Although Figure 1 embodiments include elements stored in the memory 160 of the system 100 within a motor vehicle, in some embodiments, an external computing device such as a network-connected server implements some or all of the features depicted in the system 100. Accordingly, those skilled in the art should recognize that any reference to the operation of the system 100 including the controller 148 and the memory 160 should further include the operation of server computing devices and other distributed computing components in alternative embodiments of the system 100.
[0036] In Figure 1 embodiments, the feature extractor 164 is configured to generate word sequence features having a plurality of digital elements corresponding to the content of each candidate speech recognition result, the candidate speech recognition results including speech recognition results generated by one of the speech recognition engines 162 or hybrid speech recognition results that combine words from two or more of the speech recognition engines 162. Additionally, the feature extractor 164 is further configured to generate natural language understanding features via the natural language understanding unit 166. The feature extractor 164 generates word sequence features that include elements of any one or combination of the following features: (a) trigger pairs, (b) confidence scores, and (c) individual word-level features including a bag of words with decay features.
[0037] Figure 2 is an illustration of a framework 200 for ranking hypotheses generated by an ASR engine for a speech utterance. Here, the new framework uses a deep feedforward neural network with NLU-related features to rank competing hypotheses generated by one or more ASR engines, and not only outputs the highest-ranked hypothesis as the new ASR result, but also outputs the corresponding NLU results (i.e., intent detection results and slot filling results).
[0038] The proposed framework is a deep feedforward neural network that receives as input the N (in our system, N = 10) competing hypotheses generated by one or more ASR engines for a speech utterance and predicts a ranking result for these hypotheses, optionally also predicting an intent detection result. The overall structure is illustrated as Figure 1 shown in
[0039] Features extracted from the hypotheses are fed into the input layer, while the same type of features from different hypotheses are concatenated together to facilitate learning. For one type of feature, hundreds or more features can be extracted from each hypothesis. We use two projection layers to handle such features. For each type of feature, first, the features from each hypothesis are projected into a smaller space using a shared projection matrix, and then a second conventional projection layer projects these spaces from all hypotheses into an even more compressed representation. Then, the representations achieved for each type of feature are concatenated and fed into an internal layer, which is a fully connected feedforward layer. In cases where each hypothesis can generate only one or a few features (such as confidence score features) for a certain feature type, we simply omit the projection layer for that feature type and directly feed the corresponding features extracted from all hypotheses into the internal layer by concatenating these features with the second projection layer for other feature types.
[0040] The output layer contains two parts, a main part that predicts the ranking result for the input hypotheses and an optional other part that predicts the intent detection result. The main part contains N output nodes that correspond to the N input hypotheses in the same order. Softmax activation is used to generate the output values, and then the hypotheses are ranked accordingly based on these values. To effectively rank the hypotheses, we use soft target values (instead of one - hot values) similar to those described in U.S. Patent No. 10,170,110 for training, and the soft target values are as follows:
[0041]
[0042] where di is the Levenshtein distance between the i - th hypothesis and the reference utterance. Using this definition, the target distribution preserves the ranking information of the input hypotheses, generating higher scores for the output nodes if the corresponding input hypothesis contains fewer ASR errors. The output distribution is approximated to the target distribution by minimizing the Kullback - Leibler divergence loss.
[0043] The "intent output" part in the output layer is optional. When intent is available in our experiments, jointly training the ranking task and intent detection can be beneficial because intent information can help distinguish between hypotheses. For intent-related outputs, nodes corresponding to possible intents are assigned one-hot target values (1 for the reference intent and 0 for other intents), and are trained using cross-entropy loss. When using the intent output (as in our system), we jointly train the network so as to backpropagate the costs from the ASR ranking part and the intent-related part to lower layers.
[0044] In this system, there are four main types of features extracted from each input hypothesis: trigger features, bag-of-words (BOW) features, bidirectional long short-term memory (BLSTM) features, and confidence features.
[0045] Trigger features are used to model long-distance / flexible-distance constraints. Similar to U.S. Patent No. 10,170,110, we define a trigger as a pair of linguistically significant related units in the same utterance, where a linguistic unit can be a word or a slot (i.e., <song name>). A trigger pair (e.g., "play" <song name>) captures the dependency between the two units, regardless of how far apart they are in an utterance. Given a text corpus collected in the domain of interest, we first process it by replacing corresponding texts with slots (e.g., using <song name> to replace "Poker Face"). Then, we calculate the mutual information (MI) scores of all possible trigger pairs based on Equation 2 below for,
[0046]
[0047] where refers to the event where A / B does not appear in the utterance. Then, the top n trigger pairs with the highest MI scores are selected as trigger features.
[0048] Next, the feature extraction of trigger features is extended by robustly identifying slots in each hypothesis using the NLU module to extract word / slot trigger pairs. When extracting trigger features from a hypothesis, an independent NLU module is used to detect slots in that hypothesis. If a trigger pair appears in the hypothesis, the value of the trigger feature is 1, otherwise 0.
[0049] BOW features include the definition from U.S. Patent No. 10,170,110. Given a dictionary, Equation (3) below is used to calculate the vector of BOW features for each hypothesis,
[0050]
[0051] where K is the number of words in the hypothesis, and is the one-hot representation of the i-th word in the hypothesis. is a decay factor, which is set to 0.9.
[0052] The BLSTM features are new and are related to the NLU properties of each hypothesis. Note that, as Figure 4 described in, the NLU module used in the extraction of the trigger features utilizes a bidirectional LSTM RNN to encode each hypothesis, where the last states of both the forward recurrent neural network (RNN) and the backward recurrent neural network cover the information of the entire hypothesis. We concatenate the last two states into a sentence embedding vector, called the BLSTM feature. Since the NLU module is a joint model for intent detection and slot filling, the BLSTM feature is intent-sensitive.
[0053] The confidence features include the definition from U.S. Patent No. 10,170,110. The sentence-level confidence scores are assigned to each hypothesis by the ASR engine, and the hypotheses are directly fed into the internal layer. Note that various ASR engines may use different distributions to produce the confidence scores. When the input hypotheses are generated by different ASR engines, the application of a linear regression method can be used to align the confidence scores into the same space, and then the aligned scores are used as the confidence features.
[0054] Figure 3 is a block diagram of an independent natural language understanding (NLU) module 300 that utilizes an encoder / decoder neural network architecture. The independent NLU module used in feature extraction and subsequent evaluation is implemented using state-of-the-art methods that jointly model slot filling and intent detection. The module can adopt an RNN-based encoder-decoder structure, using LSTM as the RNN cell. The pre-trained word embedding vectors for each input word can be fed into the encoder 302. When a predefined list of names is available, we further enhance the input vector by appending the named entity feature 304. The aim is to use the added name information to facilitate learning, especially for cases where the size of the training data is limited and many names appear only a few times or are not visible therein. For the named entity features, each of them corresponds to a list of names and is set to 1 if the input word is part of the names in the list, otherwise set to 0. For Figure 3In the example shown, the word "Relax" is both a song name and a playlist name. Thus, in the named entity vector 304, the corresponding features for both are set to 1. Using this information and the context knowledge captured by the RNN, the NLU module can identify "Relax" as a playlist name even if the name "Relax" has not been seen in the training data. Based on the hidden representation generated by the encoder 302 for a given input statement, the decoder 306 then generates the NLU result for the given statement, predicting the intent of the statement and detecting the slots in the statement.
[0055] An in-vehicle infotainment system that utilizes multiple different types of ASR engines is used to evaluate a framework for ranking hypotheses generated by multiple engines. The in-vehicle infotainment system includes a vehicle control system such as a driver assistance system. A speech training / adaptation / test set is recorded in the car, the speech training / adaptation / test set having relatively low noise conditions, being from a gender-balanced set of multiple speakers, and containing 9,166, 1,975, and 2,080 utterances respectively. Each utterance is decoded by two domain-specific ASR engines (trained on separate in-domain datasets using grammar and statistical language models respectively) and a general cloud engine. The three engines (two domain-specific ASR engines and one cloud ASR engine) have complementary strengths for decoding. The process includes feeding the top best hypotheses from each engine into the proposed framework for ranking, and clearing the space for additional hypotheses allowed in the input layer by setting the relevant features to 0. Most of the names involved in the system are from 16 name lists, some of which are large (e.g., the song name list contains 5,232 entries). Although other tagging methods can be used, 40 slot tags (including phone numbers, frequencies, etc., and items without a predefined list) are created following the IOB pattern, and 89 different intents (e.g., "tune radio frequency") are used.
[0056] The standalone NLU module is trained using reference statements. GloVe: Global Vectors for Word Representation version 100d is used as the word embedding for each input word. The named entity vector is constructed based on a given list of names. The NLU module is trained, however, without using an attention mechanism due to the observed limited benefits of the attention mechanism in terms of ASR hypotheses and for efficiency considerations.
[0057] Trigger features are selected based on an additional text set of 24,836 in-domain statements from the infotainment data, and 850 trigger features are utilized. In terms of BOW features, the lexicon is limited to the most frequent words in 90% of the training references, as well as entries for out-of-vocabulary words, and confidence score alignment is applied.
[0058] In the ranking framework, each shared projection matrix projects the corresponding features into a space of 50 nodes. A layer of size 50 * 10 * 3 is provided (with 500 nodes for trigger, BOW, and BLSTM features respectively), and this layer is further projected into a smaller second projection layer (200 * 3 nodes). Then, this second projection layer is concatenated with 10 confidence features to be fed into the input layer. Four internal layers are used (with 500, 200, 100, and 100 nodes respectively), and the internal layers use activation functions such as the ReLU activation function and batch normalization applied to each layer. We adopt Adam optimization to train the model in batches. When the loss of the validation set fails to improve in the last 30 iterations, early stopping is performed. Then, the model that achieves the best performance on the validation data is used in the evaluation. And the hyperparameters are selected empirically.
[0059] In an exemplary standalone NLU module, it is beneficial to feed named entity features into the encoder. In one system, the intent detection error rate is reduced from 9.12% to 5.17%, while the F1 score of slot filling for the test reference is increased from 64.55 to 90.68. This indicates that when using a large name list and limited training data, the introduced name information effectively alleviates the difficulty of learning.
[0060] For the ranking framework, we first train the framework using only the ASR ranking output, called the standalone ASR framework, and then use both the ASR ranking and the intent output to train the joint framework. The evaluation results of the two frameworks are included in Table 1.
[0061]
[0062] In Table 1, "Oracle Hypothesis" and "Highest Scoring Hypothesis" refer to the hypotheses with lower word error rate (WER) and the highest (alignment) confidence score among competing hypotheses respectively. "+NLU" represents the process of applying a standalone NLU module to the hypothesis to obtain NLU results (i.e., F1 score of slot filling and intent detection error rate). Both the WER and the NLU results are evaluated. For the ranking framework, during feature extraction, each input hypothesis is processed by the NLU module to obtain its NLU result. When the framework predicts the highest ranked hypothesis, it also retrieves the NLU result associated with that hypothesis. Note that for the joint framework, the intent-related output also predicts the intent. However, it should be noted that the predicted intent performs worse than the intent assigned to the highest ranked hypothesis, which may be due to the confusion introduced by competing hypotheses, so the latter is selected as the intent result.
[0063] Table 1 shows that, compared with the "highest-scoring hypothesis" baseline (i.e., the performance of ranking hypotheses based only on the alignment confidence scores), the standalone ASR framework alone brings a 6.75% relative decrease in WER, which is better than the performance of each individual engine. The joint framework expands the benefit to a relative WER reduction of 11.51%. Similar improvements were achieved for the NLU results.
[0064] The experiments also show that for the framework, it is important to use the proposed soft target values for ranking. For example, when using one-hot values to replace the soft target values, the WER obtained by the joint model rises to 7.21%. It was also observed that all four types of features are beneficial for in-vehicle infotainment data, and removing each one will result in worse performance. For example, removing the slot-based trigger feature from the joint framework increases the WER of the resulting model to 7.28%.
[0065] In system 100, each trigger pair stored in feature extractor 164 includes a predetermined set of two linguistic items, each of the two linguistic items can be a word or a slot detected by the independent NLU module 300. The two items of each trigger pair have previously been identified as having a strong correlation in statements from the training corpus, which represents a text copy of the expected speech input. In the speech input, there is a strong statistical likelihood that the first trigger item is followed by the second trigger item in the trigger pair, although these items may be separated by an indeterminate number of intervening words in different speech inputs. Thus, if the speech recognition result includes the trigger items, the likelihood that those trigger words in the speech recognition result are accurate is relatively high due to the statistical correlation between the first trigger item and the second trigger item. In system 100, statistical methods known in the art are used to generate trigger items based on mutual information scores. Memory 160 stores a predetermined set of N trigger pair elements in the feature vector based on the set of trigger items having high mutual information scores, the feature vector corresponding to the trigger pairs having a high correlation level between the first item and the second item. As described below, the trigger pairs provide additional features of the speech recognition result to neural network ranker 168, which enables neural network ranker 168 to rank the speech recognition result using additional linguistic knowledge beyond the word sequence information present in the speech recognition result.
[0066] The confidence score feature corresponds to a numeric confidence score value generated in conjunction with each candidate speech recognition result by the speech recognition engine 162. For example, in one configuration, a value in the range (0.0, 1.0) indicates that the speech recognition engine places the accuracy of a particular candidate speech recognition result at a probabilistic confidence level from the lowest confidence (0.0) to the highest confidence (1.0). Each of the hybrid candidate speech recognition results generated by one or more speech recognition engines is assigned a confidence score. When a speech recognition engine generates a speech recognition result, it also assigns a confidence score to it.
[0067] In system 100, the controller 148 also normalizes and whitens the confidence score values of the speech recognition results generated by different speech recognition engines to generate final feature vector elements that include the normalized and whitened confidence scores that are uniform among the outputs of the plurality of speech recognition engines 162. The controller 148 uses a normalization process to normalize the confidence scores from different speech recognition engines and then uses a prior art whitening technique to whiten the normalized confidence score values based on the mean and variance estimated for the training data. In one embodiment, the controller 148 uses a linear regression process to normalize the confidence scores between different speech recognition engines. The controller 148 first subdivides the confidence score range into a predetermined number of subsections or "bins", such as 20 unique bins for two speech recognition engines A and B. Then, the controller 148 identifies the actual accuracy rates for the various speech recognition results corresponding to each score bin based on the observed speech recognition results and the actual underlying inputs used during the training process prior to this process. The controller 148 performs a clustering operation on the confidence scores within a predetermined numeric window around the "edges" that separate the bins of each set of results from different speech recognition engines and identifies the average accuracy score corresponding to each edge confidence score value. The "edge" confidence scores are uniformly distributed along the confidence score range of each speech recognition engine and provide a predetermined number of comparison points to perform a linear regression that maps the confidence scores of the first speech recognition engine to the confidence scores of another speech recognition engine with similar accuracy rates.
[0068] The controller 148 uses the identified accuracy data for each edge score to perform a linear regression mapping that enables the controller 148 to convert a confidence score from a first speech recognition engine into another confidence score value corresponding to a peer confidence score from a second speech recognition engine. Mapping a confidence score from a first speech recognition engine to another confidence score from another speech recognition is also referred to as score alignment processing, and in some embodiments, the controller 148 determines the alignment of the confidence score from the first speech recognition engine with the confidence score of the second speech recognition engine using the following equation:
[0069]
[0070] where x is the score from the first speech recognition engine, x’ is the equivalent value within the confidence score range of the second speech recognition engine, the value x is the estimated accuracy score for different edge values that are different from the value e i and e i+1 and are closest to the value of the first speech recognition engine x (e.g., the estimated accuracy scores for edge values 20 and 25 around the confidence score 22), the value e i ’ and e i+1 ’ correspond to the estimated accuracy scores for the same relative edge values for the second speech recognition engine.
[0071] In some embodiments, the controller 148 stores the result of the linear regression in the feature extractor 164 in the memory 160 as a look-up table or other suitable data structure to enable efficient normalization of the confidence scores between different speech recognition engines 162 without having to regenerate the linear regression for each comparison.
[0072] The controller 148 also uses a feature extractor 164 to identify word-level features in the candidate speech recognition results. The word-level features correspond to data that the controller 148 places in elements of a feature vector, where the elements of the feature vector correspond to characteristics of individual words in the candidate speech recognition results. In one embodiment, the controller 148 only identifies the presence or absence of words within a plurality of predetermined vocabularies, which corresponds to individual elements of a predetermined feature vector within each candidate speech recognition result. For example, if the word "street" appears at least once in the candidate speech recognition result, the controller 148 sets the value of the corresponding element in the feature vector to 1 during the feature extraction process. In another embodiment, the controller 148 identifies the frequency of each word, where "frequency" as used herein refers to the number of times a single word appears in the candidate speech recognition result. The controller 148 places the number of times the word appears in the corresponding element of the feature vector.
[0073] In yet another embodiment, the feature extractor 164 generates a "bag of words with decay features" for elements in the feature vector corresponding to each word in a predetermined vocabulary. As used herein, the term "bag of words with decay features" refers to a numeric score that the controller 148 assigns to each word in a predetermined vocabulary based on the number of occurrences and location of the word in the result for a given candidate speech recognition result. The controller 148 generates a bag of words with decay scores for each word in the predetermined vocabulary within the candidate speech recognition result and assigns a bag of words with a decay score of zero to those words in the vocabulary that do not appear in the candidate result. In some embodiments, the predetermined vocabulary includes a special entry for representing any out-of-vocabulary word, and the controller 148 also generates a single bag of words with a decay score for this special entry based on all out-of-vocabulary words within the candidate result. For a given word in the predetermined lexicon w i , the bag of words with decay scores can be expressed according to Equation 2 above or Equation 6 below,
[0074]
[0075] where is the set of locations in the candidate speech recognition result where the word w i appears, and the term is a predetermined numeric decay factor in the range (0, 1.0), e.g., set to 0.9 in an illustrative embodiment of the system 100.
[0076] Return reference Figure 1 , in Figure 1In an embodiment, the neural network ranker 168 is a trained neural network that includes a neuron input layer that receives a plurality of feature vectors corresponding to a predetermined number of candidate speech recognition results, and a neuron output layer that generates a rank score corresponding to each of the input feature vectors. Generally, a neural network includes a plurality of nodes referred to as "neurons". Each neuron receives at least one input value, applies a predetermined weighting factor to the input value (where different input values typically receive different weighting factors), and generates an output as the sum of the weighted inputs, where in some embodiments an optional bias factor is added to the sum. The exact weighting factor for each input and the optional bias value in each neuron are generated during a training process, which will be described in more detail below. The output layer of the neural network includes another set of neurons that are specifically configured with an "activation function" during the training process. The activation function is, for example, a sigmoid function or other threshold function that produces an output value based on an input from a final hidden layer of neurons in the neural network, where the exact parameters of the sigmoid function or threshold are generated during the training process of the neural network.
[0077] In Figure 1 In a specific configuration, the neural network ranker 168 may include a feedforward deep neural network. As is known in the art, a feedforward neural network includes neuron layers connected in a single direction from an input layer to an output layer, without any loops or "feedback" loops that connect neurons in one layer of the neural network to neurons in a previous layer of the neural network. A deep neural network includes neurons in at least one "hidden layer" (and typically more than one hidden layer), which are not exposed as the input layer or the output layer. For example, a plurality of neurons in k hidden layers may be used to connect the input layer to the output layer.
[0078] Consider Figure 2, in one embodiment of neural network 200, the input layer further includes projection layers 204A, 204B, which apply a predetermined matrix transformation to the selected sets of input feature vector elements 202A, 202B. The projection layers 204A, 204B each include two different projection matrices for word sequence features, such as trigger pair feature elements, BOW feature elements, BLSTM feature elements, and word-level feature elements. The projection layer 204 generates a simplified representation of the output of the input neurons in the input layer 202 because in most actual inputs, the feature vector elements for word sequence features are "sparse", which means that each candidate speech recognition result includes only a small number (if any) of trigger pair items encoded in the structure of the feature vector and a small number of words from a large overall word set (e.g., 10,000 words). The transformation in the projection layer 204 enables the remaining layers of the neural network to include fewer neurons while still generating useful ranking scores for the feature vector inputs of the candidate speech recognition results. In an illustrative embodiment, the P f for trigger word pairs and the P w for word-level features each project the corresponding input neurons into a smaller vector space, each of which has 200 elements, which results in a projection layer of 401 neurons for each of the n input feature vectors in the neural network ranker 168 (reserving one neuron for the confidence score feature).
[0079] During operation, system 100 uses microphone 128 to receive audio input data and uses multiple speech engines 162 to generate multiple candidate speech recognition results, which in some embodiments include hybrid speech recognition results that include words selected from two or more of the candidate speech recognition results. The controller 148 uses the feature extractor 164 to extract features from the candidate speech recognition results to generate feature vectors and provides the feature vectors to the neural network ranker 168 to generate an output score for each feature vector. Then, the controller 148 identifies the feature vector and the candidate speech recognition result corresponding to the highest ranking score, and the controller 148 uses the candidate speech recognition result corresponding to the highest ranking score among the multiple candidate speech recognition results as an input to operate the automation system.
[0080] Figure 4 depicts process 400 for performing speech recognition using multiple speech recognition engines and a neural network ranker to select a candidate speech recognition result. In the following description, a reference to process 400 performing a function or action refers to the controller operating on stored program instructions to perform the function or action associated with other components in the automation system. For illustrative purposes, in conjunction withFigure 1 Processing 400 is described with respect to system 100.
[0081] Processing 400 begins with system 100 generating multiple candidate speech recognition results using multiple speech recognition engines 162 (block 404). In system 100, a user provides spoken audio input to an audio input device such as microphone 128 (block 402). Controller 148 uses multiple speech recognition engines 162 to generate multiple candidate speech recognition results. As described above, in some embodiments, controller 148 uses words selected from the candidate speech recognition results of a domain-specific speech recognition engine to generate a hybrid candidate speech recognition result to replace words selected from the candidate speech recognition results of a general speech recognition engine. The speech recognition engines 162 also generate confidence score data that system 100 uses during the generation of feature vectors in processing 400.
[0082] Processing 400 continues when system 100 performs feature extraction to generate multiple feature vectors, each feature vector corresponding to one of the candidate speech recognition results (block 406). In system 100, controller 148 uses feature extractor 164 to generate feature vectors via word sequence features - the word sequence features including one or more of the trigger pairs, confidence scores, and word-level features described above - to generate feature vectors having Figure 2 the structure of feature vectors 202 in or additional similar structures for one or more word sequence features such as trigger pairs, confidence scores, and word-level features. In Figure 4 embodiments, controller 148 uses a bag of words with an attenuation measure for word-level feature elements of the feature vectors to generate word-level features.
[0083] Block 408 processes each speech recognition result using an independent NLU module. The NLU module performs two tasks, namely slot filling and intent detection. For a focused speech recognition result, the NLU module detects the slots contained therein and detects its intent. The NLU module also stores the last state of each direction of a bidirectional recurrent neural network (RNN) in an encoder to support subsequent steps of feature extraction.
[0084] Block 410 extracts NLU-related features based on the output of the NLU module for each speech recognition result. It extracts trigger features for the focused speech recognition result based on the word sequences and (multiple) slots detected in block 208. It also concatenates the two last states of the bidirectional RNN stored in the NLU module encoder in block 408 to construct a BLSTM feature for the focused speech recognition result.
[0085] When the controller 148 provides the feature vectors for the multiple candidate speech recognition results to the neural network ranker 168 as inputs in the inference process to generate multiple ranking scores corresponding to the multiple candidate speech recognition results, the process 400 continues (block 412). In one embodiment, the controller 148 uses the trained feed-forward deep neural network ranker 168 to generate multiple ranking scores at the output layer neurons of the neural network using the inference process. As described above, in another embodiment, the controller 148 uses the wireless network device 154 to transmit the feature vector data, the candidate speech recognition results, or an encoded version of the recorded audio speech recognition data to an external server, where a processor in the server performs a portion of the process 400 to generate ranking scores for the candidate speech recognition results.
[0086] In most cases, the controller 148 generates multiple candidate speech recognition results and corresponding feature vectors n that match the predetermined number n of feature vector inputs that the neural network ranker 168 is configured to receive during the training process. However, in some cases, if the number of feature vectors for a candidate speech recognition result is less than the maximum number n, the controller 148 generates "empty" feature vector inputs that are all zero values to ensure that all neurons in the input layer of the neural network ranker 168 receive an input. The controller 148 ignores the scores for the corresponding output layer neurons for each of the empty inputs, while the neural network in the ranker 168 generates scores for the non-empty feature vectors of the candidate search recognition results.
[0087] When the controller 148 identifies the candidate speech recognition result corresponding to the highest ranking score in the output layer of the neural network ranker 168, the process 400 continues (block 414). For example, each output neuron in the output layer of the neural network can generate an output value corresponding to a ranking score for one of the input feature vectors provided by the system 100 to a predetermined set of input neurons in the input layer. The controller 148 then identifies the candidate speech recognition result with the highest ranking score based on the index of the output neuron that produces the highest ranking score within the neural network.
[0088] When the controller 148 outputs (block 416) and uses the selected highest-ranked speech recognition result as the input from the user to operate the automation system (block 418), the process 400 continues. In Figure 1In the vehicle information system 100, the controller 148 operates various systems, including, for example, a vehicle navigation system that performs vehicle navigation operations in response to a voice input from a user using the GPS 152, the wireless network device 154, and the LCD display 124 or the head-up display 120. In another configuration, the controller 148 plays music through the audio output device 132 in response to a voice command. In yet another configuration, the system 100 uses the smartphone 170 or another network-connected device to make a hands-free phone call or transmit a text message based on a voice input from a user. Although Figure 1 illustrates an embodiment of a vehicle information system, other embodiments employ an automated system that uses audio input data to control the operation of various hardware components and software applications.
[0089] Although Figure 1 the vehicle information system 100 is depicted as an illustrative example of an automated system that performs speech recognition to receive and execute commands from a user, similar speech recognition processing can be implemented in other scenarios. For example, a mobile electronic device such as the smartphone 170 or other suitable device typically includes one or more microphones and a processor that can implement a speech recognition engine, a ranker, stored triggers, and other components for implementing a speech recognition and control system. In another embodiment, a home automation system uses at least one computing device to control HVAC and household appliances in a house, the computing device receiving a voice input from a user and performing speech recognition using multiple speech recognition engines to control the operation of various automated systems in the house. In each embodiment, the system is optionally configured to use a different set of domain-specific speech recognition engines that are customized for the specific applications and operations of different automated systems.
[0090] In Figure 1 the system 100 and Figure 4 the speech recognition processing, the neural network ranker 168 is a trained feed-forward deep neural network. The neural network ranker 168 is trained to perform the speech recognition processing described above before operating the system 100. Figure 5 illustrates an illustrative embodiment of a computerized system 500 that is configured to train the neural network ranker 168, and Figure 4 illustrates a training process 400 for generating the trained neural network ranker 168.
[0091] System 500 includes a processor 502 and a memory 504. The processor 502 includes, for example, one or more CPU cores optionally connected to a parallel hardware accelerator, which is designed to train neural networks in a time- and power-efficient manner. Examples of such accelerators include, for example, a GPU having computer shader units configured for neural network training, and a specially programmed FPGA chip or ASIC hardware dedicated to training neural networks. In some embodiments, the processor 502 further includes a cluster of computing devices that operate in parallel to perform neural network training processing.
[0092] The memory 504 includes, for example, non-volatile solid-state or magnetic data storage devices and volatile data storage devices such as random access memory (RAM), which store programming instructions for the operation of the system 500. In Figure 3 this configuration, the memory 504 stores data corresponding to training input data 506, a gradient descent trainer 508 for the neural network, a speech recognition engine 510, a feature extractor 512, a natural language understanding module 514, and a neural network ranker 516.
[0093] The training data 506 includes, for example, a large number of speech recognition results generated by the same speech recognition engine 162 in the system 100 for a large number of predetermined inputs, and the speech recognition results optionally include mixed speech recognition results. The training speech recognition result data also includes confidence scores for the training speech recognition results. For each speech recognition result, the training data further includes a Levenshtein distance metric that quantifies the difference between the speech recognition result and a predetermined actual speech input training data, and the predetermined actual speech input training data represents the "correct" result in the training process. The Levenshtein distance metric is an example of an "edit distance" metric because this metric quantifies the amount of change (editing) required to transform the speech recognition result from the speech recognition engine into the actual input for the training data. In a comparison metric, both the speech recognition result and the actual speech input training data are referred to as "strings" of text. For example, the edit distance quantifies the number of changes required to transform the speech recognition result string "Sally sells seashells by the seashore" into the corresponding correct actual training data string "Sally sells seashells by the sea".
[0094] In other cases, the Levenshtein distance metric is known in the art and has several properties, including: (1) the Levenshtein distance is always at least the difference in the sizes of two strings; (2) the Levenshtein distance is at most the length of the longer string; (3) the Levenshtein distance is zero if and only if the strings are equal; (4) if the strings are the same size, the Hamming distance is an upper bound on the Levenshtein distance; and (5) the Levenshtein distance between two strings is no greater than the sum of their Levenshtein distances to a third string (triangle inequality). Further, the Hamming distance refers to a measure of the minimum number of substitutions required to change one string into another or the minimum number of errors that can transform one string into another. Although for illustrative purposes, system 500 includes training data encoded using the Levenshtein distance, in alternative embodiments, other edit distance metrics are used to describe the difference between the training speech recognition results and the corresponding actual training inputs.
[0095] In Figure 5 the embodiment, the feature extractor 512 in the memory 504 is the same as the feature extractor 164 used in the above-described system 100. In particular, the processor 502 uses the feature extractor 512 to generate a feature vector from each of the training speech recognition results using one or more of the trigger pairs, confidence scores, and word-level features described above.
[0096] The gradient descent trainer 508 includes stored program instructions and parameter data for neural network training processing. The processor 502 performs the neural network training processing to train the neural network ranker 516 using the feature vectors generated by the feature extractor 512 based on the training data 506. As is known in the art, the gradient descent trainer includes a class of related training processes that train a neural network in an iterative process by adjusting the parameters within the neural network to minimize the difference (error) between the neural network output and a predetermined objective function (also referred to as a "metric" function). Although gradient descent training is well known in the art and not discussed in detail herein, the system 500 modifies the standard training process. In particular, the training process attempts to utilize the neural network to generate an output using the training data as input to minimize the error between the neural network's output and the expected target result from the predetermined training data. In some training processes, the target value typically specifies whether a given output is "correct" or "incorrect" in a binary manner, and such a target output from the neural network ranker provides a score indicating whether the feature vector input for the training speech recognition result is 100% correct or incorrect to some extent when compared to the actual input in the training data. However, in the system 500, the gradient descent trainer 508 uses the edit distance target data in the training data 506 as a "soft" target to more accurately reflect the correctness level of different training speech recognition results, which can include an error range affecting the ranking score within a continuous range rather than just being completely correct or incorrect.
[0097] The processor 502 uses the "soft" target data in the objective function to perform the training process using the gradient descent trainer 508. For example, Figure 3 the configuration uses a "softmax" objective function of the following form:
[0098]
[0099] where d i is the edit distance of the i-th speech recognition result from a reference copy of the given speech input. During the training process, the gradient descent trainer 508 performs a cost minimization process, where "cost" refers to the cross-entropy between the output value of the neural network ranker 516 and the target value generated by the objective function during each iteration of the training process. The processor 502 provides batches of samples to the gradient descent trainer 508 during the training process, such as a batch of 180 training inputs, each training input including different training speech recognition results generated by multiple speech recognition engines. The iterative process continues until the cross-entropy of the training set does not improve over the course of ten iterations, and the trained neural network parameters that produce the lowest total entropy from all the training data form the final trained neural network.
[0100] During the training process, the processor 502 shuffles the same input feature vectors among different groups of input neurons in the neural network ranker 516 during different iterations of the training process to ensure that the positioning of a particular feature vector in the input layer of the neural network does not introduce incorrect biases in the trained neural network. As described above in the inference process, if a particular training data set does not include a sufficient number of candidate speech recognition results to provide inputs to all neurons in the input layer of the neural network ranker 516, the processor 502 generates "empty" input feature vectors with zero-valued inputs. As is known in the art, the gradient descent optimization used in the training process involves numerical training parameters, and in one configuration of the system 500, adaptive moment estimation (Adam) optimization is used in the gradient descent trainer 508, and the hyperparameters of the gradient descent trainer 508 are = 0.001, = 0.9 and = 0.999.
[0101] Although Figure 5 FIG. depicts a particular configuration of the computerized device 500 that generates the trained neural network ranker, but in some embodiments, the same system that uses the trained neural network ranker in the speech recognition process is further configured to train the neural network ranker. For example, in some embodiments, the controller 148 in the system 100 is an example of a processor that can be configured to perform a neural network training process.
[0102] Figure 6 FIG. depicts a process 600 for performing speech recognition using multiple speech recognition engines and a neural network ranker to select candidate speech recognition results. In the following description, a reference to the process 600 that performs a function or action refers to an operation in which a processor executes stored program instructions to perform a function or action associated with other components in an automated system. For illustrative purposes, the process 600 is described in connection with Figure 5 the system 500 of FIG.
[0103] The process 600 begins with the system 500 generating a plurality of feature vectors corresponding to a plurality of training speech recognition results stored in the training data 506 (block 602). In the system 500, the processor 502 uses the feature extractor 512 to generate a plurality of feature vectors, where each feature vector corresponds to one of the training speech recognition results in the training data 506. As described above, in at least one embodiment of the process 600, the controller 502 generates each feature vector, which includes one or more of trigger pair features, confidence scores, and word-level features, where the word-level features include a bag of words with decaying features.
[0104] As part of the feature extraction and feature generation processing, in some embodiments, the controller 502 generates a feature vector structure that includes specific elements mapped to trigger pair features and word-level features. For example, as described above, in system 100, in some embodiments, the controller 502 generates a feature vector having a structure that corresponds to only a portion of the words observed in the training data 506, such as the 90% most frequently observed words, while the remaining 10% of the words with the lowest occurrence frequencies are not encoded into the structure of the feature vector. The processor 502 optionally identifies the most common trigger pair features and generates a structure for the most frequently observed trigger word pairs present in the training data 506. In embodiments in which the system 500 generates the structure of the feature vector during processing 600, the processor 502 stores the structure of the feature vector together with the feature extractor data 512 and, after the training process is complete, provides the structure of the feature vector to the automated system along with the neural network ranker 516, which uses the feature vector with the specified structure as the input to the trained neural network to generate a ranking score for the candidate speech recognition results. In other embodiments, the structure of the feature vector is determined a priori based on a natural language such as English or Chinese rather than specifically based on the content of the training data 506.
[0105] When the system 500 uses the gradient descent trainer 508 to train the neural network ranker 516 based on the feature vectors of the training speech recognition results from the training data 506 and the soft target edit distance data, the processing 600 continues (block 604). During the training process, the processor 502 uses multiple feature vectors corresponding to multiple training speech recognition results as the input to the neural network ranker and trains the neural network ranker 516 based on a cost minimization process between the multiple output scores generated by the neural network ranker during the training process and the objective function having the soft scores as described above, where the soft scores are based on a predetermined edit distance between the multiple training speech recognition results and the predetermined correct inputs for each of the training speech recognitions in the multiple speech recognition results. During processing 600, the processor 502 modifies the input weighting coefficients and neuron bias values in the input layer and hidden layer of the neural network ranker 516 and uses the gradient descent trainer 508 to iteratively adjust the parameters of the activation function in the neuron output layer.
[0106] After the training process is complete, the processor 502 stores the structure of the trained neural network ranker 516 and optionally stores the structure of the feature vector in embodiments in which the feature vector is generated based on the training data in the memory 504 (block 606). The stored structure of the neural network ranker 516 and the feature vector structure are then transferred to other automated systems, such as Figure 1System 100, which uses a trained neural network ranker 516 and a feature extractor 512 to rank multiple candidate speech recognition results during a speech recognition operation, and then operates the system based on the results (block 608).
[0107] It should be appreciated that the variations and other features and functions disclosed above, or their alternatives, can desirably be combined into many other different systems, applications, or methods. Those skilled in the art can then make various substitutions, modifications, variations, or improvements that are currently unforeseen or unanticipated, and such substitutions, modifications, variations, or improvements are also intended to be covered by the appended claims.
[0108] The program code embodying the algorithms and / or methods described herein can be distributed, either alone or in combination, in a variety of different forms as a program product. The program code can be distributed using a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to execute aspects of one or more embodiments. The inherently non-transitory computer-readable storage medium can include volatile and non-volatile and removable and non-removable tangible media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer-readable storage medium can further include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid state memory technologies, portable compact disc read-only memory (CD-ROM) or other optical storage, cassette tapes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be read by a computer. The computer-readable program instructions can be downloaded from the computer-readable storage medium to a computer, another type of programmable data processing apparatus, or another device, or can be downloaded via a network to an external computer or external storage device.
[0109] The computer-readable program instructions stored in the computer-readable medium can be used to cause a computer, other type of programmable data processing apparatus, or other device to function in a particular manner, such that the instructions stored in the computer-readable medium produce an article of manufacture including instructions for implementing the functions, acts, and / or operations specified in the flowchart or diagram. In certain alternative embodiments, the functions, acts, and / or operations specified in the flowchart and diagram can be reordered, serially processed, and / or processed simultaneously in a manner consistent with one or more embodiments. Additionally, any one of the flowcharts and / or diagrams can include more or fewer nodes or boxes than those consistent with one or more embodiments.
[0110] While the invention has been illustrated by the description of various embodiments, and while these embodiments have been described in considerable detail, the applicant does not intend to restrict or in any way limit the scope of the appended claims to such details. Additional advantages and modifications will be readily apparent to those skilled in the art. Accordingly, the invention in its broader aspects is not limited to the specific details, representative apparatus and methods, and illustrative examples shown and described. Thus, departures may be made from such details without departing from the spirit or scope of the general inventive concept of the invention.
Claims
1. A method for speech recognition in a system, executed by a controller, comprising: Parsing a plurality of candidate speech recognition results from a speech input; Receiving, from a first speech recognition engine, a first plurality of feature vectors from each of the plurality of candidate speech recognition results, the first plurality of feature vectors including a first confidence score; Receiving, from a second speech recognition engine different from the first speech recognition engine, a second plurality of feature vectors from each of the plurality of candidate speech recognition results, the second plurality of feature vectors including a second confidence score lower than the first confidence score; Extracting NLU results from each of the plurality of candidate speech recognition results based on natural language understanding (NLU) information; Compressing the first plurality of feature vectors and the second plurality of feature vectors to the shared projection layer via the shared projection layer based on the NLU results and NLU-related features, wherein the NLU-related features include slot-based trigger features and semantic features representing utterance embeddings sensitive to slots / intents; Further compressing the shared projection layer to a second projection layer based on the NLU results and NLU-related features; Associating a ranking score with each of the plurality of candidate speech recognition results via a neural network ranker, the ranking score being based on the first plurality of feature vectors and the second plurality of feature vectors and the NLU results of each of the plurality of candidate speech recognition results, wherein the neural network ranker raises the second confidence score to be greater than the first confidence score based on the NLU-related features; Selecting, from the plurality of candidate speech recognition results, a speech recognition result associated with the ranking score having the highest value; And Operating the system using the speech recognition result selected from the plurality of candidate speech recognition results and corresponding to the highest ranking score as an input.
2. The method according to claim 1, wherein the neural network ranker is a deep feedforward neural network ranker.
3. The method according to claim 1, wherein the compression is performed via a shared projection matrix.
4. The method according to claim 3, further comprising bypassing, by the controller, the shared projection layer and the second projection layer in response to the first plurality of feature vectors and the second plurality of feature vectors being smaller than a threshold size, such that the second plurality of feature vectors are directly fed to the neural network ranker, wherein for each hypothesis, the threshold size of the feature vectors is less than 2 features.
5. The method according to claim 4, wherein the first plurality of feature vectors and the second plurality of feature vectors include a plurality of confidence scores, and further comprising: Performing, by the controller, a linear regression process based on the plurality of confidence scores to generate a normalized plurality of confidence scores for each of the first plurality of feature vectors and the second plurality of feature vectors, the normalized plurality of confidence scores being based on the confidence score of a predetermined candidate speech recognition result among the plurality of candidate speech recognition results.
6. The method according to claim 1, wherein the NLU information is based on slot-based trigger features or semantic features representing slot- and intent-sensitive utterance embeddings.
7. The method according to claim 6, wherein the first speech recognition engine is a domain-specific speech recognition engine, and the second speech recognition engine is a general speech recognition engine or a cloud-based speech recognition engine.
8. The method according to claim 7, wherein the first plurality of feature vectors and the second plurality of feature vectors include bidirectional long short-term memory (BLSTM) features.
9. A method for speech recognition in a system, executed by a controller, comprising: Parsing a plurality of candidate speech recognition results from a speech input; Extracting a first plurality of feature vectors from each of the plurality of candidate speech recognition results via a first speech recognition engine; Extracting a second plurality of feature vectors from each of the plurality of candidate speech recognition results via a second speech recognition engine different from the first speech recognition engine; Extracting NLU results from each of the plurality of candidate speech recognition results based on natural language understanding (NLU) information; Compressing the first plurality of feature vectors and the second plurality of feature vectors to the shared projection layer based on the NLU results and NLU-related features, wherein the NLU-related features include slot-based trigger features and semantic features representing utterance embeddings sensitive to slots / intents; Further compressing the shared projection layer to a second projection layer based on the NLU results and NLU-related features; Associating a ranking score with each of the plurality of candidate speech recognition results via a neural network ranker, the ranking score being based on the first plurality of feature vectors and the second plurality of feature vectors and the NLU results of each of the plurality of candidate speech recognition results; Selecting, from the plurality of candidate speech recognition results, the speech recognition result associated with the ranking score having the highest value; And Operating the system using the speech recognition result selected from the plurality of candidate speech recognition results corresponding to the highest ranking score as an input.
10. The method according to claim 9, wherein the neural network ranker is a deep feed-forward neural network ranker.
11. The method according to claim 9, wherein the compression is performed via a shared projection matrix.
12. The method according to claim 11, further comprising, in response to the first plurality of feature vectors and the second plurality of feature vectors being less than a threshold size, bypassing, by the controller, the shared projection layer and the second projection layer such that the second plurality of feature vectors are directly fed to the neural network ranker, wherein the threshold size of the feature vectors is less than 2 features for each hypothesis.
13. The method according to claim 12, wherein the first plurality of feature vectors and the second plurality of feature vectors include a plurality of confidence scores, and further comprising: The controller performs linear regression processing based on the plurality of confidence scores to generate a normalized plurality of confidence scores for each of the first plurality of feature vectors and the second plurality of feature vectors, the normalized plurality of confidence scores being based on the confidence score of a predetermined candidate speech recognition result among the plurality of candidate speech recognition results.
14. The method according to claim 9, wherein the NLU information is a slot-based trigger feature or a semantic feature representing a slot- and intent-sensitive utterance embedding.
15. The method according to claim 14, wherein the first speech recognition engine is a domain-specific speech recognition engine, and the second speech recognition engine is a general speech recognition engine or a cloud-based speech recognition engine.
16. The method according to claim 15, wherein the first plurality of feature vectors and the second plurality of feature vectors include bidirectional long short-term memory (BLSTM) features.
17. A speech recognition system, comprising: a microphone configured to receive speech input from one or more users; a processor in communication with the microphone, the processor being programmed to: parse a plurality of candidate speech recognition results from the speech input; receive, from a first speech recognition engine, a first plurality of feature vectors for each of the plurality of candidate speech recognition results, the first plurality of feature vectors including a first confidence score; receive, from a second speech recognition engine different from the first speech recognition engine, a second plurality of feature vectors for each of the plurality of candidate speech recognition results, the second plurality of feature vectors including a second confidence score lower than the first confidence score; extract NLU results from each of the plurality of candidate speech recognition results based on natural language understanding (NLU) information; associate a ranking score with each of the plurality of candidate speech recognition results via a neural network ranker, the ranking score being based on the first plurality of feature vectors, the second plurality of feature vectors, and the NLU results of each of the plurality of candidate speech recognition results, wherein the neural network ranker raises the second confidence score to be greater than the first confidence score based on NLU-related features, wherein the NLU-related features include a slot-based trigger feature and a semantic feature representing a slot / intent-sensitive utterance embedding; and select a speech recognition result associated with the ranking score having the highest value from the plurality of candidate speech recognition results.
18. The speech recognition system according to claim 17, wherein the processor is further programmed to operate the system using the speech recognition result selected from the plurality of candidate speech recognition results corresponding to the highest ranking score as an input.
19. The speech recognition system according to claim 17, wherein the processor is further programmed to train a neural network associated with the speech recognition system using at least the NLU results.
20. The speech recognition system according to claim 17, wherein the processor is further programmed to compress the first plurality of feature vectors and the second plurality of feature vectors to the shared projection layer via the shared projection layer based on the NLU result and the NLU-related features, and further compress the shared projection layer to a second projection layer based on the NLU result and the NLU-related features.
Citation Information
Patent Citations
System and method for ranking of hybrid speech recognition results with neural networks
US10170110B2
Evidence-Based Natural Language Input Recognition
US20160259779A1
Electronic apparatus for processing user utterance and server
US20180374482A1