Techniques for improved keyword detection
By combining an improved statistical language model with a keyword language model, the accuracy problem of keyword detection in the existing technology is solved, more efficient keyword recognition and parsing is achieved, and the user intention understanding ability of computing devices is improved.
Patent Information
- Application Number
- CN202111573180.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2016-09-23
- Filing Date
- 2017-08-17
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2037-08-17
AI Technical Summary
Existing keyword detection technologies have difficulty in effectively distinguishing and prioritizing keywords, resulting in lost or unclear keyword information in the output transcript.
The modified statistical language model and the linear interpolation of the keyword language model are adopted to train the speech recognition algorithm through the hidden Markov model to prioritize the matching of keywords. The speech parser and auxiliary agent are combined to update the confidence state to improve the accuracy of keyword detection.
Improves the accuracy and efficiency of keyword detection, ensuring more accurate recognition and parsing of keywords in output transcripts, and enabling computing devices to better understand user intent.
Smart Images

Figure CN114141241B_ABST
Abstract
Description
[0001] This application is a divisional application of PCT international application number PCT / US2017 / 047389, international application date August 17, 2017, application number 201780051396.4 entering the Chinese national phase, and entitled "Technology for Improved Keyword Detection".
[0002] Cross-reference to related U.S. patent applications
[0003] This application claims priority to U.S. utility patent application Ser. No. 15 / 274,498, filed on Sep. 23, 2016, entitled “TECHNOLOGIES FOR IMPROVED KEYWORD SPOTTING.” Background Art
[0004] Automatic speech recognition for computing devices has a wide range of applications, including providing spoken commands to a computing device or dictating documents (such as entries in a medical record). In some cases, keyword detection may be required, for example, if a piece of speech data is being searched for the presence of a specific word or group of words.
[0005] Keyword detection is typically accomplished by executing a speech recognition algorithm that is customized to match only keywords and ignore or reject words outside the keyword list. The output of the keyword detector may be only the matched keywords, with no output provided for speech data that does not match the keywords. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The concepts described herein are illustrated in the accompanying drawings by way of example and not limitation. For simplicity and clarity of illustration, the elements shown in the drawings are not necessarily drawn to scale. Where considered appropriate, reference numerals have been repeated among the drawings to indicate corresponding or similar elements.
[0007] Figure 1 is a simplified block diagram of at least one embodiment of a computing device for keyword detection;
[0008] Figure 2 Can be Figure 1 A block diagram of at least one embodiment of an environment established by a computing device;
[0009] Figure 3 is a simplified flow chart of at least one embodiment of a method for training an automatic speech recognition algorithm for keyword detection, which algorithm may be Figure 1 a computing device for execution; and
[0010] Figure 4 It is a Figure 1A simplified flowchart of at least one embodiment of a method for automatic speech recognition with keyword localization performed by a computing device of the present invention. DETAILED DESCRIPTION
[0011] While the concepts of the present invention are susceptible to various modifications and alternative forms, specific embodiments of the invention are shown by way of example in the drawings and are herein described in detail. However, it should be understood that there is no intention to limit the concepts of the present disclosure to the specific forms disclosed, but rather, the present invention is intended to cover all modifications, equivalents, and alternatives consistent with the present disclosure and the appended claims.
[0012] References in the specification to "one embodiment," "an embodiment," "an illustrative embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but each embodiment may or may not include the particular feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an embodiment, it should be understood that, whether or not explicitly described, it is within the knowledge of one skilled in the art to implement such feature, structure, or characteristic in conjunction with other embodiments. In addition, it should be understood that items included in a list of the form "at least one of A, B, and C" may refer to (A); (B); (C); (A and B); (A and C); (B and C); or (A, B, and C). Similarly, items listed in the form "at least one of A, B, or C" may refer to (A); (B); (C); (A and B); (B and C); (A or C), or (A, B, and C).
[0013] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried by or stored on one or more transient or non-transient machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. A machine-readable storage medium may be embodied as any storage device, mechanism, or other physical structure for storing or transmitting information in a machine-readable form (e.g., volatile or non-volatile memory, a media disk, or other media device).
[0014] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. On the contrary, in some embodiments, such features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of a structural or method feature in a particular drawing does not imply that such feature is required in all embodiments, and in some embodiments, such feature may not be included or may be combined with other features.
[0015] See now Figure 1 An illustrative computing device 100 includes a microphone 108 for capturing voice data from a user of the computing device 100. The computing device 100 executes an automatic speech recognition algorithm on the voice data using a statistical language model that has been modified to prioritize matching words in a list of keywords. The output transcript of the automatic speech recognition algorithm includes all transcribed words in the voice data, including words that match the keywords and words that do not match the keywords. The computing device 100 then takes action based on the output transcript, for example, by executing a command given by the user, parsing the output transcript based on keywords, or determining the user's intent based on the matched keywords.
[0016] The illustrative computing device 100 may be embodied as any type of computing device capable of performing the functions described herein. For example, the computing device 100 may be embodied as or otherwise included in, but not limited to, a smartphone, a cellular phone, a wearable computer, an embedded computing system, a system on a chip (SoC), a tablet computer, a notebook computer, a laptop computer, a server computer, a desktop computer, a handheld device, a messaging device, a camera device, a multi-processor system, a processor-based system, a consumer electronic device, and / or any other computing device.
[0017] The illustrative computing device 100 includes a processor 102, memory 104, an input / output (I / O) subsystem 106, data storage 108, and a microphone 110. In some embodiments, one or more of the illustrative components of the computing device 100 may be incorporated into, or otherwise form part of, another component. For example, in some embodiments, the memory 104 or portions thereof may be incorporated into the processor 102.
[0018] The processor 102 may be embodied as any type of processor capable of performing the functions described herein. For example, the processor 102 may be embodied as (multiple) single-core or multi-core processors, single-socket or multi-socket processors, digital signal processors, graphics processors, microcontrollers, or other processors or processing / control circuits. Similarly, the memory 104 may be embodied as any type of volatile or non-volatile memory or data storage device capable of performing the functions described herein. In operation, the memory 104 may store various data and software used during operation of the computing device 100, such as an operating system, applications, programs, libraries, and drivers. The memory 104 is communicatively coupled to the processor 102 via an I / O subsystem 106, which may be embodied as circuit systems and / or components for facilitating input / output operations with the processor 102, memory 104, and other components of the computing device 100. For example, the I / O subsystem 106 may be embodied as or otherwise include a memory controller hub, an input / output control hub, firmware devices, communication links (i.e., point-to-point links, bus links, wires, cables, optical guides, printed circuit board traces, etc.) and / or other components and subsystems for facilitating input / output operations. In some embodiments, the I / O subsystem 106 may form part of a system on a chip (SoC) and may be combined on a single integrated circuit chip with the processor 102, memory 104, and other components of the computing device 100.
[0019] Data storage 108 may be embodied as any type of one or more devices configured for short-term or long-term storage of data. For example, data storage 108 may include any one or more memory devices and circuits, memory cards, hard drives, solid-state drives, or other data storage devices.
[0020] The microphone 110 may be embodied as any type of device capable of converting sound into an electrical signal. The microphone 110 may be based on any type of suitable sound capture technology, such as electromagnetic induction, capacitance change, and / or piezoelectric.
[0021] Of course, in some embodiments, computing device 100 may include additional components typically found in computing devices 100, such as a display 112 and / or one or more peripherals 114. Peripherals 114 may include a keyboard, a mouse, communication circuitry, and the like.
[0022] Display 112 may be embodied as any type of display that can display information to a user of computing device 100, such as a liquid crystal display (LCD), a light emitting diode (LED) display, a cathode ray tube (CRT) display, a plasma display, an image projector (e.g., 2D or 3D), a laser projector, a touch screen display, a heads-up display, and / or other display technologies.
[0023] See now Figure 2 In use, the computing device 100 can establish an environment 200. The illustrative environment 200 includes an automatic speech recognition algorithm trainer 202, a voice data capturer 204, an automatic speech recognizer 206, a speech parser 208, and an auxiliary agent 210. The various components of the environment 200 can be embodied as hardware, firmware, software, or a combination thereof. Thus, in some embodiments, one or more of the components of the environment 200 can be embodied as a circuit system or collection of electrical devices (e.g., the automatic speech recognition algorithm trainer circuit 202, the voice data capturer circuit 204, the automatic speech recognizer circuit 206, etc.).
[0024] It should be understood that in such embodiments, the automatic speech recognition algorithm trainer circuit 202, the speech data capture circuit 204, the automatic speech recognizer circuit 206, etc. may form part of one or more of the processor 102, the I / O subsystem 106, the microphone 110, and / or other components of the computing device 100. Additionally, in some embodiments, one or more of the illustrative components may form part of another component and / or one or more of the illustrative components may be independent of one another. Furthermore, in some embodiments, one or more of the components of the environment 200 may be embodied as virtualized hardware components or simulated architectures that may be established and maintained by the processor 102 or other components of the computing device 100.
[0025] The automatic speech recognition algorithm trainer 202 is configured to train an automatic speech recognition algorithm. In an illustrative embodiment, the automatic speech recognition algorithm trainer 202 obtains labeled training data (i.e., training speech data with corresponding transcripts), which is used to train a hidden Markov model and generate an acoustic model. In some embodiments, the training data can be data from a specific field (such as in the medical or legal fields), and some or all keywords can correspond to terms from that field. The illustrative automatic speech recognition algorithm uses an acoustic model to match speech data to phonemes, and also uses a statistical language model to match speech data and corresponding phonemes to words based on the relative likelihood of usage of sequences of different words, such as n-grams of different lengths (e.g., unigrams, bigrams, bigrams). The illustrative statistical language model is a large vocabulary language model and can include more than, less than, or any of 1,000, 2,000, 5,000, 10,000, 20,000, 50,000, 100,000, 200,000, 500,000, and 1,000,000 words. In an illustrative embodiment, the automatic speech recognition algorithm trainer 202 includes a statistical language keyword enhancer 214 that is configured to enhance the statistical language model with a keyword language model that uses a second hidden Markov model to match words in a keyword list. The statistical language keyword enhancer 214 can enhance the statistical language model by performing linear interpolation between the statistical language model and the keyword language model. In an illustrative embodiment, the automatic speech recognition algorithm trainer 202 modifies the large vocabulary language model to prioritize matching the keyword over the similar words when the speech data reasonably matches one of the keywords or one of the similar words in the statistical language model. To this end, the automatic speech recognition algorithm trainer 202 weights the keywords, giving them a higher weight than the corresponding words in the large vocabulary language model. Keywords can include key phrases, which are more than one word, and the automatic speech recognition algorithm can treat key phrases as single words (even if they are more than one word). The number of keywords can be more, less, or any of 1, 2, 5, 10, 20, 50, 100, 200, 500, 1000, 2000, and 5000 words or phrases.
[0026] In other embodiments, a different speech recognition algorithm may be used instead of or in addition to a hidden Markov model with a corresponding different speech recognition training process. For example, the speech recognition algorithm may be based on a neural network, including a deep neural network and / or a recurrent neural network. It should be understood that in some embodiments, the computing device 100 may receive some or all parameters of an automatic speech recognition algorithm that has been trained by a different computing device and need not perform some or all of the training itself.
[0027] The voice data capturer 204 is configured to capture voice data using the microphone 110. The voice data capturer 204 may capture voice data continuously, constantly, periodically, or upon a user's command (such as a user pressing a button to start voice recognition).
[0028] Automatic speech recognizer 206 is configured to execute an automatic speech recognition algorithm trained on speech data by automatic speech recognition algorithm trainer 202. Automatic speech recognizer 206 generates an output transcript for use by, for example, an application of computing device 100, the output transcript including both words present in the keyword list and words not present in the keyword list. In an illustrative embodiment, the output transcript generated by automatic speech recognizer 206 includes both matched individual keywords and separate complete transcripts. In some embodiments, the output transcript may include only the transcribed text without any specific indication of which words are keywords.
[0029] The speech parser 208 is configured to parse the output transcript to determine semantic meaning based on a particular application. In some embodiments, the speech parser 208 may use matched keywords to determine the context of a portion of the output transcript. For example, in one embodiment, a user may dictate an entry into a medical record and may say, "Prescribe 10 ml of dapoxetine to John Smith for an allergic reaction on October 20, 2015." Prescription, Insurance ID: 7503986, Claim Number: 450934." The matching keywords may be "prescription," "insurance ID," and "claim ID." The speech parser 208 may use the matching keywords to determine the semantic context of each portion of the output transcript and determine the parameters of the medical entry, such as prescription ( , 10ml), insurance ID (7503986), claim ID (450934), etc.
[0030] The assist agent 210 is configured to conduct a conversation with a user of the computing device 100 to assist with certain tasks. The assist agent 210 includes a confidence state manager 216 that stores information about the current state of the conversation between the user and the computing device 100, such as the user's current discussion topic or current intention. The confidence state manager 216 includes a keyword analyzer 218. When the output transcript matches a keyword, the keyword analyzer 218 can update the current confidence state in response to matching the keyword, and can do so without waiting for the next transcribed word. In an illustrative embodiment, the keyword analyzer 218 can consult previous transcriptions and update and correct any ambiguities (such as by consulting a word lattice of an automatic speech recognition algorithm and searching for a potentially more suitable match based on the presence of the keyword).
[0031] Now see Figure 3 In use, the computing device 100 may perform a method 300 for training an automatic speech recognition algorithm. The method 300 begins in block 302, where the computing device trains the automatic speech recognition algorithm. In an illustrative embodiment, the computing device 100 trains an acoustic model using a hidden Markov model in block 304 and trains a statistical language model in block 306. The training is performed based on labeled training data. In some embodiments, in block 308, the computing device 100 may train the automatic speech recognition algorithm using domain-specific training data (such as data from the medical or legal fields). In addition to or instead of training an algorithm based on a hidden Markov model, in some embodiments, the computing device 100 may train an automatic speech recognition algorithm based on a neural network in block 310.
[0032] In block 312, computing device 100 augments the language model with the keywords. In the illustrative embodiment, computing device 100 does so by interpolating between the statistical language model and the keyword language model in block 314.
[0033] Now see Figure 4 In use, computing device 100 may execute method 400 for executing an automatic speech recognition algorithm. The method begins in block 402, where computing device 100 determines whether to recognize speech. Computing device 100 may perform speech recognition continuously, constantly, periodically, and / or when directed by a user of computing device 100. In some embodiments, computing device 100 may continuously monitor for the presence of speech using a speech detection algorithm and, when the speech detection algorithm detects speech, perform speech recognition. If the computing device decides not to perform speech recognition, method 400 loops back to block 402. If computing device 100 decides to perform speech recognition, method 400 proceeds to block 404.
[0034] In block 404, computing device 100 captures voice data from microphone 110. It should be understood that in some embodiments, the voice data may alternatively be captured by a different computing device and sent to computing device 100 via some communication means (such as the Internet).
[0035] In block 406 , computing device 100 performs automatic speech recognition on the captured speech data. Computing device 100 recognizes phonemes of the speech data based on the acoustic model in block 408 and recognizes words and keywords based on the statistical language model in block 410 .
[0036] In block 412, computing device 100 generates an output transcript. In an illustrative embodiment, the output transcript includes both the matched individual keywords and a separate complete transcript. In some embodiments, the output transcript may include only the transcribed text without any specific indication of which words are keywords. The output transcript may then be further processed or used by computing device 100 (such as by an application provided to computing device 100).
[0037] In block 414, computing device 1000 parses the output transcript. In block 416, computing device 100 identifies a context for a portion of the output transcript based on the identified keywords.
[0038] In block 418, in some embodiments, computing device 100 may update the confidence state of the assist agent in response to matching the keyword. In an illustrative embodiment, computing device 100 may consult a transcript of a previous conversation with the user and update and correct any ambiguities (such as by consulting a word grid of an automatic speech recognition algorithm and searching for a potentially more suitable match based on the presence of the keyword). In some embodiments, computing device 100 may update the current confidence state without waiting for the next transcribed word, even though typical behavior of computing device 100 is to wait for the next complete sentence, the next silence, and / or the like before taking any action on the output transcription.
[0039] Example
[0040] The following provides illustrative examples of the devices, systems, and methods disclosed herein. Embodiments of the devices, systems, and methods may include any one or more of the examples described below, and any combination thereof.
[0041] Example 1 includes a computing device for automatic speech recognition, the computing device comprising: an automatic speech recognition algorithm trainer, the automatic speech recognition algorithm trainer for obtaining a statistical language model for use with the automatic speech recognition algorithm, wherein the statistical language model includes a large vocabulary language model that has been modified to preferentially match words present in a plurality of keywords; a voice data capturer, the voice data capturer for receiving voice data from a user of the computing device; and an automatic speech recognizer, the automatic speech recognizer for executing the automatic speech recognition algorithm on the voice data to produce an output transcript, wherein the output transcript includes one or more keywords from the plurality of keywords and one or more words that are not in the plurality of keywords.
[0042] Example 2 includes the subject matter of Example 1, and wherein the large vocabulary language model that has been modified to preferentially match words present in multiple keywords includes a first hidden Markov model and a second hidden Markov model, wherein the first hidden Markov model is used to match words present in the large vocabulary and the second hidden Markov model is used to match words present in multiple keywords.
[0043] Example 3 includes the subject matter of any of Examples 1 and 2, and wherein weights of the plurality of keywords are higher than corresponding weights of the remainder of the statistical language model such that the statistical language model preferentially matches the plurality of keywords.
[0044] Example 4 includes the subject matter of any of Examples 1-3, and wherein the statistical language model is formed by linear interpolation of the large vocabulary language model and the keyword language model.
[0045] Example 5 includes the subject matter of any of Examples 1-4, and wherein the plurality of keywords includes fewer than fifty words and the large vocabulary includes more than one thousand words.
[0046] Example 6 includes the subject matter of any of Examples 1-5, and wherein receiving the voice data comprises capturing the voice data with a microphone of the computing device.
[0047] Example 7 includes the subject matter of any of Examples 1-6, and further includes a speech parser to identify a context of a portion of the output transcript based on the one or more keywords; and parse the output transcript according to the context of the portion of the output transcript.
[0048] Example 8 includes the subject matter of any of Examples 1-7, and wherein obtaining a statistical language model for use in an automatic speech recognition algorithm includes: training the statistical language model for a large vocabulary and enhancing the statistical language model with a keyword language model so that the statistical language model preferentially matches multiple keywords.
[0049] Example 9 includes the subject matter of any of Examples 1-8, and wherein the statistical language model has been trained using domain-specific training data.
[0050] Example 10 includes the subject matter of any of Examples 1-9, and further comprising: an auxiliary agent for updating a confidence state of the auxiliary agent in response to a match of one or more keywords.
[0051] Example 11 includes the subject matter of any of Examples 1-10, and wherein updating the confidence state in response to matching one or more keywords comprises updating the interaction context without waiting for a next recognized word of the speech data.
[0052] Example 12 includes the subject matter of any of Examples 1-11, and wherein updating the confidence state in response to matching one or more keywords comprises: searching a word lattice of an automatic speech recognition algorithm and finding a better match of the word lattice to the speech data based on the one or more keywords.
[0053] Example 13 includes the subject matter of any of Examples 1-12, and wherein at least one keyword of the plurality of keywords is a key phrase including two or more words.
[0054] Example 14 includes a method for automatic speech recognition by a computing device, the method comprising: obtaining, by the computing device, a statistical language model for use in an automatic speech recognition algorithm, wherein the statistical language model comprises a large vocabulary language model that has been modified to preferentially match words present in a plurality of keywords; receiving, by the computing device, speech of a user of the computing device; and executing, by the computing device, the automatic speech recognition algorithm on the speech data to produce an output transcript, wherein the output transcript comprises one or more keywords from the plurality of keywords and one or more words that are not in the plurality of keywords.
[0055] Example 15 includes the subject matter of Example 14, and wherein the large vocabulary language model that has been modified to preferentially match words present in multiple keywords includes a first hidden Markov model and a second hidden Markov model, wherein the first hidden Markov model is used to match words present in the large vocabulary and the second hidden Markov model is used to match words present in multiple keywords.
[0056] Example 16 includes the subject matter of any of Examples 14 and 15, and wherein weights of the plurality of keywords are higher than corresponding weights of the remainder of the statistical language model such that the statistical language model preferentially matches the plurality of keywords.
[0057] Example 17 includes the subject matter of any of Examples 14-16, and wherein the statistical language model is formed by linear interpolation of the large vocabulary language model and the keyword language model.
[0058] Example 18 includes the subject matter of any of Examples 14-17, and wherein the plurality of keywords includes fewer than fifty words and the large vocabulary includes more than one thousand words.
[0059] Example 19 includes the subject matter of any of Examples 14-18, and wherein receiving the voice data comprises capturing the voice data using a microphone of the computing device.
[0060] Example 20 includes the subject matter of any of Examples 14-19, and further comprising: identifying, by the computing device and based on the one or more keywords, a context for a portion of the output transcript; and parsing, by the computing device, the output transcript based on the context for the portion of the output transcript.
[0061] Example 21 includes the subject matter of any of Examples 14-20, and wherein obtaining a statistical language model for use in an automatic speech recognition algorithm comprises: training the statistical language model for a large vocabulary and enhancing the statistical language model with a keyword language model so that the statistical language model preferentially matches multiple keywords.
[0062] Example 22 includes the subject matter of any of Examples 14-21, and wherein the statistical language model has been trained using domain-specific training data.
[0063] Example 23 includes the subject matter of any of Examples 14-22, and further comprising updating, by the auxiliary agent of the computing device, a confidence state of the auxiliary agent in response to matching the one or more keywords.
[0064] Example 24 includes the subject matter of any of Examples 14-23, and wherein updating, by the auxiliary agent, the confidence state in response to matching the one or more keywords comprises updating, by the auxiliary agent, the interaction context without waiting for a next recognized word of the speech data.
[0065] Example 25 includes the subject matter of any of Examples 14-24, and wherein updating the confidence state by the auxiliary agent in response to matching one or more keywords comprises: searching a word lattice of an automatic speech recognition algorithm and finding a better match of the word lattice to the speech data based on the one or more keywords.
[0066] Example 26 includes the subject matter of any of Examples 14-25, and wherein at least one keyword of the plurality of keywords is a key phrase including two or more words.
[0067] Example 27 includes one or more computer-readable media including a plurality of instructions stored thereon that, when executed, cause a computing device to perform the method of any one of claims 14-26.
[0068] Example 28 includes a computing device for low-power capture of sensor values with high-precision timestamps, the computing device comprising: means for obtaining a statistical language model for use with an automatic speech recognition algorithm, wherein the statistical language model comprises a large vocabulary language model that has been modified to prioritize matching words that are present in a plurality of keywords; means for receiving speech data from a user of the computing device; and means for executing the automatic speech recognition algorithm on the speech data to produce an output transcript, wherein the output transcript comprises one or more keywords from the plurality of keywords and one or more words that are not in the plurality of keywords.
[0069] Example 29 includes the subject matter of Example 28, and wherein the large vocabulary language model that has been modified to preferentially match words present in multiple keywords includes a first hidden Markov model and a second hidden Markov model, wherein the first hidden Markov model is used to match words present in the large vocabulary and the second hidden Markov model is used to match words present in multiple keywords.
[0070] Example 30 includes the subject matter of any of Examples 28 and 29, and wherein weights of the plurality of keywords are higher than corresponding weights of the remainder of the statistical language model such that the statistical language model preferentially matches the plurality of keywords.
[0071] Example 31 includes the subject matter of any of Examples 28-30, and wherein the statistical language model is formed by linear interpolation of the large vocabulary language model and the keyword language model.
[0072] Example 32 includes the subject matter of any of Examples 28-31, and wherein the plurality of keywords includes fewer than fifty words and the large vocabulary includes more than one thousand words.
[0073] Example 33 includes the subject matter of any of Examples 28-32, and wherein the means for receiving the voice data comprises means for capturing the voice data with a microphone of the computing device.
[0074] Example 34 includes the subject matter of any of Examples 28-33, and further includes means for identifying a context of a portion of the output transcript based on the one or more keywords; and means for parsing the output transcript based on the context of the portion of the output transcript.
[0075] Example 35 includes the subject matter of any of Examples 28-34, and wherein the apparatus for obtaining a statistical language model for use in an automatic speech recognition algorithm comprises: an apparatus for training the statistical language model for a large vocabulary, and an apparatus for enhancing the statistical language model with a keyword language model so that the statistical language model preferentially matches multiple keywords.
[0076] Example 36 includes the subject matter of any of Examples 28-35, and wherein the statistical language model has been trained using domain-specific training data.
[0077] Example 37 includes the subject matter of any of Examples 28-36, and further comprising means for updating, by the auxiliary agent of the computing device, a confidence status of the auxiliary agent in response to matching one or more keywords.
[0078] Example 38 includes the subject matter of any of Examples 28-37, and wherein the means for updating, by the auxiliary agent, the confidence state in response to matching one or more keywords comprises means for updating, by the auxiliary agent, the interaction context without waiting for the next recognized word of the speech data.
[0079] Example 39 includes the subject matter of any of Examples 28-38, and wherein the means for updating the confidence state by the auxiliary agent in response to matching one or more keywords comprises: means for searching a word grid of an automatic speech recognition algorithm, and means for finding a better match of the word grid to the speech data based on the one or more keywords.
[0080] Example 40 includes the subject matter of any of Examples 28-39, and wherein at least one keyword of the plurality of keywords is a key phrase including two or more words.
Claims
1. A computing device comprising: microphone; processor; as well as a memory coupled to the processor and having instructions stored thereon that, when executed by the processor, cause the computing device to: Obtaining an enhanced language model obtained by enhancing a statistical language model with a second language model for an automatic speech recognition algorithm, wherein the statistical language model is based on the likelihood of occurrence of a word sequence, and the second language model is based on a plurality of keywords, and wherein the statistical language model is enhanced by matching words present in the plurality of keywords of the second language model; receiving voice data from the microphone; executing the automatic speech recognition algorithm on the speech data based on the enhanced language model to generate an output transcript, wherein the output transcript includes one or more keywords from the plurality of keywords and one or more words that are not from the plurality of keywords; as well as At least a portion of the automatic speech recognition algorithm is received, the automatic speech recognition algorithm not trained by the computing device.
2. The computing device of claim 1, wherein: The computing device comprises a smartphone.
3. The computing device of claim 1, wherein: At least one of the keywords includes a key phrase comprising more than one word.
4. The computing device of claim 1, wherein: The automatic speech recognition algorithm includes an algorithm based on a neural network.
5. The computing device of claim 1 , further comprising one or more of the following: Display; and Communication circuit.
6. The computing device of claim 1, wherein: The microphone is configured to capture the voice data in response to a user command.
7. The computing device of claim 1, wherein: Each of the plurality of keywords is associated with a weight in the automatic speech recognition algorithm.
8. The computing device of claim 1, wherein: The instructions, when executed by the processor, further cause the computing device to: storing information related to a confidence state of a conversation between the computing device and the user; and In response to a match of at least one keyword among the plurality of keywords, the confidence status is updated.
9. A method for speech recognition, comprising: Obtaining, by a processor of a computing device, an enhanced language model obtained by enhancing a statistical language model with a second language model for an automatic speech recognition algorithm, wherein the statistical language model is based on the likelihood of occurrence of a word sequence, and the second language model is based on a plurality of keywords, and wherein the statistical language model is enhanced by matching words present in the plurality of keywords of the second language model; receiving, by the processor, voice data from a microphone of the computing device; executing, by the processor, the automatic speech recognition algorithm on the speech data based on the enhanced language model to generate an output transcript, wherein the output transcript includes one or more keywords from the plurality of keywords and one or more words that are not from the plurality of keywords; as well as At least a portion of the automatic speech recognition algorithm is received by the processor, the automatic speech recognition algorithm not trained by the computing device.
10. The method according to claim 9, wherein The computing device comprises a smartphone.
11. The method according to claim 9, wherein At least one of the keywords includes a key phrase comprising more than one word.
12. The method according to claim 9, wherein The automatic speech recognition algorithm includes an algorithm based on a neural network.
13. The method according to claim 9, wherein Each of the plurality of keywords is associated with a weight in the automatic speech recognition algorithm.
14. A device for speech recognition in a computing device, comprising: means for obtaining an enhanced language model obtained by enhancing a statistical language model with a second language model for an automatic speech recognition algorithm, wherein the statistical language model is based on the likelihood of occurrence of a word sequence, and the second language model is based on a plurality of keywords, and wherein the statistical language model is enhanced by matching words present in the plurality of keywords of the second language model; means for receiving voice data from a microphone of said computing device; means for executing the automatic speech recognition algorithm on the speech data based on the enhanced language model to generate an output transcript, wherein the output transcript includes one or more keywords from the plurality of keywords and one or more words that are not from the plurality of keywords; as well as Means for receiving at least a portion of the automatic speech recognition algorithm, the automatic speech recognition algorithm not trained by the computing device.
15. The device according to claim 14, characterized in that The computing device comprises a smartphone.
16. The device according to claim 14, wherein At least one of the keywords includes a key phrase comprising more than one word.
17. The device according to claim 14, wherein The automatic speech recognition algorithm includes an algorithm based on a neural network.
Citation Information
Patent Citations
Voice map searching method and system
CN104008132A
Technologies for improved keyword spotting
CN109643542A