Voice input processing

The voice input processing system improves speech recognition by contextually selecting grammars and transcriptions, addressing ambiguity in user intent interpretation and reducing resource usage.

JP7713570B2Active Publication Date: 2025-07-25GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024147124
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-12-03
Filing Date
2024-08-29
Publication Date
2025-07-25
Estimated Expiration
2039-11-27

AI Technical Summary

Technical Problem

Existing speech recognition systems struggle to accurately interpret user intentions from voice input due to ambiguity in transcriptions and lack of context-based grammar selection, leading to incorrect actions and increased resource consumption.

Method used

A voice input processing system that identifies multiple grammars and transcriptions, adjusts grammar confidence scores based on context, and selects the most likely combination to match user intent, reducing latency and resource usage.

Benefits of technology

Enhances speech recognition accuracy by selecting appropriate grammars based on context, minimizing incorrect actions and reducing computational and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713570000001
    Figure 0007713570000001
  • Figure 0007713570000002
    Figure 0007713570000002
  • Figure 0007713570000003
    Figure 0007713570000003
Patent Text Reader

Abstract

To execute voice input processing.SOLUTION: The processing includes determining a context indicating that a user is listening to the music through a computing device, receiving audio data of an utterance, generating one or more candidate transcriptions for the utterance, and parsing a candidate transcription from the one or more candidate transcriptions having the highest transcription confidence score using a grammar based on the context indicating that the user is listening to the music through the computing device to identify an action to be performed by the computing device.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification generally relates to speech recognition.

Background Art

[0002] There is an increasing demand to enable interaction with a computer using voice input. This requires the development of input processing, particularly methods of programming a computer to process and analyze natural language data. Such processing may involve speech recognition, which is a field of computational linguistics that enables a computer to recognize spoken language and convert it into text.

Summary of the Invention

Problems to be Solved by the Invention

[0003] This specification generally relates to speech recognition.

Means for Solving the Problems

[0004] To enable a user to provide input to a computing device via voice, a voice input processing system can identify a plurality of grammars to apply to a plurality of candidate transcripts generated by an automatic speech recognition unit using context. Each grammar may indicate different intentions of the speaker or actions that the system performs on the same candidate transcript. The system can select a grammar and a candidate transcript based on the grammar for parsing the candidate transcript, the likelihood that the grammar matches the user's intention, and the likelihood that the candidate transcript matches the user's utterance. Next, the system can execute an action corresponding to the grammar using the details included in the selected candidate transcript.

[0005] More specifically, the speech processing system receives an utterance from a user and generates a word lattice. The word lattice is a data structure that reflects likely words in the utterance and a confidence score for each word. The system identifies, from the word lattice, a plurality of candidate transcriptions and a transcription confidence score for each candidate transcription. The system identifies the current context based on user characteristics, the system's location, the system's characteristics, the application running on the system (e.g., the currently active application or the application running in the foreground), or other similar context data. Based on the context, the system generates a grammar confidence score for each of a plurality of grammars that parse each candidate transcription. If multiple grammars may be applied to the same candidate transcription, the system may adjust some of the plurality of grammar confidence scores. The system selects a grammar and a candidate transcription based on a combination of the adjusted plurality of grammar confidence scores and the plurality of transcription confidence scores.

[0006] According to an innovative aspect of the subject matter described in this application, a method for processing voice input includes receiving, by a computing device, audio data of an utterance; generating, by the computing device, using an acoustic model and a language model, a word lattice that includes a plurality of candidate transcriptions of the utterance and a plurality of transcription confidence scores, each of the plurality of transcription confidence scores reflecting the likelihood that a respective candidate transcription matches the utterance; determining, by the computing device, the context of the computing device; identifying, by the computing device, based on the context of the computing device, a plurality of grammars corresponding to the plurality of candidate transcriptions; determining, by the computing device, for each of the plurality of candidate transcriptions, based on the current context, a plurality of grammar confidence scores that reflect the likelihood that each grammar matches the respective candidate transcription; selecting, by the computing device, a candidate transcription from among the plurality of candidate transcriptions based on the plurality of transcription confidence scores and the plurality of grammar confidence scores; and providing, by the computing device for output, the selected candidate transcription as a transcription of the utterance.

[0007] These and other implementations can optionally include one or more of the following features. The action includes determining that two or more of a plurality of grammars correspond to one of a plurality of candidate transcriptions, and adjusting a plurality of grammar confidence scores for two or more grammars based on determining that two or more of the plurality of grammars correspond to one of the plurality of candidate transcriptions. The computing device selects a candidate transcription from among the plurality of candidate transcriptions based on the transcription confidence score and the adjusted grammar confidence score. The action of adjusting a plurality of grammar confidence scores for two or more grammars includes increasing each of the plurality of grammar confidence scores by a factor for each of the two or more grammars. The action includes determining, for each of the plurality of candidate transcriptions, the product of the respective transcription confidence score and the respective grammar confidence score. The computing device selects a candidate transcription from among the plurality of candidate transcriptions based on the product of the transcription confidence score and the respective grammar confidence score. The action of determining the context of the computing device by the computing device is based on the location of the computing device, the application running in the foreground of the computing device, and the time. The language model is configured to specify probabilities for sequences of words included in a word lattice. The acoustic model is configured to identify phonemes that match a portion of the audio data. The action includes the computing device performing an action based on the selected candidate transcription and the grammar that matches the selected candidate transcription.

[0008] Other embodiments of this aspect include corresponding systems, apparatuses, and computer programs recorded on computer storage devices, each configured to perform the operations of the method.

[0009] Certain embodiments of the subject matter described herein can be implemented to realize one or more of the following advantages. A speech recognition system can use both received speech input and determined context to select a grammar used to further process the received speech input to cause a computing device to perform an action. Thus, the speech recognition system can reduce the human-machine interface latency by applying a limited number of grammars to multiple candidate transcriptions. The speech recognition system can use a vocabulary that includes all or substantially all words of a language so that the speech recognition system can output a transcription when the system receives an unexpected input.

[0010] Details of one or more embodiments of the subject matter described herein are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Modes for Carrying Out the Invention

[0012] Like reference numerals and designations in the various drawings indicate like elements. Figure 1 shows an exemplary system 100 for selecting a grammar to apply to an utterance based on context. As will be described briefly below and in more detail, user 102 makes an utterance 104. Computing device 106 detects the utterance 104 and performs an action 108 in response to the utterance. Computing device 106 can use the context 110 of computing device 106 to determine the likely intent of user 102. Based on the likely intent of user 102, computing device 106 selects an appropriate action.

[0013] More specifically, user 102 makes an utterance 104 near computing device 106. Computing device 106 can be any type of computing device configured to detect audio. For example, computing device 106 can be a mobile phone, laptop computer, wearable device, smart appliance, desktop computer, television, or any other type of computing device capable of receiving audio data. Computing device 106 processes the audio data of utterance 104 and generates a word lattice 112. The word lattice 112 represents different combinations of multiple words, each of which can correspond to an utterance and a confidence score.

[0014] As shown in FIG. 1, user 102 does not speak clearly and makes utterance 104 that sounds like "flights out". Computing device 106 detects the audio of the user's voice and uses automatic speech recognition to determine what user 102 said. Computing device 106 applies an acoustic model to the audio data. The acoustic model can be configured to determine various phonemes that are likely to correspond to the audio data. For example, the acoustic model may determine that the audio data contains the phoneme of the "f" phoneme in addition to other phonemes. Computing device 106 can apply a language model to the phonemes. The language model may be configured to identify likely words and sequences of words that match the phonemes. In some implementations, computing device 106 can send the audio data to another computing device, such as a server. The server may perform speech recognition on the audio data.

[0015] The language model generates a word lattice 112 that includes different combinations of multiple words that match utterance 104. The word lattice includes two words that could be the first word of utterance 104. Sometimes the word 114 "lights" is the first word, and sometimes the word 116 "flights" is the first word. The language model can calculate a confidence score that reflects the likelihood that the words in the word lattice are the words spoken by the user. For example, the language model can determine that the confidence score 118 for the word 114 "lights" is 0.65 and the confidence score 120 for the word 116 "flights" is 0.35. The word lattice 112 includes words that could be the second word. In this example, the language model identified only one possible word 122 "out" as the second word.

[0016] In some implementations, computing device 106 can use a sequence-to-sequence neural network, or other types of neural networks, instead of an acoustic model and a language model. The neural network can have one or more hidden layers and can be trained using machine learning and training data that includes audio data of the utterance and transcription of the sample corresponding to each sample of the utterance. In this example, the sequence-to-sequence neural network can generate a word lattice 112 and confidence scores 118 and 120. The sequence-to-sequence neural network may not generate individual confidence scores for combinations of phonemes and words, as an acoustic model or a language model would. Instead, the sequence-to-sequence neural network may generate a confidence score that is a combination of the confidence score of the phonemes generated by the acoustic model and the confidence score of the combination of words generated by the language model.

[0017] Based on word lattice 112, computing device 106 identifies two candidate transcriptions of utterance 104. The first candidate transcription is "lights out" and the second candidate transcription is "flights out". The confidence score of the first candidate transcription is 0.65 and the confidence score of the second candidate transcription is 0.35. The confidence score of each candidate transcription may be the product of the confidence scores of each word of the candidate transcription. The phrase "fights out" may be the closest acoustic match to utterance 104, but based on the language model, the combination of "fight" and "out" is unlikely to occur. Thus, "fights out" is not a candidate for transcription. In implementations where a neural network is used instead of an acoustic model and a language model, the neural network may not generate individual confidence scores using the acoustic model and the language model.

[0018] Computing device 106 can determine the context 110 of computing device 106. The context 110 may be based on any combination of factors present on or around computing device 106. For example, the context 110 may include that the computing device 106 is located at user 102's home. The context 110 may include that the current time is 10 PM and the day of the week is Tuesday. The context 110 may include that the computing device 106 is a mobile phone running a digital assistant application in the foreground of the computing device 106.

[0019] In some implementations, the context 110 may include additional information. For example, the context 110 may include data regarding the orientation of computing device 106, such as being flat on a table or in the user's hand. The context 110 may include applications running in the background. The context may include audio output by computing device 106 or data displayed on the screen of the computing device prior to receiving utterance 104. For example, a display including a prompt such as "Hello, how can I help?" may indicate that computing device 106 is running a virtual assistant application in the foreground. The context 110 may include data stored on or accessible by computing device 106, such as weather, identification information of user 102, demographic data of user 102, and contacts.

[0020] Computing device 106 can use context to select a grammar to apply to a transcription. A grammar can be any structure consisting of multiple words that can be described using a common notation such as Backus-Naur form. Each grammar may correspond to a particular user's intent. For example, the user's intent may be to issue a home automation command or a media playback command. An example of a grammar may include an alarm grammar. In the alarm grammar, the notation $DIGIT=(0|1|2|3|4|5|6|7|8|9) can be used to define a digit as 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, or 0. The alarm grammar can define time using the notation $TIME=$DIGIT$DIGIT:$DIGIT$DIGIT(am|pm), which indicates that the time is two digits for the hour, followed by a colon, followed by two digits for the minute, followed by "am" or "pm". The alarm grammar can define the mode of the alarm using the notation $MODE=(alarm|timer), which indicates whether to set the alarm to alarm mode or timer mode. Finally, the alarm grammar can define the alarm syntax as $ALARM=set$MODEfor$TIME, which indicates that the user can say "set alarm for 6:00 am" or "set timer for 20:00". Computing device 106 uses the grammar to parse the spoken or typed command and identify the action that computing device 106 should perform. Thus, the grammar provides functional data that causes computing device 106 to operate in a particular manner to parse the command.

[0021] Each grammar may become active in a specific context. For example, if the application running in the foreground of computing device 106 is an alarm application, the alarm grammar may be active. Furthermore, if the application running in the foreground of the computing device is a digital assistant application, the alarm grammar may become active. Since the user is more likely to try to set an alarm with the alarm application than with the digital assistant application, computing device 106 may assign a probability of 0.95 to the possibility of the alarm grammar that matches the command and the user's intention when computing device 106 is running the alarm application, and may assign a probability of 0.03 when computing device 106 is running the digital assistant application.

[0022] In the example shown in FIG. 1, computing device 106 determines a confidence score for each grammar based on different candidate transcriptions that can be applied to the grammar. In other words, computing device 106 identifies the grammar that parses each of the candidate transcriptions. The applied grammar 124 indicates the grammar that parses the candidate transcription "lights out", thereby showing a grammar confidence score that each grammar is the correct grammar for the user's intention. A higher grammar confidence score indicates that the grammar is likely to be based on a given transcription and context. When the transcription is "lights out", the grammar confidence score 126 of the home automation grammar is 0.35. When the transcription is "lights out", the grammar confidence score 128 of the movie command grammar is 0.25. When the transcription is "lights out", the grammar confidence score 130 of the music command grammar is 0.20. The sum of the grammar confidence scores for each transcription should be 1.0. In that case, computing device 106 can assign the remaining grammar confidence score to the default intent grammar. In this example, when the transcription is "lights out", the grammar confidence score 132 of the default intent grammar is 0.20. The default intent grammar may not be limited to a specific action or intent and can parse all or almost all transcriptions.

[0023] In some implementations, none of the grammars may be able to parse a particular transcription. In this case, the default intent grammar is applied. For example, there may not be a grammar to parse "flights out". Thus, when the transcription is "flights out", the grammar confidence score 132 of the default intent grammar is 1.0.

[0024] In some implementations, computing device 106 selects grammar and transcription based on calculating the product of a transcription confidence score and a grammar confidence score. Box 134 shows the result of selecting a transcription and grammar based on the product of the transcription confidence score and the grammar confidence score. For example, the combined confidence score 136 for the combination of "lights out" and the home automation grammar is 0.2275. The combined confidence score 138 for the combination of "lights out" and the movie command grammar is 0.1625. The combined confidence score 140 for the combination of "lights out" and the music command grammar is 0.130. The combined confidence score 142 for the combination of "lights out" and the default intent grammar is 0.130. The combined confidence score 144 for the combination of "flights out" and the default intent grammar is 0.350. In this implementation, the combined confidence score 144 is the highest score. Thus, using this approach, computing device 106 can execute an action corresponding to the command "flights out". This is likely an inappropriate result for user 102 and may result in the search engine providing the transcription "flights out".

[0025] If it is not the intention of user 102 to search the Internet for "flights out", the user needs to repeat utterance 104. By doing so, it becomes necessary for computing device 106 to consume additional computing and power resources to process the additional utterance. User 102 may manually enter the desired command into computing device 102, which can use the additional processing and power resources, by keeping the screen of computing device 102 active for an additional period.

[0026] To reduce the possibility that computing device 106 performs an action that does not match the user's intention, computing device 106 can normalize the grammar confidence score of the applied grammar 124. To normalize the grammar confidence score of the applied grammar 124, computing device 106 can calculate the factor necessary to increase the highest grammar confidence score to 1.0. In other words, the product of the factor and the highest grammar confidence score should be 1.0. Next, computing device 106 can multiply the other grammar confidence scores by the factor to normalize the other grammar confidence scores. The normalization performed by computing device 106 may be different from conventional probability normalization that adjusts the sum of probabilities to 1. Since the highest grammar confidence score is increased to 1.0, the sum of the normalized probabilities does not become 1. The increased confidence score may represent a pseudo-probability rather than a probability in the conventional sense. Similar pseudo-probabilities may also be generated in the normalization process described elsewhere in this application.

[0027] As shown in the normalized grammar confidence score 146, the computing device 106 can identify the grammar confidence score of the home automation grammar as the highest grammar confidence score. To increase the grammar confidence score 126 to 1.0, the computing device multiplies the grammar confidence score 126 by 1.0 / 0.35 = 2.857. The computing device 106 calculates the normalized grammar confidence score 150 by multiplying the grammar confidence score 128 by 2.857. The computing device 106 calculates the normalized grammar confidence score 152 by multiplying the grammar confidence score 130 by 2.857. The computing device 106 calculates the normalized grammar confidence score 154 by multiplying the grammar confidence score 132 by 2.857. The grammar confidence score of the default intent grammar when the transcription is "flights out" is 1.0, because there is no other grammar to parse the transcription "flights out". Therefore, when the transcription is "flights out", since the score is already 1.0, there is no need to normalize the grammar confidence score of the default intent grammar.

[0028] As shown in box 134, instead of multiplying the grammar confidence score of box 124 by the transcription confidence score of the word lattice 112, the computing device 106 uses the transcription confidence score from the word lattice 112 and the normalized grammar confidence score to calculate the combined confidence score. In particular, the computing device 106 multiplies each of the normalized grammar confidence scores by the transcription confidence score of each transcription. When only the default intent grammar is applied to the grammar, since 1.0 is multiplied by the transcription confidence score, the corresponding transcription confidence score does not change.

[0029] As shown in box 156, computing device 106 calculates combined confidence score 158 by multiplying grammar confidence score 148, normalized by the transcription confidence score for "lights out", to obtain a result of 0.650. Computing device 106 calculates combined confidence score 160 by multiplying grammar confidence score 150, normalized by the transcription confidence score for "lights out", to obtain a result of 0.464. Computing device 106 calculates combined confidence score 162 by multiplying grammar confidence score 152, normalized by the transcription confidence score for "lights out", to obtain a result of 0.371. Computing device 106 calculates combined confidence score 164 by multiplying grammar confidence score 154, normalized by the transcription confidence score for "lights out", to obtain a result of 0.371. Since the computing device 106 does not have a grammar to parse the transcription as "flights out", it maintains the transcription confidence score for "flights out" at 0.350.

[0030] In some implementations, the computing device 106 can adjust the confidence score within box 156 to be able to accommodate likely grammars in view of the current user context of the current context of the computing device 106. For example, user 102 may be listening to music using the computing device 106. The computing device 106 can be a media device that plays music. In this case, the computing device 106 can adjust the confidence score 162. The computing device 106 can increase the confidence score by multiplying the confidence score by a factor, by assigning a preset value to the confidence score, or by another technique. For example, the current probability of a music command is 0.9 based on the context of user 102 listening to music via the computing device 102 which is a media device. The probability of the music command within box 124 can be the confidence score 130 which is 0.20. In this case, the computing device 106 can multiply the confidence score 162 by a ratio of 0.9 / 0.2 = 4.5. The resulting confidence score 162 is 0.371 * 4.5 = 1.67. The other confidence scores in box 156 can be adjusted at a similar ratio for each respective confidence score. For example, the current probability of a home automation command may be 0.04. In this case, the ratio is 0.04 / 0.35 = 0.16 which is 0.04 divided by the confidence score 126. The computing device 106 can multiply the confidence score 158 by 0.16 to calculate an adjusted or biased confidence score of 0.10. In this case, the highest confidence score will correspond to the music command.

[0031] In some implementations, this additional adjustment step may affect the candidate transcript selected by computing device 106. For example, if the candidate transcript "flights out" is the name of a video game and the user is expected to launch the video game, computing device 106 can adjust the confidence score 166 to 0.8 based on a ratio similar to that calculated above and using the probability that the user will launch the video game, which is based on the current context of computing device 106 and / or user 102.

[0032] In some implementations, such as when multiple grammars parse the same candidate transcript, computing device 106 can improve the detection of the user's intent and speech recognition without using this rescoring step. This rescoring step may enable computing device 106 to select a different candidate transcript that is not considered likely based on the speech recognition confidence score.

[0033] In some implementations, computing device 106 can use this rescoring step without considering the confidence score of box 124. For example, the computing device can apply the rescoring step to the transcript confidence score identified from the word lattice 112. In this case, computing device 106 is not considered to perform the adjustments shown in both box 146 and box 156.

[0034] Computing device 106 selects the grammar and transcription with the highest combined confidence score. For example, a home automation transcription and the transcription "lights out" may have the highest combined confidence score of 0.650. In this case, computing device 106 executes "lights out" based on the home automation grammar. Computing device 106 can turn off the lighting in the house where computing device 106 is located or in the house of user 102. If computing device 106 uses movie command grammar, computing device 106 can play the movie "Lights Out". If computing device 106 uses music command grammar, computing device 106 can play the song "Lights Out".

[0035] As shown in FIG. 1, the display of computing device 106 may indicate the action performed by computing device 106. First, computing device 106 can display a prompt 168 for the digital assistant. Computing device 106 receives utterance 104 and indicates by displaying prompt 170 that computing device 106 is turning off the lighting.

[0036] FIG. 2 shows the components of an exemplary system 200 that selects a grammar to apply to an utterance based on context. System 200 can be any type of computing device configured to receive and process voice audio. For example, system 200 can be similar to computing device 106 of FIG. 1. The components of system 200 can be implemented on a single computing device or distributed across multiple computing devices. System 200 implemented on a single computing device may be beneficial for privacy reasons.

[0037] System 200 includes an audio subsystem 205. The audio subsystem 205 may include a microphone 210, an analog-to-digital converter 215, a buffer 220, and various other audio filters. The microphone 210 can be configured to detect ambient sounds such as voices. The analog-to-digital converter 215 can be configured to sample the audio data detected by the microphone 210. The buffer 220 can store the sampled audio data for processing by the system 200. In some implementations, the audio subsystem 205 can be continuously active. In this case, the microphone 210 may always be detecting sound. The analog-to-digital converter 215 can always sample the detected audio data. The buffer 220 can store the most recently sampled audio data, such as the last 10 seconds of sound. If other components of the system 200 do not process the audio data in the buffer 220, the buffer 220 can overwrite the previous audio data.

[0038] The audio subsystem 205 provides the processed audio data to the speech recognition unit 225. The speech recognition unit provides the audio data as input to the acoustic model 230. The acoustic model 230 may be trained to identify phonemes that may correspond to sounds in the audio data. For example, when the user says "set", the acoustic model 230 may identify phonemes corresponding to the "s" sound, the "e" vowel, and the "t" sound. The speech recognition unit 225 provides the identified phonemes as input to the language model 235. The language model 235 generates a word lattice. The word lattice includes a term confidence score for each of the candidate terms identified by the language model 235. For example, the word lattice may indicate that the first term is likely to be "set". The language model 235 may not identify other possible terms for the first term. The language model 235 may identify two possible terms for the second term. For example, the word lattice may include the terms "time" and "chime" as possible second terms. The language model 235 can assign a term confidence score to each term. The term confidence score for "time" may be 0.75, and the term confidence score for "chime" may be 0.25.

[0039] The speech recognition unit 225 can generate a candidate transcription based on the word lattice. Each candidate transcription may have a transcription confidence score that reflects the likelihood that the speaker spoke the terms in the transcription. For example, the candidate transcription may be "set time" and may have a transcription confidence score of 0.75. Another candidate transcription may be "set chime" and may have a transcription confidence score of 0.25.

[0040] While the speech recognition unit 225 generates a word lattice, candidate transcriptions, and a transcription confidence score, the context determination unit 240 may collect context data indicating the current context of the system 200. The context determination unit 240 can collect sensor data 245 from any sensor of the system. The sensors may include position sensors, thermometers, accelerometers, gyroscopes, gravity sensors, time and day of the week, and other similar sensors. The sensor data 245 may also include data related to the status of the system 200. For example, the status may include the battery level of the system 200, signal strength, nearby devices that the system 200 may be communicating with or recognizing, and other similar states of the system 200.

[0041] The context determination unit 240 can also collect system process data 250 indicating the processes being executed by the system 200. The process data 250 may indicate the memory allocated to each process, the process resources allocated to each process, the applications being executed by the system 200, the applications being executed in the foreground or background of the system 200, the content of the interface on the system's display, and similar process data.

[0042] As an example, the context determination unit 240 can receive system process data 250 and sensor data 245 indicating that the system 200 is at the user's home, the time is 6:00 PM, the day of the week is Monday, the foreground application is a digital assistant application, and the device is a tablet.

[0043] The grammar score generation unit 255 receives the context of the system 200 from the context identification unit 240 and receives a word lattice from the speech recognition unit 225. The grammar score generation unit 255 identifies a grammar 260 that parses each of the candidate transcriptions of the word lattice. In some cases, none of the grammars 260 parse the candidate transcription. In this case, the grammar score generation unit 255 assigns a grammar confidence score of 1.0 of the default intent grammar to the candidate transcription that cannot be parsed by any of the grammars 260.

[0044] In some examples, a single grammar 260 can parse the candidate transcription. In this case, the grammar score generation unit 255 may determine the grammar confidence score of the single grammar assuming that the candidate transcription is the actual transcription. Since the grammar confidence score represents a probability, the grammar confidence score may be less than 1.0. The grammar score generation unit 255 assigns the difference between 1.0 and the grammar confidence score to the grammar confidence score of the default intent grammar of the candidate transcription.

[0045] In some examples, multiple grammars 260 can parse the candidate transcription. In this case, the grammar score generation unit 255 may determine the grammar confidence score of each of the multiple grammars assuming that the candidate transcription is the actual transcription. Since the grammar confidence score represents a set of probabilities, the sum of the grammar confidence scores may be less than 1.0. The grammar score generation unit 255 assigns the difference between 1.0 and the sum of the grammar confidence scores to the grammar confidence score of the default intent grammar of the candidate transcription.

[0046] The grammar score normalization unit 265 receives the grammar confidence scores for each of the candidate transcriptions and normalizes those grammar confidence scores. In some implementations, the grammar score normalization unit 265 normalizes only the grammar confidence scores of grammars other than the default intent grammar. When the grammar score generation unit 255 generates one grammar confidence score for one grammar for a particular candidate transcription, the grammar score normalization unit 265 increases the grammar confidence score to 1.0. When the grammar score generation unit 255 does not generate grammar confidence scores for a plurality of grammars for a particular candidate transcription, the grammar score normalization unit 265 maintains the grammar confidence score of the default intent grammar at 1.0.

[0047] When the grammar score generation unit 255 generates a plurality of grammar confidence scores for each of a plurality of grammars for a particular candidate transcription, the grammar score normalization unit 265 identifies the highest grammar confidence score for the particular candidate transcription. The grammar score normalization unit 265 calculates a coefficient for increasing the highest grammar confidence score to 1.0 such that the product of the coefficient and the highest grammar confidence score is 1.0. The grammar score normalization unit 265 increases the other grammar confidence scores for the same particular transcription by multiplying each of the other grammar confidence scores by the coefficient.

[0048] The grammar and transcription selection unit 270 receives the normalized grammar confidence scores and the transcription confidence scores. Using both the grammar confidence score and the transcription confidence score, the grammar and transcription selection unit 270 identifies the grammar and transcription that are likely to closely match the received utterance and the speaker's intent by calculating a combined confidence score. The grammar and transcription selection unit 270 determines each combined confidence score by calculating the product of each normalized grammar confidence score and the transcription confidence score of the corresponding candidate transcription. When the default intent grammar is the only grammar for a particular candidate transcription, the grammar and transcription selection unit 270 maintains the transcription confidence score as the combined confidence score. The grammar and transcription selection unit 270 selects the grammar and candidate transcription having the highest combined confidence score.

[0049] The action identification unit 275 receives the selected grammar and the selected transcription from the grammar and transcription selection unit 270, and identifies the actions that the system 200 performs. The selected grammar may indicate the type of action, such as setting an alarm, sending a message, playing a song, calling a person, or other similar actions. The selected transcription may indicate the details of the type of action, such as the time to set the alarm, the recipient of the message, the song to play, the called party, or other similar details of the action. The grammar 260 may include information regarding the type of action of a particular grammar. The grammar 260 may include information on how the action identification unit 275 should parse the candidate transcription to determine the details of the type of action. For example, the grammar may be the $alarm ($ALARM) grammar, and the action identification unit 275 may parse the selected transcription and determine to set the timer to 20 minutes. The action identification unit 275 can execute the action or provide instructions to another part of the system 200, such as a processor.

[0050] In some implementations, the user interface generation unit 280 displays an indication of the action or an indication of the execution of the action, or both. For example, the user interface generation unit 280 can display a timer that counts down from 20 minutes, or a confirmation of turning off the lighting in the house. In some examples, the user interface generation unit 280 may not provide an indication of the executed action. For example, the action may be adjusting a thermostat. The system 200 may adjust the thermostat without generating a user interface for display on the system 200. In some implementations, the user interface generation unit 280 can generate an interface for the user to interact with or confirm the action. For example, the action may be calling the mother. The user interface generation unit 280 can generate a user interface for the user to confirm the action of calling the mother before the system 200 executes the action.

[0051] Figure 3 is a flowchart showing an exemplary process 300 for selecting grammar based on context. Generally, process 300 performs speech recognition on audio and identifies actions to be performed based on the transcription of the audio and the grammar that parses the audio. Process 300 normalizes the grammar confidence score to identify the actions that are most likely to be intended by the speaker. Process 300 is described as being performed by a computer system that includes one or more computers, such as computing device 106 of FIG. 1 or system 200 of FIG. 2.

[0052] The system receives audio data of the utterance (310). For example, the user may speak an utterance that sounds like "lights out" or "flights out". The system detects the utterance via a microphone or receives the audio data of the utterance. The system can process the audio data using an audio subsystem.

[0053] The system uses an acoustic model and a language model to generate a word lattice that includes multiple candidate transcriptions of the utterance and multiple transcription confidence scores, where each of the multiple transcription confidence scores reflects the likelihood that each candidate transcription matches the utterance (320). The system generates the word lattice using automatic speech recognition. The automatic speech recognition process may include providing the audio data as input to an acoustic model that identifies various phonemes that match each portion of the audio data. The automatic speech recognition process may include providing the phonemes as input to a language model that generates a word lattice that includes the confidence score for each candidate word in the utterance. The language model selects the words of the word lattice from a vocabulary. The vocabulary may include words of the language configured for the system to recognize. For example, the system may be configured for English and the vocabulary may include English words. In some implementations, the actions of process 300 do not include restricting the words within the vocabulary that the system can recognize. In other words, process 300 can generate a transcription using any word of the language configured for the system to recognize. Since the system can recognize each word in the language, the system's automatic speech recognition process can function as long as it is within the language configured for the system to recognize when the user says something unexpected.

[0054] The system can use the word lattice to generate various candidate transcriptions. Each candidate transcription may include a different transcription confidence score that reflects the likelihood that the user spoke the transcription. For example, the confidence score for the candidate transcription "lights out" may be 0.65 and the confidence score for the candidate transcription "flights out" may be 0.35.

[0055] The system determines the context of the system (330). In some implementations, the context is based on data stored in or accessible by the system, such as the location of the system, the application running in the foreground of the computing device, the demographic information of the user, contact data, previous user queries or commands, time, date, day of the week, weather, the orientation of the system, and other similar types of information.

[0056] The system identifies a plurality of grammars corresponding to a plurality of candidate transcriptions based on the context of the system (340). The grammar may include various structures for various commands that the system can execute. For example, the grammar may include command structures for setting an alarm, playing a movie, performing an Internet search, checking the user's calendar, or other similar actions.

[0057] The system determines a plurality of grammar confidence scores that reflect the likelihood that each respective grammar matches each respective candidate transcription based on the current context (350). The grammar confidence score can be considered as the conditional probability that the grammar matches the speaker's intention based on the context of the system, assuming that one of the candidate transcriptions is the transcription of the utterance. For example, the grammar confidence score of the home automation grammar when the transcription is "lights out" may be 0.35. In other words, assuming the transcription is "lights out", the conditional probability of the home automation grammar is 0.35. The system may generate a grammar confidence score for each grammar for which it parses the candidate transcription. In some implementations, due to the context, the system may generate grammar confidence scores only for some of the grammars for which it parses the candidate transcription. The system may assign the remaining probability to the default intent grammar such that the sum of the grammar confidence scores for each candidate transcription is 1.0. out)”, the conditional probability of the home automation grammar is 0.35. The system may generate a grammar confidence score for each grammar for which it parses the candidate transcription. In some implementations, due to the context, the system may generate grammar confidence scores only for some of the grammars for which it parses the candidate transcription. The system may assign the remaining probability to the default intent grammar such that the sum of the grammar confidence scores for each candidate transcription is 1.0.

[0058] The system selects one candidate transcription from a plurality of candidate transcriptions based on a transcription confidence score and a grammar confidence score (360). In some implementations, the system normalizes the grammar confidence score by increasing the number of candidate transcriptions that match the plurality of grammars. For each candidate transcription, the system can multiply the highest grammar confidence score by a factor necessary to normalize the grammar confidence score to 1.0. The system can multiply the same factor for other grammar confidence scores for the same candidate transcription. The system can multiply the transcription confidence score by the normalized grammar confidence score to generate a combined confidence score. The system selects the grammar and transcription with the highest combined confidence score.

[0059] The system provides the selected candidate transcription as the transcription of the utterance for output (370). In some implementations, the system performs an action based on the grammar and the candidate transcription. The grammar may indicate the action to take. The action may include playing a movie, calling a contact, turning on the lighting, sending a message, or other similar types of actions. The selected transcription may include the contact to call, the message to send, the recipient of the message, the name of the movie, or other similar details.

[0060] In some implementations, the system uses a finite state transducer (FST) to tag the word lattice of a given set of grammars. In some implementations, the system may constrain the grammar as the union of all grammars that match the current context. In some implementations, the system may perform the collation offline in advance or dynamically at runtime.

[0061] In some implementations, process 300 can assign weights to the edges of a word lattice. The system can compile the grammar into a weighted finite state transducer that enforces the grammar. The weight of an arc, i.e., the weight of an edge of this finite state transducer, may encode the amount of probability of a word of a given grammar, which may be a negative logarithmic weight. In some implementations, all of almost all grammars associated with the word lattice may be integrated together.

[0062] Process 300 can continue with a system that determines the context-dependent probability of each grammar, which, given the context, can be the probability of the grammar. The system can determine the context-dependent probability by the provided components and corresponding weights, which may be encoded in an arc or edge in a grammar constraint section such as an opening decorator arc.

[0063] Process 300 can continue with a system that clears the weights of the word lattice. The system may use a grammar constraint section to construct the word lattice, whereby spans that match the grammar may be marked. For example, a span that matches the grammar may be surrounded by start and end decorator tags such as <media_commands>play song< / media_commands> for the span (or a transcription of "play song").

[0064] Process 300 may continue to normalize the probabilities. The system generates a copy of the lattice to which tags with the decorator tags removed are attached. The system determines and minimizes the arc costs of the tropical semiring, resulting in each unique word path being obtained, and the word path encodes the highest probability through the word path. The probability may be the probability obtained by multiplying the probability of the word path by the probability of the grammar given the grammar. The probability can be inverted by inverting the sign of the arc weights using negative logarithm weights, and this inverted lattice may be composed of tagged lattices. Composing with the inverted lattice may be substantially equivalent to performing division by the weights of that lattice (in this case, for each word sequence or path, multiplying the maximum amount of the probability of the word path by the probability of the grammar given the grammar). Thus, the system divides the probability of the optimal tagged path by itself, and each may become 1.0. In some implementations, non-optimal paths receive a lower pseudo-probability and generate a lattice that includes the desired amount that is the pseudo-probability of the grammar for a given word path.

[0065] In some implementations, process 300 may include discarding grammars with low conditional probabilities by using beam pruning. For example, when there are too many grammars that match the utterance "lights out" (where the probability of home automation is 0.32, movie command 0.25, music command 0.20, general search 0.10, online business 0.03, social group 0.03, and online encyclopedia search 0.02). The system can reduce the processing burden by removing grammars that fall below a threshold or below a threshold percentage of the most likely interpretation, i.e., the grammar. For example, the system may remove grammars with a probability less than one-tenth of the most likely grammar. Next, the system may remove grammars with a probability less than 0.032 for a given transcription. The theoretical basis behind pruning is that since the probability of low-scoring tagging is very low, there is no rationality that even with bias, that interpretation will be selected over a more likely interpretation. In pruning, the system may need to add a finite-state transducer that operates on the lattice obtained with the desired pruning weight, which specifies how late the best assumption that tagging has for removal is.

[0066] According to the advantages of process 300, it enables the system to bias grammars directly on the lattice. Process 300 has less complexity than other word lattice tagging algorithms. Process 300 resolves the probability fragmentation caused by tagging ambiguity by implementing maximum normalization. Process 300 enables discarding relatively unlikely proposed taggings by using beam pruning.

[0067] In some implementations, process 300 can extract the n best hypotheses from the lattice, individually tag all hypotheses and possibilities, and optionally recombine them back into the lattice at the end. The system may assume that the taggers are identical in other respects. This may result in the same outcome of increased latency and / or decreased recall. If the n best hypotheses after the top N (such as N = 100) are removed, recall may decrease. This may be useful for latency management. Since the number of individual sentences in the lattice can be exponential with respect to the number of words in it, some workarounds may be exponentially slow in the worst case unless the aforementioned top N limit (which may limit the number of options to examine) is used.

[0068] FIG. 4 shows an example of a computing device 400 and a mobile computing device 450 that may be used to implement the techniques described herein. Computing device 400 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. Mobile computing device 450 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, and other similar computing devices. The components shown in this specification, the connections and relationships between the components, and the functions of the components are only intended to be exemplary and are not intended to be limiting.

[0069] Computing device 400 includes a processor 402, a memory 404, a storage device 406, a high-speed interface 408 that connects to the memory 404 and a plurality of high-speed expansion ports 410, and a low-speed interface 412 that connects to a low-speed expansion port 414 and the storage device 406. Each of the processor 402, the memory 404, the storage device 406, the high-speed interface 408, the high-speed expansion ports 410, and the low-speed interface 412 are interconnected via various buses and may be implemented on a common motherboard or in other manners as needed. The processor 402 is capable of processing instructions for execution within the computing device 400, including instructions stored in the memory 404 or the storage device 406, to display graphic information of a GUI on an external input / output device such as a display 416 connected to the high-speed interface 408. In other implementations, multiple memories and types of memories, along with multiple processors and / or multiple buses as needed, may be used. Also, multiple computing devices may be connected and each device may provide a portion of the required operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).

[0070] The memory 404 stores information within the computing device 400. In some implementations, the memory 404 is one or more volatile memory units. In some implementations, the memory 404 is one or more non-volatile memory units. The memory 404 may also be another form of computer-readable medium, such as a magnetic disk or an optical disk.

[0071] The memory device 406 can provide the computing device 400 with large-capacity storage. In some implementations, the memory device 406 may be a floppy (registered trademark) disk device, a hard disk device, an optical disk device, or a tape device, a computer-readable medium such as a flash memory or other similar solid-state memory device, or an array of devices including a storage area network or other devices in other configurations, or may include the same. Instructions may be stored on an information carrier. When executed by one or more processing devices (e.g., the processor 402), the instructions perform one or more of the methods as described above. The instructions may be stored in one or more storage devices such as a computer-readable medium or a machine-readable medium (e.g., the memory 404, the memory device 406, or the memory on the processor 402).

[0072] The high-speed interface 408 manages the bandwidth-intensive operations of the computing device 400, while the low-speed interface 412 manages the more low-bandwidth-intensive operations. Such a function assignment is just an example. In some implementations, the high-speed interface 408 is connected to a high-speed expansion port 410 that can receive the memory 404, the display 416 (e.g., through a graphics processor or an accelerator), and various expansion cards (not shown). In an implementation, the low-speed interface 412 is connected to the memory device 406 and the low-speed expansion port 414. The low-speed expansion port 414 may include various communication ports (e.g., USB, Bluetooth (registered trademark), Ethernet (registered trademark), wireless Ethernet), and may be connected to one or more input / output devices such as a keyboard, a pointing device, a scanner, or a network device such as a switch or a router, for example, through a network adapter.

[0073] As shown in the figure, the computing device 400 may be implemented in many different forms. For example, it may be implemented as a standard server 420 or multiple times in a group of such servers. It may also be implemented in a personal computer such as a laptop computer 422. It may also be implemented as part of a rack server system 424. Alternatively, components from the computing device 400 may be combined with other components within a mobile device (not shown) such as a mobile computing device 450. Each of such devices may include one or more of the computing device 400 and the mobile computing device 450, and the entire system may be composed of multiple computing devices that communicate with each other.

[0074] Among other components, the mobile computing device 450 includes a processor 452, a memory 464, input / output devices such as a display 454, a communication interface 46, and a transceiver 468. The mobile computing device 450 may include a storage device such as a microdrive or other device to provide additional storage. The processor 452, the memory 464, the display 454, the communication interface 466, and the transceiver 468 are interconnected with each other via various buses, and the multiple components may be implemented on a common motherboard or in other ways as needed.

[0075] The processor 452 can execute instructions within the mobile computing device 450, including instructions stored in the memory 464. The processor 452 may be implemented as a chipset of separate and multiple analog and digital processors. The processor 452 may provide adjustments to other components of the mobile computing device 450, such as, for example, user interfaces, applications executed by the mobile computing device 450, and control of wireless communication by the mobile computing device 450.

[0076] Processor 452 may communicate with a user through a control interface 458 and a display interface 456 connected to a display 454. The display 454 may be, for example, a TFT (Thin Film Transistor Liquid Crystal Display) display, an OLED (Organic Light Emitting Diode) display, or other suitable display technology. The display interface 456 may include appropriate circuitry for driving the display 454 to present graphics and other information to the user. The control interface 458 may receive commands from the user and convert the commands for supply to the processor 452. Further, an external interface 462 may provide communication with the processor 452 to enable near-field communication between the mobile computing device 450 and other devices. The external interface 462 may provide, for example, wired communication in some implementations and wireless communication in other implementations, and multiple interfaces may be used.

[0077] Memory 464 stores information within computing device 450. Memory 464 can be implemented as one or more of a computer-readable medium, a volatile memory unit, or a non-volatile memory unit. Extended memory 474 may be provided and connected to mobile computing device 450 via expansion interface 472. Expansion interface 472 may include, for example, a SIMM (Single In-line Memory Module) card interface. Extended memory 474 may provide additional storage area for mobile computing device 450, or may store applications or other information of mobile computing device 450. Specifically, extended memory 474 may include instructions to execute or complement the above-described processes, and may include security information. Thus, for example, extended memory 474 may be provided as a security module of mobile computing device 450 and may be programmed with instructions to enable secure use of mobile computing device 450. Further, a security application may be provided along with additional information such as placing identification information on the SIMM card in a hacking-proof manner via the SIMM card.

[0078] The memory may include, for example, flash memory and / or NVRAM memory (non-volatile random access memory) as discussed below. In some implementations, the instructions are stored on an information carrier. When the instructions are executed by one or more processing devices (e.g., processor 452), one or more of the methods as described above are performed. The instructions may be stored in one or more storage devices such as one or more computer-readable media or one or more machine-readable media (e.g., memory 464, extended memory 474, or memory on processor 452). In some implementations, the instructions can be received as a propagated signal, for example, through transceiver 468 or external interface 462.

[0079] The mobile computing device 450 may communicate wirelessly through a communication interface 466 that may include a digital signal processing circuit as needed. The communication interface 466 may provide communication under various modes or protocols, such as, in particular, GSM (registered trademark) voice calls (Global System for Mobile Communications), SMS (Short Message Service), EMS (Enhanced Messaging Service), MMS messaging (Multimedia Messaging Service), CDMA (Code Division Multiple Access), TDMA (Time Division Multiple Access), PDC (Personal Digital Cellular), WCDMA (registered trademark) (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service). Such communication may occur, for example, through a transceiver 468 that uses radio frequencies. Additionally, short-range communication may occur through the use of Bluetooth, WiFi (registered trademark), or other such transceivers (not shown). Additionally, a GPS (Global Positioning System) receiver module 470 may provide additional navigation and location-related wireless data to the mobile computing device 450 that may be used as appropriate by applications running on the mobile computing device 450.

[0080] The mobile computing device 450 may communicate audibly using an audio codec 460, which may receive spoken information from a user and convert it into usable digital information. The audio codec 460 may similarly generate audible sounds for the user, for example, through a speaker within the handset of the mobile computing device 450. Such sounds may include the sounds of a voice call, may include recorded sounds (such as voice messages, music files, etc.), and may also include sounds generated by an application operating on the mobile computing device 450.

[0081] As shown in the figures, the mobile computing device 450 may be implemented in many different forms. For example, the mobile computing device 450 may be implemented as a mobile phone 480. The mobile computing device 450 may also be implemented as part of a smartphone 482, a personal digital assistant, or other similar mobile devices.

[0082] The various implementations of the systems and techniques described herein can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that are executable and / or translatable on a programmable system including at least one programmable processor coupled to receive and transmit data and instructions to and from a storage device, at least one input device, and at least one output device.

[0083] These computer programs (also known as programs, software, software applications or code) include machine language instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages, and / or in assembly / machine language. As used herein, the terms machine-readable medium and computer-readable medium refer to any computer program product, apparatus and / or device (e.g., magnetic disks, optical disks, memory, programmable logic device (PLD)) used to provide machine language instructions and / or data to a programmable processor, including a machine-readable medium that receives machine language instructions as a machine-readable signal. The term machine-readable signal refers to a signal used to provide machine language instructions and / or data to a programmable processor.

[0084] To provide interaction with a user, the systems and techniques described herein may be implemented on a computer including a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices may be used to provide interaction with the user. For example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and the input from the user may be captured in any form including acoustic input, voice input, or tactile input.

[0085] The systems and techniques described herein may be implemented in a computer system having backend components (e.g., a data server), a computer system having middleware components (e.g., an application server), a computer system having frontend components (e.g., a client computer having a graphical interface or a web browser by which a user can interact with an implementation of the systems and techniques described herein), or a computer system having any combination of such backend, middleware, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include LANs (local area networks), WANs (wide area networks), and the Internet. In some implementations, the systems and techniques described herein may be implemented on an embedded system in which speech recognition and other processing are performed directly on the device.

[0086] A computing system may include a client and a server. The client and the server are generally separated from each other and typically interact through a communication network. The relationship between the client and the server is created by computer programs that are executed on each computer and have a client-server relationship with each other.

[0087] Although several implementations have been described in detail, other modifications are possible. For example, the client application has been described as accessing a delegate, but in other implementations, the delegate may be used by other applications implemented by one or more processors, such as an application running on one or more servers. Further, the logical flow shown in the figures does not require the particular order or sequential order shown to obtain the desired result. Additionally, other operations may be provided to or removed from the described flow, and other components may be added to or removed from the described system. Accordingly, other implementations are within the scope of the following claims.

Claims

1. A method implemented by a computer, which, when executed by data processing hardware, receives audio data of speech uttered by a user and captured by a computing device associated with the user; determines a plurality of phonemes corresponding to the speech from the audio data; generates one or more candidate transcriptions for the speech using the plurality of phonemes corresponding to the speech, wherein each of the one or more candidate transcriptions has a corresponding transcription confidence score; determines a context indicating that the user is listening to music through the computing device; selects a grammar corresponding to a specific user intention of issuing a media playback command based on the context indicating that the user is listening to music through the computing device; performs parsing on a candidate transcription among the one or more candidate transcriptions having the highest transcription confidence score using the selected grammar corresponding to the specific user intention of issuing the media playback command to identify an action to be executed by the computing device; causes the data processing hardware to execute operations including the above; wherein the grammar includes a specified structure consisting of a plurality of terms.

2. The operations further include instructing the computing device to execute the identified action. The method according to claim 1.

3. The operations further include selecting the grammar from a plurality of grammars based on the context of the computing device. The method according to claim 1.

4. The grammar is a music command grammar. The operations further include selecting the grammar from a plurality of grammars based on the context of the computing device. The plurality of grammars include the music command grammar and a movie command grammar. Identifying the action to be performed by the computing device includes, based on the candidate transcription among the one or more candidate transcriptions having the highest transcription confidence score, identifying, when using the music command grammar, playing the song indicated by the candidate transcription, and when using the movie command grammar, identifying playing the movie indicated by the candidate transcription. The method according to claim 1.

5. The computing device includes a mobile phone. The method according to claim 1.

6. The computing device includes a wearable device. The method according to claim 1.

7. Each of the one or more candidate transcriptions includes a plurality of terms. The method according to claim 1.

8. Each of the plurality of terms includes a corresponding term confidence score. The method according to claim 7.

9. The transcription confidence score indicates the likelihood that the candidate transcription matches the utterance made by the user. The method according to claim 1.

10. The data processing hardware is on the computing device. The method according to claim 1.

11. A system comprising: data processing hardware; and memory hardware that communicates with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, receive audio data of an utterance made by a user and captured by a computing device associated with the user; determine a plurality of phonemes corresponding to the utterance from the audio data; generate one or more candidate transcriptions for the utterance using the plurality of phonemes corresponding to the utterance, each of the one or more candidate transcriptions having a corresponding transcription confidence score; determine a context indicating that the user is listening to music through the computing device; select a grammar corresponding to a specific user intent to issue a media playback command based on the context indicating that the user is listening to music through the computing device. For the candidate transcript among the one or more candidate transcripts having the highest transcription confidence score, parsing is performed using the selected grammar corresponding to the intention of the specific user to issue the media playback command, and identifying an action to be performed by the computing device; causing the data processing hardware to execute operations including; The grammar includes a specified structure consisting of a plurality of terms, a system.

12. The operations further include instructing the computing device to execute the identified action. The system according to claim 11.

13. The operations further include selecting the grammar from a plurality of grammars based on the context of the computing device. The system according to claim 11.

14. The grammar is a music command grammar, The operations further include selecting the grammar from a plurality of grammars based on the context of the computing device, The plurality of grammars include the music command grammar and the movie command grammar, Identifying the action to be performed by the computing device includes, based on the candidate transcript among the one or more candidate transcripts having the highest transcription confidence score, identifying playing the song indicated by the candidate transcript when using the music command grammar, while identifying playing the movie indicated by the candidate transcript when using the movie command grammar. The system according to claim 11.

15. The computing device includes a mobile phone or a wearable device. The system according to claim 11.

16. The grammar includes a default grammar. The system according to claim 11.

17. Each of the one or more candidate transcripts includes a plurality of terms. The system according to claim 11.

18. Each of the plurality of terms includes a corresponding term confidence score. The system according to claim 17.

19. The transcription confidence score indicates the likelihood that the candidate transcript matches the utterance made by the user. The system according to claim 11.

20. The data processing hardware is on the computing device. The system according to claim 11.

Citation Information

Patent Citations

  • Speech recognizing device

    JP1994102896A

  • Developer voice actions system

    WO2017151215A1