Processing intent of a pause in speech
The apparatus with a streaming-language processor and hesitation model addresses out-of-domain utterances by dynamically extending time-out periods and prompting intent prediction, improving user interaction and reducing latency in digital assistants.
Patent Information
- Application Number
- PCT/US2024/058419
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2024-12-04
- Publication Date
- 2025-07-17
AI Technical Summary
Modern digital assistants in vehicles often produce out-of-domain utterances due to syntactically unresolved pauses or incomplete utterances, leading to user frustration and prolonged latency.
An apparatus with a streaming-language processor and hesitation model that processes utterances in real time, dynamically extends time-out periods based on user hesitation, and prompts intent prediction to resolve incomplete commands.
Reduces the likelihood of out-of-domain utterances by providing timely prompts and suggestions, enhancing user interaction and reducing latency.
Smart Images

Figure US2024058419_17072025_PF_FP_ABST
Abstract
Description
PROCESSING INTENT OF A PAUSE IN SPEECHCross-Reference to Related Applications
[0001] This application claims the benefit of the January 9, 2024 filing date of U.S. Provisional Application No. 63 / 618,942, the content of which is hereby incorporated by reference in its entirety.Background
[0002] Modern digital assistants are able to use a voice interface. This is particularly useful in vehicles because a user who attempts to use a non-voice interface while driving the vehicle will tend to be distracted by the effort of doing so.
[0003] A user’s utterance falls into one of two classes: those that can be responded to and those that cannot. These are referred to as “in-domain” and “out-of-domain” utterances, respectively.
[0004] To avoid frustrating a user, it is useful to avoid having too many out-of-domain utterances. Accordingly, it is useful to observe how out-of-domain utterances arise.
[0005] A significant cause of an out-of-domain utterance is a syntactically unresolved utterance that fails to resolve within a pre-defined time-out period. This failure to syntactically resolve the utterance often arises because of a pause in speech. Such a pause is either filled by silence or by non-lexical utterances that are commonly used to avoid dead airtime during speech (e.g., “um” and “er” in English) or by repetitions of previously uttered words. Another significant cause of an out-of-domain utterance is one that is semantically incomplete and therefore fails to include enough information to interpret.
[0006] For example, a user who says “Get me directions to go to...” may realize he does not have the address at his fingertips. As a result, the digital assistant is kept hanging while the user rummages for the address. In most cases, a digital assistant awaits resolution of this syntactically unresolved utterance for a brief pre-defined timeout period. But if the user fails to resolve it in the pre-defined time-out period, the utterance becomes invalid as a result of the digital assistant having realized that there is not enough information in the utterance for it to be interpreted. It is possible to simply extend the timeout period to avoid hanging up on the request. But if the time-out period is extendedfor too long, the user could experience a long latency while waiting for the digital assistant’s response.Summary
[0007] In one aspect, the invention features an apparatus for reducing occurrences of out- of-domain utterances resulting from static intervals within an utterance that is made by a user. Such an apparatus includes an infotainment system that is in the vehicle and that executes a digital assistant. The digital assistant includes a streaming-language processor and a hesitation model. The automotive assistant is configured to be executed by the infotainment system for receiving an utterance from a user and taking action based on the utterance and on information from the streaming-language processor. The streaminglanguage processor is configured to process the utterance in real time as it is being received from the user via a microphone. The utterance includes a growth interval, a static interval, and a transition from the growth interval to the static interval. The hesitation model is configured to interact with the streaming-language processor to determine whether a static interval represents either an intent to end the utterance or an intent to begin a new growth interval after the static interval. The streaming-language processor is configured to respond to detection of the second intent by promoting resolution of the utterance based on information from the hesitation model.
[0008] An utterance exists in either a first state or a second state. The second state is the state of being “resolved.” The first state is the state of being “unresolved.” An utterance undergoes a transition from the first state to the second state. Such a transition is called “resolution.” The time at which resolution occurs cannot be predicted. Indeed, it is possible that resolution will not occur. However, it is possible to interact with the user in a manner that increases the likelihood that a resolution will occur during a time interval that follows the end of an utterance that is in the first state. Such an interaction will be referred to herein as “promoting” or “catalyzing” the resolution of the utterance.
[0009] By way of example, which, being an example, is inherently not limiting, an utterance such as, “I need directions to...” is in the first state. It is possible that the utterance will eventually be resolved, given enough time. However, it is possible to interact with the utterer by saying something like, “Do you need directions to the arena?” Such an interaction is one that would reasonably be expected to promote or catalyze resolution of the utterance.
[0010] In some embodiments, the user is in a vehicle, in which case the digital assistant is an automotive assistant, and the microphone is in the vehicle. However, embodiments include those in which the apparatus operates in a smartphone, a smart speaker, or a device that has a speech interface.[Oil] Among the embodiments are those that include a timer. This timer starts a timeout period upon a transition from the growth interval to the static interval. The automotive assistant is configured to promote resolution by extending the time-out period to grant the user more time to begin the growth interval.
[0012] Among the embodiments that include such a timer are those in which the timer extends different static intervals by different amounts. In such embodiments, the growth interval is a first growth interval of the utterance, and the static interval is a first static interval of the utterance. The timer starts a first time-out period and a second time-out period. It starts the first time-out period when the utterance transitions from the first growth interval into the first static interval and starts the second time-out period when the utterance transitions from the second growth interval into the second static interval. The first time-out period and the second time-out period differ in length.
[0013] Also, among the embodiments that include a timer are those in which the timer starts a time-out period upon a transition from the growth interval to the static interval. In such embodiments, the streaming-language processor promotes resolution by extending the time-out period to grant the user more time to begin a new growth interval in response to the hesitation model having inferred that the static interval represents a hesitation by the user and that the user intends to begin this new growth interval after the static interval.
[0014] Also, among the embodiments that include a timer are those in which the timer starts a time-out period upon a transition from the growth interval to the static interval and the hesitation model infers that the static interval is intended to end the utterance. In such embodiments, the streaming-language processor ends the utterance prior to resolution thereof.
[0015] Still other embodiments include an intent-prediction model. Among these are embodiments in which the streaming-language processor promotes resolution of the utterance by prompting the intent-prediction model to infer a candidate intent and providing the candidate intent to the user for verification thereof.
[0016] Also among these are embodiments in which the streaming-language processor promotes resolution of the utterance by prompting the intent-prediction model to infer a candidate intent based on one or more of context data, user preferences, and historical records. Among these embodiments are those in which the intent-prediction model includes an intention auto-filler that is based on a large language model
[0017] In still other embodiments, the streaming-language processor comprises an automatic speech recognizer and a natural-language understander, both of which operate in streaming audio mode. The automatic speech recognizer partitions the utterance into partial speech-recognition results, each of which corresponds to a corresponding time interval within the utterance. It also provides the partial speech-recognition results sequentially to the natural -language understander.
[0018] Other embodiments include those in which the streaming-language processor comprises a segmentation model that derives segments of natural language from partial speech-recognition results that are provided thereto in real time, those in which the streaming-language processor comprises an intent separator that identifies first and second distinct overt intents in the utterance based at least in part on partial speechrecognition results being provided in real time, each of which represents a portion of the utterance, and those in which the streaming-language processor comprises a slot identifier. In these embodiments, the utterance has a syntax that includes a filled slot and an empty slot and the slot identifier identifies the empty slot that is later to be filled.
[0019] Still other embodiments include those that include a user interface that is configured to display an unresolved utterance and at least one candidate resolution for the unresolved utterance. This candidate resolution is one that has been inferred based on either user preferences, context information, or both. The user interface also includes a visual indicator to draw attention to said unresolved utterance and to prompt resolution thereof. Preferably, the visual indicator is animated so as to draw more attention to the need to resolve the utterance.
[0020] A method and apparatus as described and claimed herein reduces the likelihood of avoidable out-of-domain utterances and ultimately results in a more natural voice interaction that mimics that which one might experience with an actual human being. This includes mimicking a human being’s ability to predict a person’s intent based on content and to thus prompt the user with suggestions on how to resolve a syntactically incomplete utterance when a user’ s hesitation suggests a difficulty in articulating his intent.
[0021] These and other features of the invention will be apparent from the following detailed description and the accompanying figures, in which:Description of Drawings
[0022] FIG. 1 shows a vehicle that includes an automotive assistant that uses a streaming language processor to promote resolution of syntactically unresolved utterances
[0023] FIG. 2 shows a confidence profile with no static intervals prior to resolution of the utterance;
[0024] FIG. 3 shows a confidence profile like that in FIG. 2 but with a static interval occurring prior to resolution;
[0025] FIG. 4 shows a confidence profile with alternating static intervals and growth intervals.
[0026] FIG. 5 shows a user interface with an unresolved utterance and a candidate resolution for that utterance; and
[0027] FIG. 6 shows details of the streaming language processor of FIG. 1.Detailed Description
[0028] FIG. 1 shows a vehicle 10 having an infotainment system 12 in which an automotive assistant 20 receives voice input through a microphone 16 and provides voice output through a loudspeaker 18 and uses its streaming-language processor 14 to interact with a user 22.
[0029] The streaming-language processor 14 processes streaming speech. This means that the streaming-language processor 14 attempts to ascertain the user’s intent as the user 22 reveals it through overt acts, such as speech. This intent is thus considered to be an “overt” intent.
[0030] The raw material provided to the streaming-language processor 14 is an utterance 21. An utterance 21 is a combination of lexical content, i.e., words and non-lexical content, such as “er” and “urn” that are used as filler. The lexical content tends to advance the utterance 21 towards a state in which the streaming-language processor 14 understands the user’s intent. The non-lexical content tends to add nothing to the understanding of the user’s intent.
[0031] As an utterance 21 continues, the increasing amounts of lexical content acquired enable the streaming-language processor 14 to grow increasingly confident about the user’s intent. FIGS. 2-4 show how the confidence with which the streaming-language processor 14 is able to discern the user’s intent evolves over time. Relationships as shown in FIGS. 2-4 are referred to herein as “confidence profiles.” In all the foregoing figures, the vertical axis represents a “confidence score” that indicates the confidence with which the streaming-language processor 14 is able to discern the user’s intent based on the lexical content of the utterance 21 and whatever context is available.
[0032] FIG. 2 shows a simple confidence profile. At the utterance’s beginning, the streaming-language processor 14 has no idea what the user’s intent might be. Thus, the streaming-language processor’s confidence is at a minimum. As the user 22 speaks, the on-going utterance 21 provides the streaming-language processor 14 with progressively more information for resolving the user’s intent. This starts a growth interval 23, during which the streaming-language processor’s confidence grows. The growth interval 23 continues until time “T2.”
[0033] During the growth interval 23, the streaming-language processor 14 develops increasing confidence in its ability to discern the user’s intent. At time “Tl,” its confidence crosses a confidence threshold. As a result, the streaming-language processor 14 considers the utterance 21 to be “resolved.”
[0034] At time “T2,” the lexical content in the utterance 21 has dwindled and possibly stopped. As a result, the streaming-language processor’s confidence no longer changes with time. This results in a static interval 25 during which the streaming-language processor’s confidence no longer grows. However, since the utterance 21 has been resolved, this does not pose a difficulty. The streaming-language processor 14 is able to act on the user’s overt intent.
[0035] A notable feature of the confidence profile shown in FIG. 2 is the absence of any intra- utterance static intervals 25. This suggests a user 22 of considerable fluency, with speech unsullied by hesitations or repetitions.
[0036] In FIG. 3, a confidence profile indicates that an utterance 21 has failed to resolve. In this case, the static interval 25 occurs before the streaming-language processor’s confidence has reached the threshold. The static interval 25 can occur for a variety of reasons, such as cessation of speech or unintelligibility thereof.
[0037] In FIGS. 2 and 3, the transition between a growth interval 23 and a static interval 25 implies the existence of a “non-overt” intent, namely the user’s intent to end the utterance 21. The streaming-language processor 14 infers this by observing non-lexical information, such as the elapsed time since the end of the preceding growth interval 23, and various non-lexical cues, such as the space fillers that are sometimes used while a user 22 thinks of what to say next. The ability to also infer this non-overt intent is particularly useful for avoiding out-of-domain commands.
[0038] For example, in FIG. 2, it is possible to infer, from the onset of the static interval 25 after time “T2,” that the user 22 intends to end the utterance 21. After all, the streaming-language processor 14 has already divined the user’s overt intent. Similarly, in FIG. 3, one might infer the static interval 25 that begins at time “T1 ” as also manifesting an intent to end the utterance 21.
[0039] However, both of the foregoing inferences could be incorrect. In the case of FIG.2, it is quite possible that the static interval 25 has arisen only because the user 22 has simply paused to gather his thoughts on how to express a second intent. In the case of FIG. 3, it is quite possible that the static interval 25 has arisen because the user has paused to retrieve information required to resolve the utterance 21.
[0040] In both cases, the result is sub-optimal. In the case of FIG. 3, the error in inferring non-overt intent will require a new utterance 21 for the second intent. In the case of FIG.3, the premature truncation of the utterance 21 results in an unnecessary out-of-domain command. Both of these are likely to contribute to an adverse perception of the user experience.
[0041] In recognition of the possibility that a static interval 25 will ultimately prove to be a brief pause in the utterance 21, it is useful for the streaming-language processor 14 will wait for some pre-determined time-out period after the static interval 25 begins before inferring that the user intended to end the utterance 21.
[0042] FIG. 3 shows a confidence profile that has been found to be common in practice. It turns out that many users 22 do not plan ahead when expressing overt intent. This results in static intervals 25 alternating with growth intervals 23.
[0043] During a static interval 25, the user 22 hesitates, either with a cessation of speech altogether or, more commonly, by uttering filler words, During growth intervals 23, the user 22 advances the utterance 21 towards resolution. In FIG. 4, there are two static intervals 23 of differing length. However, in the end, the user 22 continues to articulatehis intent during growth intervals 23 after each static interval 25 until the utterance 21 achieves resolution during a final growth interval 23.
[0044] Once a static interval 25 begins, there will be two possible non-overt outcomes;(1) the user 22 will end the static interval 25 and begin another growth interval 23 and (2) the user 22 will never end the static interval 25. The streaming-likelihood processor 14 in FIG. 1 is configured to assess the likelihood of each of these outcomes and to promote or catalyze resolution of the utterance 21 when the streaming-language processor 14 determines that the user 22 is likely to begin another growth interval 23.
[0045] Referring back to FIG. 1, the streaming-language processor 14 includes a timer 24, a streaming controller 26, which provides full duplex streaming audio communication with the user 22, an automatic speech-recognizer 28 and a natural-language understander 30, the latter two being configured to process streaming audio.
[0046] Based on information provided by the hesitation model 32, the streaming controller 26 recognizes that the user 22 is in the midst of a static interval 25. Under such circumstances, the streaming controller 26 signals the automatic speech-recognizer 28 to dynamically extend a time-out period set by the timer 24. This extended time-out period reduces the likelihood of an out-of-domain command by giving the user 22 more time to resolve the utterance 21.
[0047] In some embodiments, the extension added to the time-out period is a fixed value. It has been determined experimentally that three seconds avoids excessive latency while capturing most transitions from a static interval 25 back to a growth interval 23. In other embodiments, the length of the extension depends on the user’s historical speech patterns. For example, for a user 22 who is habitually laconic in speech or who has a speech impediment, the extension may be on the order of four seconds.
[0048] In some cases, a user will resolve an utterance before the end of a time-out period. In such cases, a brief hesitation within the time-out period is irrelevant since the utterance will have resolved by the end of the time-out period anyway. Therefore, to avoid unnecessary computation, it is useful to have the streaming controller 26 wait until close to the end of a time-out interval before invoking the hesitation model 32 but far enough from the end of the time-out interval for the hesitation model 32 to provide useful results. For example, for a time-out interval of length T, it is useful to wait for %-r before invoking the hesitation model 32 so that the hesitation model can provide useful resultswithin the remaininglA-x and allow the controller 26 to process and the timer 24 to extend the time-out interval.
[0049] An example of a static interval 25 is one that arises when a user 22 utters, “Get directions to.. and pauses, for example as a result of having temporarily forgotten the name of the desired location. In this case, analysis of the utterance’s syntactic structure in real time results in the recognition of an imperative mood verb (i.e., “get”) followed by its direct object (i.e., “directions”), which is then followed by a preposition (“to”) that starts a prepositional phrase that is missing its object. The hesitation model 32, having been trained based on the user’s speech pattern, recognizes a nascent hesitation and communicates this fact to the streaming controller 26. The automatic speech-recognizer 28 is then notified that the utterance 21 is pending resolution and, if necessary, the timer24 is made to extend a time-out period to allow the user 22 time to complete the utterance. Embodiments include those in which the streaming controller 26 causes these events to occur.
[0050] To determine an extent to which the time-out period should be extended, the streaming controller 26 relies on a configurable hesitation model 32. This hesitation model 32 is one that has been trained based on patterns observed in actual in-car utterances 21 by the relevant user 22. The hesitation model 32 thus provides the streaming controller 26 with a basis for predicting an expected value of a static interval25 based on the content of the utterance 21 preceding that static interval 25 and any non- lexical cues occurring during the static interval 25. In a preferred embodiment, the hesitation model 32 is a transformer-based model that detects user hesitation and provides a basis for dynamically extending the time-out interval upon detection of hesitation.
[0051] In some implementations, the hesitation model 32 uses observed user speech patterns to construct a probability distribution for hesitation-interval lengths and uses that distribution as a basis for determining how long to extend a static interval 25.
[0052] The natural-language understander 30 provides information concerning the detected hesitation to the streaming controller 26, which uses it as a basis for providing the timer 24 with a suitable extension period.
[0053] The technical effect of the foregoing interaction is that of implementing a dynamic feedback loop in which a controlled variable, i.e., the length of a time-out period, depends on what has gone on before, i.e., the syntactic structure of preceding speech and historical speech patterns as embodied in the hesitation model 32. In someimplementations, the hesitation model 32 is a segmentation model that generates a hesitation tag in response to detecting a hesitation in an utterance 21. Among the embodiments are those in which the segmentation model is a transformer model. In some of these embodiments, the hesitation tag includes an expected pause length.
[0054] Once the syntax of the utterance 21 has been resolved and its content understood, the streaming-language processor 14 provides it to the automotive assistant 20, which then formulates an appropriate response and provides that response back to the streaming-language processor 14. The streaming-language processor 14 uses a text-to- speech unit 34 to provide suitable speech for communication via the loudspeaker 18.
[0055] The user 22 has no practical way to know whether the streaming-language processor 14, having dynamically extended the time-out period, is still waiting for input. In some embodiments, such as that shown in FIG. 5, a graphical user-interface 36 provides a visual cue 37, such as an animated image, to indicate that this is the case.
[0056] FIG. 5 shows details of a graphical-user interface in action. As shown in FIG. 5, an unresolved utterance 39 has been detected and reproduced in the graphical userinterface 36 with ellipsis 41 to indicate the utterance’s lack of resolution. Following the ellipsis 41 is a candidate resolution 43 for the unresolved utterance 39. This candidate resolution 43 is one that is inferred based on one or both user preferences and context. The candidate resolution is thus one that is calculated to promote or catalyze resolution of the utterance. The candidate resolution 43 in this example has been typeset to suggest that it is only a candidate resolution 43 and requires ratification by the user. In the illustrated embodiment, only one candidate resolution 43 is provided. However, embodiments include those in which additional candidate resolutions 43 are presented. These include embodiments in which the candidate resolutions are presented all at once and those in which they are presented one at a time, with the most likely one being presented first and the next most likely candidate resolution being presented in case the user rejects the first candidate resolution.
[0057] In some cases, even after the lapse of an extended time-out period, the user 22 may not have completed the command. In some embodiments, the hesitation model 32 signals the timer 24 to further extend the time-out period.
[0058] However, in other embodiments, the automotive assistant 20 includes an intentprediction model 38 that attempts to predict the intent of an unresolved utterance.
[0059] The intent-prediction model 38 is implemented as a large language model that receives prompts from the natural-language understander 30 and leverages context data 40, user preferences 42, and historical records 44 (such as user logs) to either infer the user’ s intent or to offer the user 22 suggestions on how to complete the command. The natural-language understander 30 then relays this information to the streaming controller 26, which provides it to the automotive assistant 20. The automotive assistant 20 then uses the output of the intent-prediction model 38 as a basis for prompting the user 22 to resolve the utterance 21. This act of prompting the user 22 to resolve the utterance 21 is an example of promoting resolution of the utterance, catalyzing resolution of the utterance, urging the user 22 to resolve the utterance, increasing the likelihood of resolution of the utterance, and / or reducing the expected time to resolution of the utterance.
[0060] As a concrete example, consider a user 22 who has uttered the syntactically- unresolved utterance 21: “Get the directions to...” Following lapse of the extended timeout period, the streaming-language processor 14 causes the intent-prediction model 38 to provide one or more guesses for how the user 22 might wish to move towards resolving the utterance 21. These guesses are useful for promoting resolution of the utterance, catalyzing resolution of the utterance, urging the user 22 to resolve the utterance, increasing the likelihood of resolution of the utterance, and / or reducing the expected time to resolution of the utterance.
[0061] In response, the intent-prediction model 38 consults the user preferences 42, the historical records 44, and the context data 40. Based on its consultation, the intentprediction model 38 recognizes that it is now Monday morning and that the historical records 44 indicate that for the past four Monday mornings, the user 22 has driven to rehearsals at Carnegie Hall. According to the context data 40, the vehicle 10 is traveling east towards the Lincoln Tunnel. Based on this information, the intent-prediction model 38 makes the not-unreasonable inference that it is the user’s intent to resolve the syntactically-unresolved utterance 21 by specifying “Carnegie Hall” as the missing object of the preposition “to” in the utterance “Get the directions to...” Upon receiving this information from the intent-prediction model 38, the automotive assistant 20 formulates a suitable prompt to assist in propelling the utterance 21 towards resolution, such as: “Did you want directions to Carnegie Hall?” This is an example of the automotive assistant 20 promoting resolution of the utterance, catalyzing resolution of the utterance, urging the user 22 to resolve the utterance, increasing the likelihood of resolution of the utterance, and / or reducing the expected time to resolution of the utterance. Since the automotiveassistant 20 has spoken in the middle of an utterance 21, the streaming-language processor 14 is in fact now communicating in duplex mode.
[0062] Referring to FIG. 6, the automatic speech-recognizer 28 and the natural-language understander 30 collaborate with each other while operating in “streaming mode.” When operating in streaming mode, the automatic speech-recognizer 28 receives streaming audio 46 that carries the utterance 21. It then divides the streaming audio 46 into partial speech-recognition results 48 that accumulate to form a final speech-recognition result 50.
[0063] A recognition stage 52 promotes the natural-language understander’s ability to operate on partial speech-recognition results 48. This recognition stage 52 includes a segmentation model 54 in communication with a hesitation recognizer 56 and an intent separator 58.
[0064] The segmentation model 54 processes partial speech-recognition results 48 one at a time. Upon receiving the partial speech-recognition results 48, the segmentation model 54 provides them to a hesitation recognizer 56. The hesitation recognizer 56 interacts with the hesitation model 32 in an attempt to identify hesitation during a static interval 25. The hesitation recognizer 56 communicates the existence of instances of hesitation to the natural-language understander 30.
[0065] The segmentation model 54 also provides partial speech-recognition results 48 to the intent separator 58. The intent separator’s role arises from the fact that as an utterance 21 proceeds, it can evolve from a “simple utterance,” which has only a single expression of intent, into a “compound utterance,” which carries more than one expression of intent.
[0066] An example of a compound utterance 21 is one such as, “Find the nearest gas station and take me there.” In the beginning, the utterance 21 appeal's to be a simple utterance with only one expression of intent, namely that of finding a gas station. As the utterance 21 continues over time, it evolves into a compound utterance 21 with a second expression of intent, namely that of navigating to the gas station. In the context of FIGS. 2-4, a second intent would correspond to a new confidence profile corresponding to that second intent. It is the job of the intent separator 58 to identify these distinct intents so that they can be separately acted upon by the automotive assistant 20.
[0067] In some cases, the hesitation model 32 notifies the streaming controller 26, which then signals the timer 24 to extend the time-out period. In other cases, the naturallanguage understander 30 prompts the user 22 to resolve the hesitation.
[0068] The natural-language understander 30 also obtains assistance from a surfing stage 60. The surfing stage 60 includes a slot identifier 62 and an intent detector 64, both of which interact with the intent-prediction model 38. The surfing stage 60 harnesses one or more large language models in an effort to infer the user’ s intent based on what the user has uttered thus far and stored context data 40, user preferences 42, and historical records 44, all of which are available to the intent-prediction model 38. The output of the surfing stage 60, which includes partial results from the natural-language understander 30 that have been tagged with hesitation labels is provided to the streaming controller 26.
[0069] The slot identifier 62 operates when a user 22 has difficulty in completing an utterance 21 that, as a result of its syntactic structure, is manifestly incomplete. This turns out to be the cause of many hesitations. For example, in the course of asking for directions to a location, the user 22 may stumble in attempting to articulate the location. This results in a brief interval during which the syntactic structure of the utterance 21 has an empty slot, namely the slot that should hold a destination. In the case of a fluent utterance 21, such as that shown in FIG. 2, this slot would be detected almost immediately by the slot identifier 62. However, in a disfluent utterance 21, such as that shown in FIG. 4, this slot remains empty for an extended period as the user 22 attempts to articulate what should be in it.
[0070] In such cases, the intent-prediction model 38 is prompted to propose one or more candidate fillers for the empty slot, thereby assisting the user 22 in resolving the utterance 21 more quickly. The result is then provided to the streaming controller 26 in real time.
[0071] It is apparent that the ability to process streaming audio 46 takes advantage of the time that the user 22 spends hesitating to carry out certain operations that are expected to be required in short order. This results in harnessing a normally negative feature, i.e., the time wasted as the user 22 hesitates, into something useful, namely extra time available for carrying out processing that anticipates what the user 22 will ultimately need.
[0072] For example, in the case of the utterance 21, “Find the nearest gas station and...” it is not unreasonable for a model to infer that the user 22 will also want to navigate there. As a result, navigation instructions can be requested in anticipation of the need to provide them, thus creating, for the user 22, the illusion of rapid processing.
[0073] In the case of a simple utterance 21, it is still possible that the expression of intent is insufficient to act upon. In such cases, the intent detector 64 likewise prompts theintent-prediction model 38 in an effort to resolve the ambiguity in the user’s expression of intent.
[0074] The activity carried out by the surfing stage 60 often requires additional input from the user 22 in real time. Accordingly, in the illustrated embodiment, proposals for resolution an ambiguous expression of intent as well as proposals for filling in slots are communicated to the streaming controller 26.
[0075] To assist it in determining what action to take, the streaming controller 26 includes a listening timer 66 that determines how much time has been spent listening to the user 22, an intent checker 68 that determines whether or not the intent expressed by the user or inferred based on the user’s utterance 21 is actionable, a label checker 70 to identify instances of hesitation by inspecting for portions of the utterance 21 that have been tagged with a hesitation tag, and a stream checker 72 that evaluates the effectiveness of the streaming audio 46 itself.
[0076] The intent-prediction model 38 includes a large-language model 74 that receives historical user information from a user-log repository 76. Upon completion of the utterance 21, the intent-prediction model provides its best guess 78 for the user’s intent to a natural-language generator 80 for composition of a suitable reply to be used for the automotive assistant 20 to provide to the text-to-speech unit 34. A result 82 is then played over the loudspeaker 18.
[0077] To avoid having to input an excessive number of of historical records 44 and user logs 76 as well as to reduce token size, it is desirable to use the output of the naturallanguage understander 30 to filter context data 40, user preferences 42, and historical records 44. In particular, it is preferable that only data that is both within an appropriate domain and that was acquired at an appropriate time be used to generate a prompt to the large-language model 74.
[0078] Requests to a large-language model 74 often generate costs both in time and for license fees. It is therefore desirable to avoid making unnecessary requests. Accordingly, when the natural-language understander 30 predicts a user’s intent, it is desirable to see whether that intent is within a pre-defined whitelist of intents. The user’s intent as ascertained by the natural-language understander 30 triggers a request to the large- language model 74 only if the intent is within that whitelist of intents.
[0079] For example, consider an incomplete utterance, "Navigate to ...." In such a case, the automotive assistant 20 recognizes the user’s intent, i.e., to navigate to some location.However, the utterance is incomplete because one cannot act on a navigation request without knowing the destination.
[0080] The intent-prediction model 38 selects those logs from the user-log repository 76 that relate to navigation and that occurred within some pre-defined interval from the current time, such as thirty minutes from the current time. This selection relies on the assumption that such logs will be more relevant to the user’s intent. These selected logs will then be used to generate a prompt that promotes dialog between the automotive assistant 20 and the user 22. Successful completion of a dialog turn in this dialog then causes an update to the user-log repository 76.
[0081] Implementations of the approaches described above may include hardware and / or software. Software can include non-transitory machine-readable media with instructions stored thereon. The instructions, when executed by using one or more physical processors cause the processors to perform the operations described herein. The instructions may be expressed in a high-level programming language, or as instructions for a physical or virtual processor. The processors may include local and / or remote processors. The hardware can include such processors and may also include application specific integrated circuits (ASICS) or configurable circuits such as field programmable gate arrays (FPGAs). The software and / or hardware can be integrated into an infotainment system that is in a vehicle and functions provided by the software and / or hardware may include functionality of a digital automotive assistant of the infotainment system.
[0082] Having described the invention and a preferred embodiment thereof, what is claimed as new and secured by letters patent is:
Claims
CLAIMS1. An apparatus for reducing occurrences of out-of-domain utterances resulting from static intervals within an utterance that is made by a user within a vehicle, said apparatus comprising an infotainment system in said vehicle and an automotive assistant that comprises a streaming-language processor and a hesitation model, wherein said automotive assistant is configured to be executed by said infotainment system for receiving an utterance from a user and taking action based on said utterance and on information from said streaming-language processor, wherein said streaming-language processor is configured to process said utterance in real time as said utterance is being received from said user via a microphone in said vehicle, said utterance comprising a growth interval, a static interval, and a transition from said growth interval to said static interval, wherein said hesitation model is configured to interact with said streaming-language processor to determine whether a static interval represents one of a first intent and a second intent, said first intent being an intent to end said utterance and said second intent being an intent to begin a new growth interval after said static interval, and wherein said streaming-language processor is configured to respond to detection of said second intent by promoting resolution of said utterance based on information from said hesitation model.
2. The apparatus of claim 1, wherein said growth interval is a first growth interval, wherein said apparatus further comprises a timer that starts a time-out period upon a transition from said growth interval to said static interval, and wherein said streaming-language processor is configured to promote resolution by extending said time-out period to grant said user more time to begin an additional growth interval after said static interval.
3. The apparatus of claim 1, wherein said growth interval is a first growth interval, wherein said static interval is a first static interval, wherein said utterance further comprises a second growth interval and a second static interval, wherein said apparatus further comprises a timer that starts a first time-out period and a second time-out period, wherein said timer starts said first time-out period when saidutterance transitions from said first growth interval into said first static interval, wherein said timer starts said second time-out period when said utterance transitions from said second growth interval into said second static interval, and wherein said first time-out period and said second time-out period differ length.
4. The apparatus of claim 1, wherein said growth interval is a first growth interval, wherein said apparatus further comprises a timer that starts a time-out period upon a transition from said first growth interval to said static interval, wherein said streaming-language processor is configured to promote resolution by extending said time-out period to grant said user more time to begin a second growth interval after said static interval in response to said hesitation model having inferred that said static interval represents a hesitation by said user and that said user intends to begin said second growth interval after said static interval,5. The apparatus of claim 1, wherein said apparatus further comprises a timer that starts a time-out period upon a transition from said growth interval to said static interval, wherein said hesitation model infers that said static interval is intended to end said utterance, and wherein said streaming-language processor ends said utterance prior to resolution thereof.
6. The apparatus of claim 1, further comprising an intent-prediction model, wherein said streaming-language processor is configured to promote resolution of said utterance by prompting said intent-prediction model to infer a candidate intent and providing said candidate intent to said user for verification thereof.
7. The apparatus of claim 1, further comprising an intent-prediction model, wherein said streaming-language processor is configured to promote resolution of said utterance by prompting said intent-prediction model to infer a candidate intent based on one or more of context data, user preferences, and historical records.
8. The apparatus of claim 1, wherein said streaming-language processor comprises: an automatic speech recognizer and a natural-language understander, both ofwhich operate in streaming audio mode, said automatic speech recognizer being configured to partition said utterance into partial speech-recognition results, each of which corresponds to a corresponding time interval within said utterance and to provide said partial speech-recognition results sequentially to said naturallanguage understander.
9. The apparatus of claim 1, wherein said streaming-language processor comprises a segmentation model that derives segments of natural language from partial speech-recognition results that are provided thereto in real time.
10. The apparatus of claim 1, wherein said streaming-language processor comprises an intent separator that identifies first and second distinct overt intents in said utterance based at least in part on partial speech-recognition results being provided in real time, each of which represents a portion of said utterance.
11. The apparatus of claim 1, wherein said streaming-language processor comprises a slot identifier, wherein said utterance has a syntax that includes plural slots, said slots including a filled slot and an empty slot, and wherein said slot filler is configured to identify said empty slot.
12. The apparatus of claim 1, further comprising a user interface that is configured to display an unresolved utterance, at least one candidate resolution for said unresolved utterance, said candidate resolution having been inferred based on at least one of user preferences and context information, and a visual indicator to draw attention to said unresolved utterance and to prompt resolution thereof.
Citation Information
Patent Citations
Automated speech recognition using a dynamically adjustable listening timeout
US20190348065A1
Dynamic contextual dialog session extension
US20210082397A1