Mitigating latency of verbal input guided item selection

By using only verbal input and visual output when the user interacts with the computing system, combining semantic parsers and display-dependent parsers to quickly match and select subsets of items, the problem of large interaction delay in the prior art is solved, and interaction efficiency and flexibility are improved.

CN119948560APending Publication Date: 2025-05-06GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071608.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-13
Filing Date
2023-10-11
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art When users interact with computing systems, the delay in guiding users to select subsets of items from supersets of candidates is large, resulting in extended resource usage and limited throughput.

Method used

By using only verbal input in user interaction with the computing system and providing visual output in response to verbal input, omitting audible synthetic verbal output, streaming transcripts are processed in parallel with semantic parsers and display-dependent parsers to quickly match and select subsets of items.

Benefits of technology

Reduces the delay in formulating subset selection, improves the efficiency and flexibility of interactions, and is suitable for users with limited flexibility and in environments that are not suitable for touch input, such as the drive-thru lane of fast service restaurants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948560A_ABST
    Figure CN119948560A_ABST
Patent Text Reader

Abstract

Latency to guide a user to select a subset of items from a superset of candidate items and cause further actions to be performed based on the selected subset of items during interaction between the user and a computing system is mitigated. In directing a user to select a subset of items, various implementations enable the user to provide only verbal input when selecting a subset of items, and provide visual output responsive to the verbal input and directing the user to select an item. In some of those various implementations, there is no (or only a very small amount) audible spoken synthesis spoken outputs rendered by the computing system when directing the user to select a subset of items.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Various computer-based methods have been proposed for guiding a user to select a subset of items from a superset of candidate items. For example, computer-based methods have been proposed for guiding a user to select a subset of specific vehicle features for a vehicle from a superset of candidate vehicle features for the vehicle. As another example, computer-based methods have been proposed for guiding a user to select a subset of tickets for an item at a location from a superset of available tickets for the item.

[0002] As a specific example, methods for at least partially automating food ordering at a quick service restaurant (QSR) have been proposed. For example, some QSRs implement an ordering kiosk with a touch screen that enables a user to provide touch-based input to browse a hierarchy of menu items and select a subset of those menu items to be incorporated into an order. However, the use of a touch screen and / or hierarchical navigation may be high latency, for example, due to the need for the user to thoroughly review the options presented on each screen before making a selection and / or inadvertently browse in the wrong branch of the hierarchy (e.g., and need to browse back in the hierarchy). This may result in extended resource usage of a computing device implementing the ordering kiosk and / or may result in a constrained throughput of a fixed set of ordering kiosks. Further, the use of a touch screen may be impractical or impossible for users with limited flexibility and / or in various situations (e.g., when a user stays in a car at a drive-thru restaurant).

[0003] In addition, for example, it has been proposed to utilize a turn-based audible dialogue in which a user provides spoken utterances and a computer system provides audible synthesized verbal responses to attempt to formulate an order. However, such techniques may be high latency due to the time required to render the audible synthesized verbal responses. For example, if a user's spoken utterance in a turn is ambiguous and partially matches five different menu items, an audible synthesized verbal response describing the five different items will be rendered to enable the user to disambiguate the ambiguous initial input through further spoken utterances. Rendering such audible verbal responses requires time and computing device resource utilization, and the user may need to wait for the audible verbal response to be rendered before providing further spoken utterances to disambiguate. Summary of the invention

[0004] The implementations described herein are intended to mitigate the latency of guiding a user to select a subset of items from a superset of candidate items and causing further actions to be performed based on the selected subset of items during an interaction between the user and a computing system. Further actions performed based on the selected subset of items may include, for example, transmitting the selected subset of items via an application programming interface (API) to cause the selected subset to be added to a list (e.g., added to an order) and / or causing other fulfillment actions to be performed based on the selected subset of items. Therefore, implementations propose various techniques for guiding a user to complete a technical task (e.g., transmitting the selected subset via an API or other backend) during an interaction between the user and a computer system and enabling the technical task to be completed with low latency.

[0005] When guiding a user to select a subset of items, various implementations enable the user to provide only verbal input when selecting a subset of items, and provide a visual output that responds to the verbal input and guides the user to select the item. In some of those various implementations, there is no (or only a very small amount) of audible synthetic verbal output rendered by the computing system when guiding the user to select a subset of items. Instead, the guidance of the user is achieved via a visual output, such as an image of an item (e.g., at least those items in a subset) and optionally a visual natural language description of the item, an annotation rendering descriptor of the item, and / or a non-natural language earmark (e.g., a short first positive "ding" to indicate a selection and / or a short second non-positive ding to indicate a cancellation of a selection). In some of those various implementations, a non-verbal and non-touch selection of items included in a subset of items can be utilized additionally or alternatively. For example, an image from a camera can be processed to detect a display area to which a user is pointing and / or to which a user's gaze is directed, and items in the area can be selected to be included in a subset in response to a user pointing and / or directing their gaze to the area.

[0006] By omitting any (or only including a very small amount) of audible synthetic oral output when guiding the user, the time delay of formulating the selection of the subset can be alleviated. For example, this may be due to the fact that the visual output corresponding to the project can be rendered faster relative to the rendering of the synthetic oral output, and can be rendered simultaneously on the display (while the synthetic oral output must be rendered sequentially over time). Additionally or alternatively, this may be due to the fact that humans understand the visual output faster than the synthetic oral output. Further, the time delay can be reduced by implementing the selection of the subset by verbal input (for example, only verbal input is provided by the user, and humans can provide verbal input faster than a series of touch inputs for complex layered navigation interfaces). Further, implementing the selection of the subset by verbal input enables the user with constrained flexibility to perform the selection and / or enables the selection to be performed in the case where the touch and / or other types of input to the corresponding computing device are not practical or impossible. Even further, additionally or alternatively implementing the implementation of the non-verbal and non-touch selection of the items in the subset enables the user with speech impairment to perform the selection and / or enables the selection to be performed in the case where the verbal input is impossible or impractical (for example, the case where there is a high level of noise in the environment).

[0007] The implementation disclosed herein may include or interact with a visual display, via which a visual output is rendered. The visual display may include, for example, a television, a monitor, or other visual display. The visual display may be controlled by a computing device, incorporated as part of the visual display or communicated with the visual display. For example, the computing device may include a browser or other applications that interact with a graphical user interface (GUI) system, and the GUI system may specify what the application causes to be rendered on the display. The implementation may further include a microphone that may at least selectively detect audio data or interact with it, and the audio data stream detected via the microphone may at least selectively be provided to a streaming automatic speech recognition (ASR) component. The microphone may be incorporated as part of the display, or may be separated from the display but close to the display. In various implementations, in response to detecting the possible or actual presence of verbal input, such as detecting voice activity (e.g., via a voice activity detection model), detecting a vehicle within a threshold proximity of a microphone and / or a corresponding display, and / or detecting a human within a threshold proximity of a microphone and / or a corresponding display, the audio data stream is provided to a streaming ASR component.

[0008] The GUI system communicates with the streaming ASR component and can receive the streaming transcription when the streaming transcription is generated by the streaming ASR component. The ASR component can generate the streaming transcription by processing the audio data stream detected via a microphone and capturing the user's spoken words.

[0009] The GUI system may include a semantic parser that processes the streaming transcription as it is received to generate one or more instances of structured representations that match the streaming transcription (each based on a portion of the streaming transcription received so far) and a confidence measure for each of the structured representations. The semantic parser may optionally interact with a large language model (LLM) to generate the structured representations and / or corresponding confidence measures. For example, the portion of the streaming transcription received so far may be processed using the LLM to generate a representation output indicating a semantic representation of the portion of the streaming transcription received so far. The representation output may then be processed by the semantic parser to determine which of the multiple candidate structured representations that each correspond to an item in the superset match the representation outputs, and to generate a corresponding confidence measure for each candidate structured representation. Each of the candidate structured representations may be, for example, in JavaScript Object Notation (JSON) format (or other structured format), and may indicate, for example, a semantic identifier of the corresponding item and properties of the item. Attributes of an item that may be indicated by a corresponding structured representation may include the quantity of the item (e.g., one of the item, two of the item, etc.), modification options for the item (e.g., for a cheeseburger item, modification options may include adding and / or removing: mustard, ketchup, pickles, and / or onions), and / or related options for the item (e.g., for a cheeseburger item, related options may include "make it a combo," "add fries," and / or "add a drink").

[0010] When the structured representation generated by the semantic parser has an associated confidence measure that satisfies a threshold value (e.g., an absolute threshold value and / or a threshold value relative to other structured representations of confidence measures), only the corresponding item can be selected. In response, the visual output of the corresponding item can be caused to be rendered in the GUI via a display. The visual output can include, for example, a pre-stored image of the corresponding item, a visual natural language descriptor of the corresponding item, the price of the corresponding item, and / or other visual outputs. The visual output of the corresponding item can be the only visual output rendered in the GUI via a display, and / or can be rendered to indicate that the corresponding item is ready to be included in the list by indicating (e.g., a border around some / all of the visual output of the corresponding item). If passive confirmation (e.g., the passage of a threshold amount of time) or active confirmation (e.g., saying "add it", "confirm", etc.) occurs during the rendering of the visual output of the corresponding item, the corresponding item can be added to the list (e.g., via interaction with the fulfillment API). Optionally, during rendering of the visual output for the corresponding item, the streaming transcription may continue to be monitored by a semantic parser and / or a display-dependent parser (as described herein) for further verbal input, e.g., which modifies the item according to a modification option, adds other items according to relevant options for the item, and / or cancels the inclusion of the corresponding item to the list.

[0011] When a structured representation generated by a semantic parser has an associated confidence metric that fails to meet a threshold (e.g., an absolute threshold and / or a threshold relative to confidence metrics of other structured representations), multiple structured representations in the structured representation may be selected, such as N with the highest confidence metrics and / or those structured representations with confidence metrics that meet a secondary threshold. In response, a corresponding visual output of each item in the item corresponding to the selected structured representation may be caused to be rendered in the GUI via a display. Optionally, the position of the visual output of the item within the GUI may be determined based on the confidence metric of the structured representation of the item (as determined by the semantic parser) and / or based on other metrics of the item. Such other metrics of the item may include a popularity metric of the item, such as a popularity metric indicating the frequency with which the item has been included in a list by a group of users over a certain period of time (e.g., over the past day, over the past week, etc.).

[0012] In various implementations, when rendering a corresponding visual output for each of the multiple items, at least one corresponding annotation descriptor may be rendered in the GUI along with the visual output for each of the multiple items. An annotation descriptor for an item may be a descriptor that does not semantically describe the item. For example, assume that corresponding visual outputs are provided for the following two items: a bacon cheeseburger and a cheeseburger. The annotation descriptor for the bacon cheeseburger may include a number (e.g., "1"), a letter (e.g., "A"), and / or a color (e.g., "yellow") rendered along with the bacon cheeseburger visual output (e.g., on top of, next to, or around the bacon cheeseburger image). The annotation descriptor for the cheeseburger is selected to be different from those annotation descriptors for the bacon cheeseburger, and may include a number (e.g., "2"), a letter (e.g., "B"), and / or a color (e.g., "red") rendered along with the cheeseburger visual output.

[0013] When rendering the visual output of multiple projects simultaneously in the GUI, the association between each project and their rendering descriptors can be defined. The rendering descriptor of the visual output of a project may include a position descriptor and / or an annotation descriptor. The position descriptor of the visual output of a project may describe the relative position of the visual output in the GUI, that is, the relative position relative to the visual output of other projects. For example, if the visual output of a project is presented above the visual output of all other projects, the position descriptor of the project may include "top", "first" and / or "upper". The annotation descriptor of the visual output of a project may describe the annotations rendered in the GUI together with the visual output of the project. For example, if the image of the project is rendered adjacent to "A", and the image is framed with "yellow", the annotation descriptor of the project may include "A" and "yellow".

[0014] In various implementations, the GUI system also includes a display-dependent parser that selectively processes streaming transcription in parallel with the semantic parser at least. For example, the display-dependent parser can process streaming transcription in parallel with the semantic parser at least when rendering the visual output of multiple items in the GUI via a display. The display-dependent parser can use the current rendering descriptor to determine whether the current portion of the streaming transcription matches a rendering descriptor in the rendering descriptor. If matched, the display-dependent parser can cause the selection of the item stored in association with the matching rendering descriptor. For example, assuming that the image of the item is currently being rendered above the images of all other items being displayed, the image is adjacent to "A" and the image is framed with "yellow". The display-dependent parser can cause the selection of the item in response to a portion of the streaming transcription that temporarily corresponds to such rendering (including any one of "A", "yellow", "top" and / or "first"). When image-based inputs (such as pointing inputs and / or gaze-based inputs) are also enabled, the display-dependent parser may additionally or alternatively cause selection of items in response to those inputs associated with position descriptors of the items. For example, an item with a position descriptor of "top" may be selected in response to detecting that a user is pointing and / or directing their gaze toward a "top" area of ​​the display.

[0015] As noted above, the semantic parser can process streaming transcription in parallel with the parser that relies on display, and can select corresponding items in response to generating a corresponding structured representation with a threshold confidence measure. Therefore, in the case where the user provides a verbal word that semantically describes the corresponding item, the semantic parser can cause the selection of the corresponding item, and such verbal words will not cause the parser that relies on display to cause the selection of the corresponding item (because such verbal words will not match the rendering descriptor). On the contrary, in the case where the user provides a verbal word that references the position of the item in the GUI and / or an annotation rendered together with the visual output of the item in the GUI, the parser that relies on display can cause the selection of the corresponding item, and such verbal words will not cause the semantic parser to cause the selection of the corresponding item (because such verbal words will not semantically describe the actual corresponding item). In other words, the semantic parser and the parser that relies on display can complement each other when processing streaming transcription in parallel, because the semantic parser can cause the selection of the item in response to the words that semantically describe the item, and the parser that relies on display can cause the selection of the item in response to the words that match the rendering descriptor associated with the current display of the visual output of the item in the GUI.

[0016] In these and other ways, through the parallel operation of the semantic parser and the display-dependent parser, the selection of items can be enabled for a more robust range of terms included in the spoken input. This provides the corresponding user with the flexibility to provide a spoken output that truly semantically represents the desired item, or to provide a spoken output that only semantically represents how the item is currently displayed in the GUI. Therefore, this can enable the user to say what resonates most with the user, thereby enabling faster speaking and faster selection of the corresponding item. Further, speaking a term corresponding to a rendering descriptor of an item (e.g., "A" or "top") can often be faster than speaking a term that truly semantically represents the item (e.g., "the one with bacon").

[0017] The above description is provided as an overview of only some of the implementations disclosed herein. Those implementations and other implementations are described in more detail herein.

[0018] It should be understood that the techniques disclosed herein may be implemented locally on a client device, remotely by a server connected to the client device via one or more networks (e.g., in the "cloud" by a cluster of remote servers), and / or both. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A block diagram illustrating an example environment in which implementations disclosed herein may be implemented.

[0020] Figure 2 Example methods according to various implementations are illustrated.

[0021] Figure 3 A display illustrating an example initial state of rendering a GUI for guiding a user to select a subset of items from a superset of candidate items.

[0022] Figure 3A1 Example rendering in Figure 3 A display of an example next state of the GUI following an initial state, and illustrating example spoken words of a user that may cause the example next state to be rendered.

[0023] Figure 3A2 Example rendering in Figure 3A1 A display of an example further next state of the GUI following the next state and illustrating example spoken words of a user that may cause the example further next state to be rendered.

[0024] Figure 3B1 Example rendering in Figure 3A display of an example alternative next state of the GUI following an initial state and illustrating example spoken utterances of a user that may cause the example alternative next state to be rendered.

[0025] Figure 3B2 Example rendering in Figure 3B1 The alternative next state is followed by display of an example further alternative next state of the GUI, and illustrates example spoken utterances of a user that may cause the example further alternative next state to be rendered.

[0026] Figure 4 Depicted are example architectures for computing devices in accordance with various implementations. DETAILED DESCRIPTION

[0027] Initially go to Figure 1 , depicts a block diagram of an example environment 100 that illustrates various aspects of the present disclosure and in which implementations disclosed herein may be implemented.

[0028] The example environment 100 includes a client device 110, a streaming ASR engine 130, an interactive GUI system 140, and a fulfillment system 150. The components of the example environment can be communicatively coupled to each other via one or more networks, such as one or more wired or wireless local area networks ("LANs", including Wi-Fi LANs, mesh networks, Bluetooth, near field communications, etc.) or wide area networks ("WANs", including the Internet). In some implementations, the components of the streaming ASR engine 130 and / or the interactive GUI system are implemented in a "cloud" via a cluster of high-performance servers.

[0029] The client device 110 is illustrated as including a microphone 114, a speaker 116, an application 118, and a display 112. Figure 1 , microphone 114, speaker 116, application 118, and display 112 are all depicted within a rectangle representing client device 110. However, it should be understood that in various implementations, the components of client device 110 will not be housed as part of a single structure, but may be positionally distributed throughout the environment. It should also be understood that in various implementations, the components of client device 110 are not shown in the figure for simplicity. Figure 1 The additional components shown in can be housed as part of a single structure and / or distributed positionally in the environment. A non-limiting example of such additional components is a presence sensor, which is described in more detail herein and can be used to distinguish subsequent users.

[0030] As an example, the client device 110 may include a structure (e.g., a thin client device) that includes, for example, a processor, a memory, a network interface component, and / or other hardware components (not shown in FIG. Figure 1 ), and app 118 can be executed by the processor utilizing the memory. App 118 can be utilized to generate the GUI described herein, and display 112 can render the GUI generated by app 118. However, display 112 can be located remotely from a structure containing a processor, memory, etc., but in communication with a structure containing a processor, memory, etc. For example, communication between the structure and display 112 can be wireless communication or wired communication (e.g., via a High Definition Multimedia Interface (HDMI) connection). Similarly, microphone 114 and / or speaker 116 can optionally be located remotely from a structure containing a processor, memory, etc., but in communication with a structure containing a processor, memory, etc. As a specific example, display 112, microphone 114, and speaker can be located in an external environment adjacent to a drive-thru lane of a QSR, and the structure can be located inside the QSR.

[0031] The audio data stream detected via microphone 114 may be at least selectively provided to streaming ASR engine 130. For example, the audio data stream may be provided to streaming ASR engine 130 in response to detecting the possible presence or actual presence of spoken input. For example, client device 110 may detect possible or actual voice input based on detecting voice activity using a voice activity detection model, based on detecting a vehicle within a threshold proximity of microphone 114 and / or display 112, and / or based on detecting a human within a threshold proximity of microphone 114 and / or display 112. Detecting a vehicle and / or detecting a human may be based on information from a presence sensor coupled to the client device (not in the Figure 1 ), such as a passive infrared (PIR) sensor, a weight sensor (e.g., detecting the weight of an indicated vehicle), an underground magnetic ring sensor, a laser beam sensor (e.g., detecting the presence or passage of a vehicle), and / or other sensors.

[0032] The ASR engine 130 can use a streaming ASR model to process the audio data stream that captures the spoken utterance and is generated by the microphone 114 to generate a streaming transcription of the spoken utterance. For example, the ASR model can include, for example, a recurrent neural network (RNN) model, a transformer model, and / or any other type of machine learning model. As the streaming transcription of the spoken utterance is generated (e.g., on a part-by-part basis), the streaming transcription is provided to the interactive GUI system 140 by the ASR engine 130.

[0033] The interactive GUI system 140 is illustrated as including a semantic parser 142 , a display-dependent parser 144 , a resolution engine 146 , a GUI engine 148 , and a state management engine 149 .

[0034] The semantic parser 142 processes the streaming transcription as it is received to generate one or more instances of a structured representation that matches the streaming transcription (each based on a portion of the streaming transcription received so far) and a confidence measure for each of the structured representations in the structured representations. In some implementations, the semantic parser 142 may interact with a large language model (LLM) 143 to generate the structured representations and / or corresponding confidence measures. In some implementations, the semantic parser 142 may additionally or alternatively interact with an alternative machine learning model (e.g., a neural network model) and / or may utilize a text matching heuristic to generate the structured representations and / or corresponding confidence measures. For example, an alternative machine learning model and / or text matching heuristic may optionally be used when there is a relatively small superset of items and / or when the semantically meaningful descriptors of the items are relatively constrained.

[0035] As an example, before implementing the interactive GUI system 140 for an entity (e.g., an entity associated with a QSR), the entity may provide (e.g., via an API) for each item in the entity's superset of items: (a) a structured representation of the item (e.g., a structured representation that conforms to the grammar of a fulfillment system 150, which may be managed by the entity), (b) an image of the item, and (c) a natural language descriptor of the item (e.g., a menu name for the item). The information provided by the entity may be stored in an item database 155. Further, the semantic parser 142 may generate a corresponding semantic representation for each item in the items and store it in association with the (a) structured representation (e.g., in the item database 155). For example, the semantic parser 142 may process the (c) natural language descriptor using the LLM 143 to generate an LLM representation output (e.g., an embedding of a vector of values ​​of an output layer of the LLM 143), and use the generated LLM representation output as the semantic representation of the item.

[0036] Thereafter, the semantic parser 142 may receive a portion of the streaming transcription and process the portion (and optionally, the preceding portion) using the LLM to generate an LLM representation output of the portion. The semantic parser 142 may then compare the LLM representation output with the semantic representation of each item in the project. For example, the semantic parser 142 may generate a cosine distance metric, each of which is a corresponding cosine distance between the LLM representation and the semantic representation of the corresponding item. For an item, the cosine distance metric may indicate whether the item matches the streaming transcription received so far, and further, may indicate a confidence metric for the match (i.e., a closer distance metric indicates a greater confidence than a farther distance metric). The semantic parser 142 may select some of the items with the closest distance metric (or not select the item with the closest distance metric), and output the stored structured representation of those items, optionally together with the corresponding confidence metric of each item (e.g., a confidence metric based on the distance metric). For example, if a given distance metric satisfies a threshold (e.g., absolute and / or relative to other distance metrics), the semantic parser 142 can select the corresponding item and output only its structured representation. As another example, if all distance metrics in the distance metric fail to meet a threshold (e.g., an absolute threshold and / or a threshold relative to other distance metrics), the semantic parser 142 can select multiple items in the item, such as N with the closest distance metric and / or those items with a distance metric that satisfies a secondary threshold. The semantic parser 142 can then output a structured representation of each item in the selected item, and optionally output a corresponding confidence metric for each item (e.g., which conforms to or is based on the corresponding distance metric).

[0037] Display-dependent parser 144 may selectively process the streaming transcript at least in parallel with semantic parser 142. For example, display-dependent parser 144 may process the streaming transcript in parallel with the semantic parser at least while GUI engine 148 is causing visual output of a plurality of items to be rendered in the GUI via display 112. Display-dependent parser 142 may determine whether a current portion of the streaming transcript matches one of the rendering descriptors using a current rendering descriptor provided by GUI engine 148. If so, display-dependent parser 142 may cause selection of an item stored in association with the matching rendering descriptor.

[0038] The resolution engine 146 can work in conjunction with the semantic parser 142, the display-dependent parser 144, and / or the GUI engine 146. For the item selected by the parser 142 or 144, and when the visual output corresponding to the item is being rendered, the resolution engine 146 can determine whether the item should be added to the list, whether any modification is to be made to the item (before being added to the list), and / or whether any related options of the item are also to be selected and added to the list. When determining whether to add the item to the list for the selected item being rendered, the resolution engine 146 can determine to add the item to the list in response to passive confirmation. For example, passive confirmation can include the passage of a threshold duration after rendering the selected item, optionally rendering the visual descriptor of the selected item together with the indication of the upcoming selection (e.g., the border around the rendering item) at the same time. The resolution engine 146 can additionally or alternatively determine to add the item to the list in response to active confirmation, such as the user saying "add it", "confirm", "done (completed)" etc. When the user speaks the confirmation portion of the utterance, the corresponding portion of the transcription can be provided to the semantic parser 142, and the semantic parser 142 can determine the confirmation intent, provide the confirmation intent to the resolution engine 146, and the resolution engine 146 can determine the active confirmation in response to receiving the confirmation intent.

[0039] When determining whether to make any modifications to the selected item being rendered and / or whether any related options should also be selected and added to the list, the resolution engine 146 can reference the structured representation of the item, which can define possible modifications and / or related options. Further, the semantic parser 142 and / or the display-dependent parser 144 can be utilized to determine whether a part of the transcription refers to the modification and / or selection of the related options, and the corresponding indication can be provided to the resolution engine 146 for the resolution engine 146 to make such determinations. For example, assuming that the related option "make it a combo" is rendered in the GUI, together with the annotation "A" of the related option. If the user says "A", the display-dependent parser 144 can determine that this is related to the related options, provide the indication that this is related to the related options to the resolution engine 146, and the resolution engine 146 can add the "combo (set meal)" item to the list (e.g., "fries and a drink (fries and drinks)"). If the user says "combo it," semantic parser 142 may determine that this matches the intent of "make it a combo," provide an indication that this matches the intent of "make it a combo" to resolution engine 146, which may determine that this is an active intent because it is for related options, and resolution engine 146 may add the "combo" item to the list.

[0040] The GUI engine 148 may interact with the application 118 and cause the GUI rendered by the application 118 to be dynamically updated throughout interaction with the user and dependent upon output from the semantic parser 142, the display-dependent parser 144, and / or the resolution engine 148. In some implementations, the application 118 may be a browser and / or the GUI engine 148 may render the GUI via a web page (such as via a script in an HTML document).

[0041] The state management engine 149 can maintain a list of ongoing interactions and can receive items to be added to the list of ongoing interactions and / or items to be removed from the list of ongoing interactions from the resolution engine 146. Further, once the ongoing interaction is completed, the state management engine 149 can interact with the fulfillment system 150 to cause one or more actions to be performed based on the final state of the list. For example, the state management engine 149 can transmit a structured representation and / or other identifier of the items in the list to the fulfillment system 150 via an API. In response, the fulfillment system 150 can, for example, cause the items to be queued for preparation, queued for delivery, and / or can perform other actions. In some implementations, the state management engine 149 can be omitted from the interactive GUI system 140 and can be combined with the fulfillment system 150 instead. In some of those implementations, the resolution engine 146 can interact with the fulfillment system 150 (e.g., via an API) to add and / or remove items from the list of ongoing interactions.

[0042] although Figure 1 1 is described with respect to a single client device 110, but it should be understood that this is for purposes of example and is not intended to be limiting. For example, a given entity may utilize multiple client devices. For example, a given QSR location may include two or more client devices that each interact with the interactive GUI system. As another example, multiple distinct entities may each have a corresponding client device and may each interact with the interactive GUI system 140 (or another instance thereof). However, different entities will be associated with different items, and those items may be stored in the item database 155 and provided by the entities.

[0043] Now go to Figure 2 , which illustrates a flow chart of an example method 200 of some implementations disclosed herein. For convenience, the operations of the method 200 are described with reference to a system that performs the operations. The system includes a computing device (e.g., an interactive GUI system 140, Figure 4 The method 200 may be performed by one or more processors, memories, and / or other components of a computing device 410, one or more servers, and / or other computing devices. In addition, although the operations of method 200 are shown in a particular order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.

[0044] At block 202, the system receives a streaming ASR transcription. The streaming ASR transcription may be generated by an ASR engine based on processing an audio data stream from a microphone associated with a display device using a streaming ASR model. The audio data stream may capture one or more spoken utterances of a user.

[0045] At block 204 , the system processes the current transcription (ie, the transcription received so far of the streaming ASR transcription) using the semantic parser.

[0046] At block 206, the system determines whether there are any matching items based on the semantic parsing of block 204. For example, the system can determine whether the semantic parsing of block 204 indicates any items in a superset of items corresponding to the streaming ASR transcription. For example, if the semantic parsing generates structured representations for respective corresponding items with corresponding confidence scores that satisfy a threshold, the system can determine that there are matching items. If at block 206 the system determines that there are no matching items, the system re-advances to block 202 and receives additional portions of the streaming ASR transcription. If at block 206 the system determines that there are matching items, the system advances to block 208.

[0047] At block 208, the system determines whether there are multiple matching items, or conversely, only a single matching item, based on the semantic parsing of block 204. For example, the system can make this determination based on whether the semantic parsing has generated multiple structured representations and / or based on confidence scores of the multiple structured representations. If the system determines at block 208 that there is only a single item, the system proceeds to block 212. If the system determines at block 208 that there are multiple matching items, the system proceeds to block 210.

[0048] At block 210, the system displays a corresponding image for each of the plurality of items, and displays the corresponding images simultaneously. For example, the system may cause the application GUI to render the corresponding images for the plurality of items. The system may optionally display additional information for each of the plurality of items, such as a visual natural language descriptor for the item, a visual price for the item, and / or other additional information.

[0049] Box 210 optionally includes an optional sub-box 210A, in which the system displays and / or defines a corresponding rendering descriptor for each item. For example, the system can display an annotation descriptor "A" next to a first image of a first item, an annotation descriptor "B" next to a second image of a second item, and so on, and define corresponding associations between those annotation descriptors and corresponding items. In addition, for example, the system can define corresponding position descriptors for each item in the project. Each position descriptor can describe the relative position of the corresponding image (and optional additional information) of the item in the display.

[0050] After block 210, the system re-advances to block 202 and receives additional portions of the streaming ASR transcription.

[0051] During at least some iterations of blocks 204, 206, 208, 210, and / or 212, the system can perform block 222 in parallel. For example, the system can perform block 222 at least while a rendering descriptor is defined for the current display (e.g., via iterations of subblock 210A). At block 222, the system processes the current transcript using a display-dependent parser. The display-dependent parser can utilize the current rendering descriptor to determine whether a current portion of the streaming transcript matches one of the rendering descriptors.

[0052] At block 224, the system determines whether the display-dependent parser indicates that the current portion of the streaming transcription matches one of the rendering descriptors. If not, the system re-advances to block 202. If yes, the system advances to block 212.

[0053] At box 212, the system displays an image of the single item and, optionally, displays additional information for the single item. The system may optionally render an indication of impending selection along with the image of the single item (e.g., highlighting around the image). When box 212 is encountered from a "yes" determination at box 224, the system displays the single item based on a determination at box 224 indicating that the current portion of the streaming transcription matches one of the rendering descriptors associated with the single item. When box 212 is encountered from a "no" determination at box 208, the system displays the single item based on a determination at box 208 indicating that there is only a single matching item. Box 212 optionally includes sub-box 212A, in which the system displays related options for one or more dependent items (e.g., other related items to be added to the list along with the selected item).

[0054] At box 214, the system determines whether there is a selection of a single item of box 212, and optionally, whether there is also a selection of a related option of a dependent item in the related options of the dependent item displayed at subbox 212A. For example, the system can determine the passive selection of a single item in response to the passage of a threshold amount of time without receiving any further verbal input from the user. As another example, the system can determine the active selection of a single item in response to a further affirmative verbal input from the user (e.g., as determined in another iteration of box 204, not shown for simplicity). As another example, the system can determine that there is also a selection of a related option of a dependent item in the related options of the dependent item in response to a further affirmative verbal input from the user, and the further affirmative verbal input semantically references the related option or references the rendering descriptor of the related option (e.g., as determined in another iteration of box 204 and / or another iteration of box 222, not shown for simplicity).

[0055] If the decision at box 214 is no (e.g., in response to receiving further negative verbal input), the system re-advances to box 202, optionally also rendering an audible and / or visual negative prompt, such as a negative ear icon and / or an "X" or other "cancel" symbol. If the decision at box 214 is yes, the system advances to box 216 and performs one or more corresponding actions, such as interacting with a state management engine and / or a fulfillment system to add the item to the list. Box 216 optionally includes box 216A, in which the system renders an audible and / or visual affirmative prompt to indicate that the item is added to the list, such as a positive ear icon and / or a "check mark" or other "positive" symbol.

[0056] At block 218, the system determines whether the current interaction with the user is complete. This can be based on, for example, performing another iteration of block 204 and / or another iteration of block 222 (not shown for simplicity) to determine whether the user has provided further verbal input, and if so, whether the further verbal input indicates that the interaction is complete or, conversely, that the interaction should continue. If the determination at block 218 is "yes," the system re-advances to block 202. If the decision at block 220 is "no," the system advances to block 220 and the iteration of method 200 ends.

[0057] Figure 3 An example initial state 180 illustrating rendering of a GUI for guiding a user to select a subset of items from a superset of candidate items Figure 1 The display 112. Figure 3 and Figure 3A1 , Figure 3A2 , Figure 3B1 and Figure 3B2 In the example of , the superset of candidate items is the menu items of the QSR. However, it should be understood that the technology disclosed herein can be utilized in addition or alternatively with other types of items.

[0058] The initial state 180 may be an initial state that is initially displayed, for example, after a previous order is completed and / or when a new vehicle is detected within proximity of the display 112. The initial state 180 includes a first visual output of a first item, a second visual output of a second item, and a third visual output of a third item. These three items are a subset of the menu items and may be selected for display on the initial screen based on various criteria. For example, one or more of the items may be selected based on being popular, being part of a current promotion, or the selection of one or more of the items may be random.

[0059] The first visual output of the first project includes an image 182A, a natural language descriptor 184A, and an annotation 186A. The annotation 186A includes the number "1" and includes a "yellow" coloring around the "1", where the "yellow" is indicated by a vertical hatch. The image 182A and the descriptor 184A can be provided by the QSR and retrieved from a corresponding database. The annotation 186A can be automatically generated by the GUI engine 148 based on a template or otherwise, for example. Further, the GUI engine 148 can associate a rendering descriptor with the first project, such as an annotation descriptor "1", "yellow" and a position descriptor such as "top" and "first".

[0060] The second visual output of the second project includes image 182B, natural language descriptor 184B and annotation 186B. Annotation 186B includes number "2", and includes "blue (blue)" coloring around "2", wherein "blue" is indicated by horizontal hatching. Image 182B and descriptor 184B can be provided by QSR and retrieved from corresponding database. Annotation 186B can be automatically generated by GUI engine 148 based on template or other means, for example. Further, GUI engine 148 can associate rendering descriptor with the second project, such as annotation descriptor "2", "blue" and position descriptors such as "middle (middle)" and "second (second)".

[0061] The third visual output of the third project includes image 182C, natural language descriptor 184C and annotation 186C. Annotation 186C includes number "3", and includes "green (green)" coloring around "3", wherein "green" is indicated by diagonal hatching. Image 182C and descriptor 184C can be provided by QSR and retrieved from corresponding database. Annotation 186C can be automatically generated by GUI engine 148 based on template or other means, for example. Further, GUI engine 148 can associate rendering descriptor with the third project, such as annotation descriptor "3", "green" and position descriptors such as "bottom (bottom)" and "third (third)".

[0062] Initially go to Figure 3A1 and Figure 3A2 , illustrating the Figure 3 One possible progression starting with the initial graphical interface.

[0063] Figure 3A1 Example rendering in Figure 3The display 112 shows an example next state 180A1 of the GUI after the initial state 180 of the GUI, and illustrates example spoken utterances 190A1A, 190A1B, 190A1C, and 190A1D of the user, any of which when displayed Figure 3 When state 180 is provided, it can cause the example next state 180A1 to be rendered. Figure 3A1 In the example above, the second item ("mystery sub") has been selected as the specific item, and in Figure 3A1 The next state 180A1 of exemplifies a visual output 182B1 of an image of "mystery sub" having a box around the image to indicate that "mystery sub" is about to be added to the list.

[0064] When displaying Figure 3 When spoken utterances 190A1A ("two"), 190A1B ("blue"), or 190A1C ("middle one") are provided at state 180 of the display-dependent parser 144, the selection of "mystery sub" (and the transition to the next state 180A1) may be based on output from the display-dependent parser 144 when processing the corresponding transcription. For example, because the rendering descriptors "2", "blue", and "middle" are associated with the second item when displaying state 180, any of spoken utterances 190A1A, 190A1B, and 190A1C ("middle one") will cause the display-dependent parser 144 to determine, based on the corresponding transcription, that "mystery sub" is selected.

[0065] Notably, semantic parser 142 does not resolve any of spoken utterances 190A1A, 190A1B, and 190A1C as “mystery sub” because they do not semantically describe an actual “mystery sub.” However, when spoken utterance 190A1D (“mystery sub”) is provided, selection of “mystery sub” (and transition to next graphical interface 180A1) may be based on output from semantic parser 144 when processing the corresponding transcription.

[0066] Figure 3A1Also illustrated is a visual output of a first related option for "mystery sub" that also adds drinks and fries to the list. The visual output of the first related option includes an image 182D, a natural language descriptor 184D, and an annotation 186D. The annotation 186D includes the number "1" and includes "yellow" coloring around the "1". The image 182D and the descriptor 184D can be provided by the QSR and retrieved from a corresponding database. The annotation 186D can be automatically generated by the GUI engine 148 based on a template or otherwise, for example. Further, the GUI engine 148 can associate a rendering descriptor with the first related option, such as the annotation descriptors "1" and "yellow". Still further, the display-dependent parser 144 can utilize such rendering descriptors when the graphical interface 180A1 is rendered to determine whether the verbal input refers to any of those rendering descriptors, and if so, the addition of the item of the first related option to the list can be caused.

[0067] The visual output of the second related option includes an image 182E, a natural language descriptor 184E, and an annotation 186E. An annotation 186E includes the number "2" and includes "blue" coloring around "1". The image 182E and the descriptor 184E can be provided by the QSR and retrieved from the corresponding database. The annotation 186E can be automatically generated by the GUI engine 148 based on a template or otherwise, for example. Further, the GUI engine 148 can associate rendering descriptors with the second related option, such as the annotation descriptors "2" and "blue". Further still, the display-dependent parser 144 can utilize such rendering descriptors when the graphical interface 180A1 is rendered to determine whether the verbal input refers to any one of those rendering descriptors.

[0068] Figure 3A2 Example rendering in Figure 3A1 The display of the example further next state 180A2 after the next state 180A1 and the example spoken utterances 190A2A, 190A2B and 190A2C of the user are illustrated, any of which when displayed Figure 3A1 When the next state 180A1 is provided, it can cause the example further next state 180A2 to be rendered. The further next state 180A2 includes Figure 3A1 The same visual output 182B1 and also includes a natural language indication 182B1A that "mystery sub" has been added to the order. Further, an audible affirmative ding 189A2 may be rendered to indicate that "mystery sub" was added to the order.

[0069] Spoken utterance 190A2A actually represents the lack of any further spoken input from the user. This lack of any further spoken input for a threshold duration can be a passive confirmation that "mystery sub" was added to the order (e.g., by resolution engine 146). Spoken utterances 190A2B ("done") and 190A2C ("sub only") are examples of positive confirmations that "mystery sub" was added to the order. They can be viewed as positive confirmations based on the output from semantic parser 142 when processing the corresponding transcription.

[0070] Now go to Figure 3B1 and Figure 3B2 , illustrating the Figure 3 The alternative possible progression from the initial state 180 is as will be appreciated, for example with reference to Figure 3B1 and Figure 3B2 Descriptions of alternative possible progressions may be based on alternative spoken utterances provided by the user.

[0071] Figure 3B1 Example rendering in Figure 3 The display of the GUI in the next state 180B2 after the initial state and the example spoken utterances 190B1A and 190B1B of the user are shown, either of which when displayed Figure 3 When state 180 is provided, it can cause the example next state 180B1 to be rendered.

[0072] Figure 3B1 The fourth visual output of the fourth item, the fifth visual output of the fifth item, and the sixth visual output of the sixth item are included. Figure 3B1 In alternative next state 180B1 , based on being selected, based on output from semantic parser 142 when processing a transcription from spoken utterance 190B1A or spoken utterance 190B1B, the fourth item, the fifth item, and the sixth item are selected for display.

[0073] The fourth visual output of the fourth project includes an image 182F, a natural language descriptor 184F, and an annotation 186F. The annotation 186F includes the number "1" and includes "yellow" coloring around the "1". The image 182F and the descriptor 184F can be provided by the QSR and retrieved from a corresponding database. The annotation 186F can be automatically generated by the GUI engine 148 based on a template or otherwise, for example. Further, the GUI engine 148 can associate a rendering descriptor with the fourth project, such as the annotation descriptor "1", "yellow", and a position descriptor such as "top" and "first".

[0074] The fifth visual output of the fifth project includes an image 182G, a natural language descriptor 184G, and an annotation 186G. The annotation 186G includes the number "2" and includes "blue" coloring around the "2". The image 182G and the descriptor 184G can be provided by the QSR and retrieved from a corresponding database. The annotation 186G can be automatically generated by the GUI engine 148 based on a template or otherwise, for example. Further, the GUI engine 148 can associate a rendering descriptor with the fifth project, such as the annotation descriptors "2", "blue", and position descriptors such as "middle" and "second".

[0075] The sixth visual output of the sixth item includes an image 182H, a natural language descriptor 184H, and an annotation 186H. The annotation 186H includes the number "3" and includes "green" around the "3". The image 182H and the descriptor 184H can be provided by the QSR and retrieved from a corresponding database. The annotation 186H can be automatically generated by the GUI engine 148 based on a template or otherwise, for example. Further, the GUI engine 148 can associate a rendering descriptor with the sixth item, such as the annotation descriptors "3", "green", and position descriptors such as "bottom" and "third".

[0076] Figure 3B2 Example rendering in Figure 3B1 The example GUI of the alternative next state 180B1 is followed by a further next alternative state 180B2 of the display 112, and illustrates example spoken utterances 190B2A, 190B2B, 190B2C of the user, any of which when displayed Figure 3B1 The next state 180B1 when provided may cause a further alternative next state 180B1 of the example to be rendered.

[0077] exist Figure 3B2 , the fourth item ("eggs and bacon platter") has been selected as a specific item, and Figure 3B2 A further next graphical interface 180B2 of exemplifies an output 182F1 of an image of “eggs and bacon platter” having a box around the image visually indicating that “eggs and bacon platter” has been added to the order. Figure 3B2Also included is a natural language indication 182F1A that "eggs and bacon platter" has been added to the order. Further, an audible affirmative ding 189B2 (via speaker 116) may be rendered to indicate that "eggs and bacon platter" has been added to the order. Based on the fact that none of the relevant options are defined for "eggs and bacon platter", Figure 3B2 There are no options for "eggs and bacon platter" that are instantiated.

[0078] When displaying Figure 3B1 When the spoken utterance 190B2A ("one") or 190B2B ("yellow") is provided when display-dependent parser 144 displays the alternative next state 180B1, the selection of "eggs and bacon platter" (and the transition to the alternative further next state 180B2) may be based on the output from display-dependent parser 144 when processing the corresponding transcription. For example, because the rendering descriptors "1" and "yellow" are associated with the fourth item when displaying the alternative next state 180B1, either of the spoken utterances 190B2 and 190B2B will result in display-dependent parser 144 determining the selection of "eggs and bacon platter" based on the corresponding transcription.

[0079] Notably, semantic parser 142 does not resolve either of spoken utterances 190B2A and 190B2B to "eggs and bacon platter" because they do not semantically describe the actual "eggs and baconplatter." However, when spoken utterance 190B2C ("plate") is provided, the selection of "eggs and baconplatter" (and the transition to further alternative next state 180B2) may be based on the output from semantic parser 144 when processing the corresponding transcription. Although Figure 3B2 , but in some implementations, the transition to the further alternative next state 180B2 may additionally or alternatively be based on, for example, detecting that the user is pointing and / or directing their gaze toward an area of ​​the alternative next state 180B1, and the semantic parser 142 determining that the area is associated with a location descriptor of “eggs and bacon platter.”

[0080] As noted above, the alternative further next state 180B2 includes a visual output 182F1 of an image of "eggs & bacon platter" with a box around the image to indicate that "eggs and bacon platter" has been added to the order, and is illustrated without the Figure 3B1 Any visual output corresponding to the alternative next state 180B1 of "baconcheeseburger" or "crab cake & bacon w / bun". However, another alternative further next state may instead match Figure 3B1 180B1, but includes a visual indication to indicate that "eggs & bacon platter" is selected. In other words, the further alternative next state may not only include visual output corresponding to "eggs & bacon platter", but may also include visual output corresponding to "bacon cheeseburger" and "crab cake & bacon w. bun" (even if these are not selected). For example, the visual indication indicating selection may include a box or circle around the image 182F, the natural language descriptor 184F, and / or the annotation 186F. Further, for example, the visual indication indicating selection may additionally or alternatively include an "X" or strikethrough rendered atop the visual output of "bacon cheeseburger" and "crab cake & bacon w / bun". An audible earmark indicating selection may also be rendered in this further alternative next state.

[0081] Now go to Figure 4 , depicts a block diagram of an example computing device 410 that can optionally be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more of a client device, a cloud-based automated assistant component, and / or other components can include one or more components of the example computing device 410.

[0082] The computing device 410 typically includes at least one processor 414 that communicates with a number of peripheral devices via a bus subsystem 412. These peripheral devices may include a storage subsystem 424 (including, for example, a memory subsystem 425 and a file storage subsystem 426), a user interface output device 420, a user interface input device 422, and a network interface subsystem 416. The input and output devices allow for user interaction with the computing device 410. The network interface subsystem 416 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.

[0083] User interface input device 422 may include a keyboard, a pointing device (such as a mouse, trackball, touchpad, or graphics tablet), a scanner, a touch screen integrated into a display, an audio input device (such as a voice recognition system, a microphone), a camera for gesture detection, a camera for detecting pointed items, or a camera for detecting visual focus, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways to input information into computing device 410 or onto a communication network.

[0084] User interface output device 420 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanisms for creating a visible image. The display subsystem may also provide a non-visual display such as via an audio output device. Generally speaking, the use of the term "output device" is intended to include devices and methods for outputting information from computing device 410 to all possible types of users or another machine or computing device.

[0085] The storage subsystem 424 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 424 may include logic to perform selected aspects of the methods disclosed herein and to implement the various components depicted in the figures.

[0086] These software modules are typically executed by the processor 414 alone or in combination with other processors. The memory 425 used in the storage subsystem 424 may include multiple memories, including a main random access memory (RAM) 430 for storing instructions and data during program execution and a read-only memory (ROM) 432 in which fixed instructions are stored. The file storage subsystem 426 can provide persistent storage for program and data files, and can include a hard drive, a solid-state disk drive or other memory chip, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular implementation can be stored in the storage subsystem 424 by the file storage subsystem 426, or in other machines accessible to the processor 414.

[0087] Bus subsystem 412 provides a mechanism for the various components and subsystems of computing device 410 to communicate with each other as intended. Although bus subsystem 412 is schematically shown as a single bus, alternative implementations of bus subsystem 412 may use multiple busses.

[0088] Computing device 410 may be of different types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the Figure 4 The description of the computing device 410 depicted in FIG. 4 is intended only as a specific example for illustrating some implementations. Many other configurations of the computing device 410 are possible, and these configurations are similar to those of FIG. Figure 4 The computing devices depicted in FIG. 1 may have more or fewer components than those depicted in FIG.

[0089] In situations where the systems described herein collect or otherwise monitor personal information about a user or can make use of the personal information and / or monitored information, the user may be given the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, the user's preferences, or the user's current geographic location), or to control whether and / or how content that may be more relevant to the user is received from a content server. In addition, certain data may be processed in one or more ways before it is stored or used so that personally identifiable information is removed. For example, the user's identity may be processed so that the user's personally identifiable information cannot be determined, or, in the case of obtaining geographic location information, the user's geographic location may be generalized (such as to a city, zip code, or state level) so that the user's specific geographic location cannot be determined. Thus, the user may control how information is collected about the user and / or how it is used.

[0090] In some implementations, a method implemented by a processor is provided, and the method includes, while spoken utterance is being provided by a user and captured in an audio data stream via one or more microphones: receiving a portion of a streaming transcription of the spoken utterance, the streaming transcription being generated using streaming automatic speech recognition; and processing the portion using a semantic parser to determine, from a defined superset of items, a subset including a plurality of items in the superset. The method further includes, while the spoken utterance is being provided and captured and in response to determining the subset: selecting a corresponding image for each of the items in the subset; causing the corresponding image to be rendered simultaneously on a display visible to the user and without simultaneously rendering any image of any other item in the superset that is not included in the items in the subset; and for each item in the subset, defining a corresponding association of the item with one or more corresponding rendering descriptors of the item. The method further includes, while the spoken utterance is being provided and captured and during concurrent rendering of the corresponding image: receiving an additional portion of a streaming transcription of the spoken utterance, the additional portion based on a portion of the spoken utterance provided during concurrent rendering of the corresponding image; processing the additional portion using a semantic parser and processing the additional portion using a display-dependent parser; determining a specific item in the items of the subset based on processing the additional portion of the transcription; and in response to determining the specific item based on processing the additional portion of the transcription, performing further actions specific to the specific item. Processing the additional portion using the display-dependent parser may include utilizing a corresponding rendering descriptor in response to the additional portion being based on a portion of the spoken utterance provided during concurrent rendering of the corresponding image.

[0091] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0092] In some implementations, no synthesized speech is provided as output during provision of the spoken utterance.

[0093] In some implementations, determining a specific item in the items in the subset based on processing the additional portion of the transcript includes: determining the specific item based on the specific item being explicitly indicated by one of the following: (a) using a semantic parser to process the additional portion and (b) using a display-dependent parser to process the additional portion. In some versions of those implementations, processing the additional portion using the display-dependent parser includes: determining whether the additional portion matches any one of the corresponding rendering descriptors. In some versions of those versions, determining the specific item in the items in the subset based on processing the additional portion of the transcript includes: in response to determining that the additional transcript portion matches a given rendering descriptor in the corresponding rendering descriptors and the corresponding association for the specific item in the items in the subset is with the given rendering descriptor, selecting the specific item. The given rendering descriptor can be, for example, a position descriptor or an annotation descriptor, such as a position descriptor that describes a relative position of a corresponding image of the specific item on a display, or an annotation descriptor that describes an annotation rendered on a display together with the corresponding image of the specific item. For example, the given rendering descriptor can be an annotation descriptor, and the annotation descriptor can be a number, a letter, a code, or a color.

[0094] In some implementations, processing the additional portion using a semantic parser includes: generating, based on the portion and the additional portion, using the semantic parser, a structured representation corresponding to a specific item and a confidence measure of the structured representation. In some of those implementations, determining the specific item in the subset based on processing the additional portion of the transcription includes: determining the specific item based on the specific item corresponding to the structured representation and based on the confidence measure of the structured representation satisfying a threshold.

[0095] In some implementations, performing the further action includes causing a corresponding image of the particular item to be rendered on a display without simultaneously rendering any other corresponding image of the corresponding images of any other item in the items in the subset. In some versions of those implementations, the method further includes determining acceptance of the particular item after performing the further action; and in response to determining acceptance, interacting with a fulfillment application programming interface (API) to add the particular item to a list maintained via the fulfillment API. In some of those versions, determining acceptance is based on no further verbal input being received within a threshold time period after causing the corresponding image of the particular item to be rendered on the display without simultaneously rendering any other corresponding image of the corresponding images of any other item in the items in the subset. In some additional or alternative versions of those versions, the method further includes, in response to determining acceptance: causing an audible affirmative earcon to be rendered via one or more speakers on or near the display.

[0096] In some implementations, performing the further action includes interacting with a fulfillment application programming interface (API) to add the particular item to a list maintained via the fulfillment API.

[0097] In some implementations, the one or more processors are one or more remote servers in network communication with the display.

[0098] In some implementations, a method implemented by a processor is provided, and the method includes, while spoken utterances are being provided by a user and captured in an audio data stream via one or more microphones: processing a portion of the audio data stream using a streaming automatic speech recognition (ASR) model to generate a transcribed portion of a streaming transcription of the spoken utterance; processing the transcribed portion to determine, from a defined superset of items, a subset that includes a plurality of items in the superset; and in response to determining the subset: selecting a corresponding image for each item in the items in the subset; causing the corresponding images to be rendered simultaneously on a display visible to the user and without simultaneously rendering any images of any other items in the superset that are not included in the items in the subset; and for each item in the items in the subset, defining a corresponding association of one or more corresponding rendering descriptors for the item. The method further includes, during concurrent rendering of the corresponding image: processing an additional portion of the audio data stream using the streaming ASR model to generate an additional transcribed portion of a streaming transcription of the spoken utterance, the additional portion of the audio data stream including audio data captured during concurrent rendering of the corresponding image; determining that the additional transcribed portion matches a given rendering descriptor in the rendering descriptors; in response to determining that the additional transcribed portion matches the given rendering descriptor, selecting a specific item in the items in the subset; and in response to selecting the specific item, performing further actions specific to the specific item.

[0099] These and other implementations of the techniques disclosed herein may include one or more of the following features.

[0100] In some implementations, processing the transcribed portion to determine the subset includes determining at least a threshold degree of match between the first transcribed portion and each of the items in the subset.

[0101] In some implementations, the one or more corresponding rendering descriptors each describe a corresponding location of a corresponding one of the rendering items.

[0102] In some implementations, the one or more corresponding rendering descriptors each describe a corresponding annotation applied to a rendering of a corresponding image for a corresponding one of the projects.

[0103] In some implementations, processing a transcribed portion to determine a subset from a defined superset of items includes, based on the portion, using a semantic parser to generate: a plurality of structured representations, each corresponding to a corresponding one of the items in the subset, and a corresponding confidence metric for each of the structured representations; and determining the subset based on the corresponding confidence metrics of the structured representations of the items in the subset satisfying a threshold.

[0104] In some implementations, causing the corresponding images to be simultaneously rendered on the display includes causing the corresponding images to be rendered in an arrangement determined based on the corresponding confidence metric for each of the structured representations. In some of those implementations, causing the corresponding images to be simultaneously rendered on the display includes causing the corresponding images to be rendered in an arrangement determined based on: the corresponding confidence metric for each of the structured representations; and the corresponding popularity metrics for the items in the subset.

[0105] In some implementations, no synthesized speech is provided as output during provision of the spoken utterance.

[0106] In some implementations, performing the further action includes causing a corresponding image of the particular item to be rendered on a display without simultaneously rendering any other corresponding image of the corresponding images of any other item in the items in the subset. In some versions of those implementations, the method further includes determining acceptance of the particular item after performing the further action; and in response to determining acceptance, interacting with a fulfillment application programming interface (API) to add the particular item to a list maintained via the fulfillment API. In some of those versions, determining acceptance is based on no further verbal input being received within a threshold time period after causing the corresponding image of the particular item to be rendered on the display without simultaneously rendering any other corresponding image of the corresponding images of any other item in the items in the subset. In some additional or alternative versions of those versions, the method further includes, in response to determining acceptance: causing an audible affirmative earcon to be rendered via one or more speakers on or near the display.

[0107] In some implementations, performing the further action includes interacting with a fulfillment application programming interface (API) to add the particular item to a list maintained via the fulfillment API.

[0108] In some implementations, the one or more processors are one or more remote servers in network communication with the display.

[0109] In addition, some implementations include one or more processors (e.g., central processing units (CPUs), graphics processing units (GPUs), and / or tensor processing units (TPUs)) of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in an associated memory, and wherein the instructions are configured to cause performance of any of the above methods. Some implementations also include one or more non-transitory computer-readable storage media storing computer instructions that can be executed by one or more processors to perform any of the above methods. Some implementations also include a computer program product that includes instructions that can be executed by one or more processors to perform any of the above methods.

Claims

1. A method implemented by one or more processors, the method comprising: While spoken words are being provided by a user and captured in an audio data stream via one or more microphones: receiving a portion of a streaming transcription of the spoken utterance, the streaming transcription generated using streaming automatic speech recognition; processing the portion using a semantic parser to determine, from a defined superset of items, a subset comprising a plurality of items from the items in the superset; In response to determining the subset: selecting a corresponding image for each of the items in the subset; causing the corresponding images to be rendered simultaneously on a display visible to the user and without concurrently rendering any image of any other item in the superset that is not included in the items in the subset; as well as defining, for each item in the items in the subset, a corresponding association of the item with one or more corresponding rendering descriptors for the item; During concurrent rendering of said corresponding images: receiving an additional portion of the streamed transcription of the spoken utterance, the additional portion being based on a portion of the spoken utterance provided during concurrent rendering of the corresponding image; processing the additional portion using the semantic parser and processing the additional portion using a display-dependent parser, wherein processing the additional portion using the display-dependent parser comprises utilizing the corresponding rendering descriptor in response to the additional portion being based on the portion of the spoken utterance provided during concurrent rendering of the corresponding image; determining, based on processing the additional portion of the transcript, a particular item among the items in the subset; as well as In response to determining the specific item based on processing the additional portion of the transcript, further actions specific to the specific item are performed.

2. The method of claim 1, wherein no synthesized speech is provided as output during providing the spoken utterance.

3. The method of any preceding claim, wherein determining the particular item in the subset based on processing the additional portion of the transcript comprises: The specific item is determined based on being explicitly instructed by one of: (a) processing the additional portion using the semantic parser and (b) processing the additional portion using the display-dependent parser.

4. The method of claim 3, wherein processing the additional portion using the display-dependent parser comprises: A determination is made as to whether the additional portion matches any one of the corresponding rendering descriptors.

5. The method of claim 4, wherein determining the particular one of the items in the subset based on processing the additional portion of the transcript comprises: In response to determining that the additional transcription portion matches a given one of the corresponding rendering descriptors and the corresponding association for a particular one of the items in the subset is with the given rendering descriptor, selecting the particular item.

6. The method of claim 5, wherein the given rendering descriptor is a position descriptor that describes a relative position of the corresponding image of the particular item on the display.

7. The method of claim 5, wherein the given rendering descriptor is an annotation descriptor that describes an annotation that is rendered on the display along with the corresponding image of the particular item.

8. The method of claim 7, wherein the annotation descriptor is a number, a letter, a code or a color.

9. The method of any preceding claim, wherein processing the additional portion using the semantic parser comprises: A structured representation corresponding to the specific item and a confidence measure of the structured representation are generated using the semantic parser based on the portion and the additional portion.

10. The method of claim 9, wherein determining the particular item in the subset based on processing the additional portion of the transcript comprises: The specific item is determined based on the specific item corresponding to the structured representation and based on the confidence measure of the structured representation satisfying a threshold.

11. The method of any preceding claim, wherein performing the further action comprises: The corresponding image of the particular item is caused to be rendered on the display without concurrently rendering any other corresponding image of the corresponding images of any other item of the items in a subset.

12. The method of claim 11, further comprising: determining acceptance of the particular item after performing the further action; as well as Responsive to determining the acceptance, interacting with a fulfillment application programming interface (API) to add the particular item to a list maintained via the fulfillment API.

13. A method as claimed in claim 12, wherein determining the acceptance is based on no further verbal input being received within a threshold time period after causing the corresponding image of the particular item to be rendered on the display without simultaneously rendering any other corresponding image of the corresponding images of any other item in the items in the subset.

14. The method of claim 12 or claim 13, further comprising, in response to determining the acceptance: An audible affirmative earcon is caused to be rendered via one or more speakers on or near the display.

15. The method of any preceding claim, wherein performing the further action comprises interacting with a fulfillment application programming interface (API) to add the particular item to a list maintained via the fulfillment API.

16. A method as claimed in any preceding claim, wherein the one or more processors are one or more remote servers in network communication with the display.

17. A method implemented by one or more processors, the method comprising: While spoken words are being provided by a user and captured in an audio data stream via one or more microphones: processing a portion of the audio data stream using a streaming automatic speech recognition (ASR) model to generate a transcribed portion of a streaming transcription of the spoken utterance; processing the transcribed portion to determine, from a defined superset of items, a subset comprising a plurality of items from among the items in the superset; In response to determining the subset: selecting a corresponding image for each of the items in the subset; causing the corresponding images to be rendered simultaneously on a display visible to the user and without concurrently rendering any image of any other item in the superset that is not included in the items in the subset; as well as defining, for each item in the items in the subset, a corresponding association of one or more corresponding rendering descriptors for the item; During concurrent rendering of said corresponding images: processing an additional portion of the audio data stream using the streaming ASR model to generate an additional transcribed portion of the streaming transcription of the spoken utterance, the additional portion of the audio data stream comprising audio data captured during simultaneous rendering of the corresponding image; determining that the additional transcription portion matches a given one of the rendering descriptors; responsive to determining that the additional transcription portion matches the given rendering descriptor, selecting a particular item of the items in the subset; as well as In response to selecting the particular item, further actions specific to the particular item are performed.

18. The method of claim 17, wherein processing the transcribed portion to determine the subset comprises determining at least a threshold degree of match between the first transcribed portion and each of the items in the subset.

19. The method of claim 17 or claim 18, wherein each of the one or more corresponding rendering descriptors describes rendering a corresponding position of a corresponding one of the items.

20. The method of any one of claims 17 to 19, wherein each of the one or more corresponding rendering descriptors describes a corresponding annotation applied to the rendering of the corresponding image of a corresponding one of the items.

21. The method of claim 17, wherein processing the transcribed portion to determine the subset from the defined superset of items comprises: Based on the part using a semantic parser to generate: a plurality of structured representations, each of the plurality of structured representations corresponding to a respective one of the items in the subset, and a corresponding confidence measure for each of the structured representations; as well as The subset is determined based on the corresponding confidence measures of the structured representations of the items in the subset satisfying a threshold.

22. The method of claim 21 , wherein causing the corresponding images to be rendered simultaneously on the display comprises: The corresponding images are caused to be rendered in an arrangement determined based on the corresponding confidence measure for each of the structured representations.

23. The method of claim 21 , wherein causing the corresponding images to be rendered simultaneously on the display comprises: causing the corresponding images to be rendered in an arrangement, the arrangement being determined based on: the corresponding confidence measure for each of the structured representations; and Corresponding popularity metrics for the items in the subset.

24. A method as claimed in any one of claims 17 to 23, wherein no synthesized speech is provided as output during provision of the spoken utterance.

25. The method of any one of claims 17 to 24, wherein performing the further action comprises: The corresponding image of the particular item is caused to be rendered on the display without concurrently rendering any other corresponding image of the corresponding images of any other item of the items in a subset.

26. The method of claim 25, further comprising: determining acceptance of the particular item after performing the further action; as well as Responsive to determining the acceptance, interacting with a fulfillment application programming interface (API) to add the particular item to a list maintained via the fulfillment API.

27. A method as claimed in claim 26, wherein determining the acceptance is based on no further verbal input being received within a threshold time period after causing the corresponding image of the particular item to be rendered on the display without simultaneously rendering any other corresponding image of the corresponding images of any other item in the items in the subset.

28. The method of claim 27, further comprising, in response to determining the acceptance: An audible affirmative earcon is caused to be rendered via one or more speakers on or near the display.

29. The method of any one of claims 17 to 28, wherein performing the further action comprises interacting with a fulfillment application programming interface (API) to add the particular item to a list maintained via the fulfillment API.

30. The method of any one of claims 17 to 29, wherein the one or more processors are one or more remote servers in network communication with the display.

31. A system comprising: a memory storing instructions; One or more processors operable to execute said instructions to cause performance of a method as claimed in any preceding claim.

32. One or more computer-readable storage media storing computer instructions executable by one or more processors to perform the method of any one of claims 1 to 30.