Voice control input of content in a graphical user interface
Through the adaptive selection of verbatim explanation or substitute content into the input field, the long interaction time and waste of equipment resources caused by user verbatim narration are solved, and more efficient interaction and energy-saving effects are achieved.
Patent Information
- Application Number
- CN201980097603.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-15
- Filing Date
- 2019-12-13
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2039-12-13
AI Technical Summary
In the prior art, when a user interacts with an automation assistant, the input content needs to be described word by word, resulting in long interaction time, high equipment power consumption and waste of resources.
By adaptively selecting whether to incorporate verbatim or alternative content of spoken discourse into the input field, the amount of speech recognition processing is reduced, the alternative content is determined using situation data and natural language understanding, and a public interface is provided for interaction.
It shortens the interaction time between users and automation assistants, reduces device power and resource consumption, and improves interaction efficiency.
Smart Images

Figure CN114144789B_ABST
Abstract
Description
Background Art
[0001] Humans can engage in human-machine conversations using an interactive software application herein referred to as an "automation assistant" (also known as a "digital agent", "chatbot", "interactive personal assistant", "intelligent personal assistant", "assistant application", "conversational agent", etc.). For example, humans (who can be referred to as "users" when they interact with the automation assistant) can provide commands and / or requests to the automation assistant using spoken natural language input (i.e., utterances) and / or by providing text (e.g., typed) natural language input. In some cases, the spoken natural language input can be converted into text and then processed. The automation assistant responds to the request by providing a responsive user interface output, which can include audio and / or visual user interface outputs.
[0002] As mentioned above, an automation assistant can convert audio data corresponding to a user's spoken utterance into corresponding text (or other semantic representation). For example, audio data can be generated based on the detection of a user's spoken utterance via one or more microphones of a client device, which includes an assistant interface that enables the user to interact with the automation assistant. The automation assistant can include a speech recognition engine that attempts to identify various features of the spoken utterance captured in the audio data, such as the sounds (e.g., phonemes) produced by the spoken utterance, the order of the produced sounds, the rhythm of the speech, the intonation, etc. In addition, the speech recognition engine can identify the text words or phrases represented by these features. The text can then be further processed by the automation assistant when determining the response content of the spoken utterance (e.g., using a natural language understanding (NLU) engine and / or a dialogue state engine). The speech recognition engine can be implemented by the client device and / or by one or more automation assistant components that are remote from the client device but communicate with the client device over a network
[0003] Separate from the automation assistant, certain keyboard applications can enable users to provide typed keyboard input by interacting with virtual keys presented by the keyboard application. Some keyboard applications also enable users to provide text input by dictating. For example, a user can select a "microphone" graphical element on the keyboard and then provide a spoken utterance. The keyboard application can then convert any audio data characterizing the spoken utterance into corresponding text and utilize the corresponding text as input to the application. Thus, these keyboard applications enable users to dictate instead of type and use strict verbatim dictation as text for the application. Summary of the Invention
[0004] Some implementations disclosed herein relate to processing audio data that captures a user's spoken utterance to generate recognized text of the spoken utterance and determining whether to provide any of the following for incorporation into an input field: (1) the recognized text itself, or (2) alternative content determined based on the recognized text. The alternative content is not an alternative speech-to-text recognition of the spoken utterance, but rather alternative content determined based on the recognized text. When determining whether to incorporate the recognized text or the alternative content into the input field, one or more characteristics of the recognized text and / or the alternative content may be considered.
[0005] In some implementations, when determining whether to incorporate the recognized text of the alternative content into the input field, one or more characteristics of the recognized content and / or the alternative content may be compared with parameters related to the context in which the user provided the spoken utterance. As an example, in one context, a user may interact with a web page to provide contact information for receiving a service provided by the web page. The user may select a portion of the web page corresponding to an input field identified as a "street address" field. Thereafter, the user may select a keyboard element to invoke an automated assistant (or the automated assistant may be invoked in any other manner, such as automatically in response to the keyboard application being brought to the foreground) and provide a spoken utterance, such as "my home". In response, the automated assistant may process the spoken utterance and cause the keyboard application to incorporate alternative content into the input field in place of the recognized text "my home", where the alternative content is the user's complete home address (e.g., "1111 West Muhammed Ali Blvd"). This may be based on, for example, the automated assistant determining that the input field is a "street address" input field and that the alternative content itself is a street address. In other words, the automated assistant may compare context-related parameters (e.g., the input field is a "street address" field) with characteristics of the alternative content (e.g., the alternative content is a street address) when determining to incorporate the alternative content in place of the recognized text. The automated assistant may utilize various techniques to determine that the input field is an address, such techniques as, for example, considering the XML or HTML of the web page (e.g., the XML or HTML tags for the input field), considering text and / or graphical elements near the input field, and / or other techniques.
[0006] However, when a user is communicating with another person via a messaging application, the user may receive a message from the other person such as, "Where are we meeting tonight?" In response, the user may invoke the automated assistant and provide a spoken utterance such as, "My home." In response, the automated assistant may process the spoken utterance and cause the keyboard application to incorporate a verbatim interpretation of the spoken utterance (i.e., the text "My home") rather than the user's full home address. The decision to incorporate the verbatim interpretation may be based on contextual data and / or historical interaction data characterizing previous instances of communication between the user and the other person. Additionally or alternatively, this decision may be based on whether the user has provided their full home address in a similar context and / or when the user is messaging the other person. Based on this information, the verbatim interpretation may be biased toward the full street address when determining whether to incorporate verbatim content or alternative content.
[0007] Providing an automated assistant that adaptively selects content that differs from a verbatim interpretation of a spoken utterance may allow for shortened interaction times between a user and the automated assistant. This may reduce the amount of automatic speech recognition (ASR) processing that must be performed, and may also reduce the amount of time a device that receives and / or processes spoken utterances remains on, thereby saving power and other computing resources that may be expended by the device. For example, the amount of time it takes a user to narrate their home address (e.g., "1111 West Muhammed Ali Blvd") may be much longer than the amount of time it takes a user to narrate the phrase "my home." Similarly, recognizing a complete address may require more ASR processing than recognizing a shorter phrase. By allowing compressed spoken utterances to replace otherwise lengthy spoken utterances, the overall duration of a user / device interaction may be reduced, thereby saving power and / or other resources that would otherwise be consumed by the extended duration of the user / device interaction. For example, a user / device interaction may end more quickly, enabling the device to transition to a reduced power state more quickly upon cessation of the interaction.
[0008] In addition, the implementations disclosed herein provide a common interface through which a user can provide spoken utterances, and the common interface automatically (e.g., without any further user input) distinguishes utterances that should result in providing recognized text for an input field from utterances that should result in providing alternative content for the input field. This can lead to improved human-computer interaction, where no further user interface input is required to alternate between an "oral dictation mode" and an "alternative content" mode. In other words, a user can utilize the same interface to provide a spoken utterance for which corresponding recognized text is provided to be inserted into a corresponding input field, and to provide a spoken utterance based on which alternative content is determined and provided to be inserted into the corresponding input field. Whether to provide "oral dictation" of a spoken utterance or "alternative content" based on a spoken utterance can be based on one or more considerations described herein and can be automatically determined without any explicit user input indicating which should be provided.
[0009] In addition, there is no need to launch a separate application and / or a separate interface on a computing device to identify alternative content through an extended user interaction with the separate application and / or interface. For example, in response to a message of "When does your flight leave / arrive tomorrow", when the keyboard application is active for a message reply input field, a user can provide a spoken utterance of "Insert details of my flight tomorrow". In response, alternative content including the user's flight details (e.g., departure airport and departure time, arrival airport and arrival time) can be determined and provided to the keyboard application for insertion into the reply input field by the keyboard application, instead of the recognized text "Insert details of my flight tomorrow". The alternative content is determined and provided for insertion without the user having to open a separate application, use the separate application to search for flight information, copy the flight information, and then return to the keyboard application to insert the flight information. Eliminating the need for the user to interact with a separate application and / or interface can reduce the duration of human-computer interaction and / or prevent the separate application from being executed, thus saving various computer resources.
[0010] In various implementations, the alternative content can be automated assistant content generated based on further processing of the recognized text using an automated assistant. For example, the natural language understanding (NLU) engine of the automated assistant can be used to process the recognized text to determine the automated assistant intent and / or the value of the intent, and the alternative content can be determined based on the response to the intent and / or the value.
[0011] In some implementations, the automated assistant can be an application separate from the keyboard application, and the automated assistant can interface with the keyboard application via an application programming interface (API) and / or any other software interface. The keyboard application can provide a keyboard interface. Optionally, the keyboard interface allows the user to invoke the automated assistant by, for example, tapping on a keyboard element presented at the keyboard interface. Additionally or alternatively, the automated assistant can be automatically invoked in response to the keyboard interface being exposed and optionally other conditions being met (e.g., voice activity being detected), and / or the automated assistant can be invoked in response to other automated assistant invocation user inputs (such as the detection of a wake phrase (e.g., "OK Assistant", "Assistant", etc.), certain touch gestures, and / or certain touchless gestures). When the automated assistant is invoked and the keyboard is exposed in the GUI of another application (e.g., a third-party application), the user can provide spoken utterances. The automated assistant can process the spoken utterances in response to the automated assistant being invoked, and the automated assistant can determine whether to provide an oral dictation of the spoken utterances or whether to provide alternative content based on the spoken utterances. The automated assistant can then provide a command to the keyboard application that causes the oral dictation or alternative content to be inserted into the corresponding input field by the keyboard. In other words, the automated assistant can determine whether to provide an oral dictation or alternative content and then communicate only the corresponding content to the keyboard application (e.g., via the keyboard application API or the operating system API). For example, when the keyboard interface is being presented and the automated assistant is invoked, providing a spoken utterance such as "my address" can cause the automated assistant to provide a command to the keyboard application based on the context in which the spoken utterance is provided. The command can, for example, cause the keyboard application to provide text that is a verbatim interpretation of the spoken utterance (e.g., "my address") or a different interpretation that causes the keyboard to output other content (e.g., "1111 West Muhammed Ali Blvd"). Allowing the automated assistant to interface with the keyboard application in this way provides a more lightweight keyboard application that does not need to instantiate all of the automated assistant functionality in memory and / or incorporate the automated assistant functionality into the keyboard application itself (thereby reducing the storage space required by the keyboard application). Instead, the keyboard application can rely on API calls from the automated assistant to effectively expose the automated assistant functionality. Additionally, allowing the automated assistant to determine whether to provide an oral dictation or alternative content based on the spoken utterances and communicate the corresponding data to the keyboard application enables this enhanced functionality to be used with any of a variety of different keyboard applications. As long as the keyboard application and / or the underlying operating system support supplying content to the keyboard application to be inserted by the keyboard, the automated assistant can interface with any of a variety of keyboard applications.
[0012] Some additional and / or alternative implementations described herein relate to an automated assistant that can infer what a user may be instructing the automated assistant to do without the user explicitly reciting verbatim what the user intends. In other words, the user may provide an oral utterance that refers to what the user is trying to have the automated assistant identify, and in response, the automated assistant may cause certain operations to be performed on words in the oral utterance that refer to the content. As an example, a user may interact with an application such as a web browser to register for a particular service provided by an entity such as another person whose website the user is accessing. When the user selects a particular input field presented at the interface of the application, a keyboard interface and / or the automated assistant may be initialized. The user may then choose to provide an oral utterance to have certain content incorporated into the particular input field that they selected.
[0013] In some implementations, the user may provide an oral utterance that is directed to the automated assistant and refers to what the user is trying to incorporate into the input field—but does not recite the content verbatim. As an example, when the input field is an address field that is intended to have the user provide a property address into the input field, the user may intend for the input field to be their location—their home address (e.g., 2812 First St., Lagrange, KY). To incorporate the home address into the input field, the user may provide an oral utterance such as “Assistant, where am I”. In response to receiving the oral utterance, the audio data representing the oral utterance may be processed to determine whether one or more portions of the verbatim content of the oral utterance (e.g., “where am I...”) are incorporated into the input field, or whether other content that the user may be referring to is incorporated. To make this determination, the context data representing the context in which the user provided the oral utterance may be processed.
[0014] In some implementations, the context data may represent metadata stored in association with the interface that the user is accessing. For example, the metadata may characterize the input field as an “address” field that requires a house number (e.g., “2812”). Thus, in response to receiving the oral utterance “Assistant, where am I”, the audio data may be processed to determine whether the user has explicitly recited any numbers. When it is determined that the oral utterance has no numbers, the audio data and / or the context data may be processed to determine whether other content is incorporated into the input field (e.g., without incorporating the verbatim recited content that is at least a portion of the oral utterance). In some implementations, the audio data may be processed to identify a synonymous interpretation of the oral utterance so that other content may be supplemented and considered for incorporation into the input field. Additionally or alternatively, the audio data may be processed to determine whether the user has historically referred to the expected content in similar oral utterances and / or via any other input to the computing device and / or application. For example, the audio data may be processed in view of other actions that can be performed by the automated assistant.
[0015] As an example, the automated assistant may include functionality for presenting navigation data when a user provides a spoken utterance such as "Assistant, when will the bus arrive at my location?". In response to the spoken utterance, the automated assistant may identify the address of the user's current location in order to identify a public transportation route that includes the user's current location. Similarly, the automated assistant may process the natural language content "my location" of the spoken utterance to determine that the spoken utterance may refer to an address at least based on previous instances when the user has used the phrase "my location", so that the automated assistant uses a command incorporating the user's current address as a parameter. Thus, in response to the spoken utterance, other content identified via the automated assistant may include the address of the user's current location (e.g., 2812 First St., Lagrange, KY). The other content identified may be further processed based on context data indicating that the input field is intended to have a certain amount of numbers (e.g., a house number). Thus, since the other content has numbers and satisfies the input field, and is at least more satisfactory than one or more parts of a verbatim interpretation of the spoken utterance, the other content may be incorporated into the input field (e.g., "Address input: 2812 First St., Lagrange, KY").
[0016] In some implementations, a user may provide a spoken utterance that includes a portion for input into a selected input field and another portion that is provided as a query to the automated assistant. For example, the user may interact with a third-party application and select an input field to facilitate entering text into the input field. In response to selecting the input field, a keyboard interface may be presented at the graphical user interface (GUI) of the third-party application. The keyboard interface may include GUI elements that, when selected by the user, cause the automated assistant to be invoked. For example, when the user attempts to respond to a recently received text message (e.g., "Jack: When are you going to the movies?"), the user may then select the text response field so that the keyboard interface is revealed at the display panel of the computing device. The user may then select a GUI element (e.g., a microphone graphic) presented with the keyboard interface so that the automated assistant is invoked. When the automated assistant is invoked, the user may provide a spoken utterance that includes the content in response to the received text message and content as an inquiry to the automated assistant (e.g., "I'm on my way now. Also, assistant, what's the weather like?").
[0017] In response to receiving the spoken utterance, one or more speech-to-text models may be used to process the audio data that is generated to represent the spoken utterance. The text resulting from the processing may be further processed to determine whether all of the text is incorporated into the selected input field, some of the text is incorporated into the input field, and / or other content is generated based on one or more portions of the resulting text. In some implementations, one or more models may be used to process the resulting text, and the one or more models may classify portions of the spoken utterance as being directed to the selected input field or being provided as a query to the automated assistant. Additionally or alternatively, context data representative of the context in which the user provided the spoken utterance may be used to perform the processing of the resulting text. For example, with the prior permission of the user, the context data may be based on prior input from the user, text accessible via the currently active messaging interface, and metadata stored in association with the selected input field and / or any other input fields accessible via third-party applications and / or separate applications on the computing device. The processing of portions of the resulting text may generate a bias that characterizes a portion of the text as being suitable for the selected input field and / or not suitable as a query to the automated assistant. For example, a portion of the resulting text may be designated as not suitable as an inquiry to the automated assistant but may be designated as suitable for incorporation into the selected input field based on the context data.
[0018] When the resulting text has been processed, a portion of the text may be designated for incorporation into the selected input field, and another portion of the text may be designated as a query to the automated assistant. For example, the first portion of the text may include "I am currently traveling to", which may be incorporated into the selected input field, and the second portion of the text may include "Also, assistant, what's the weather like", which may be processed as a query via the automated assistant. Thus, the automated assistant may cause the first portion of the text to be incorporated into the selected input field and may generate a response to the second portion of the text, such as "Tomorrow the weather will be cloudy, with a high of 72 and a low of 65".
[0019] In some implementations, when the first portion of the text (e.g., "I am currently traveling to") has been included in the text field for a response to a message, the automated assistant may wait for confirmation from the user that the user wants to send the message. For example, the first portion of the text may occupy the text field for responding to an incoming message, and the user may provide a spoken utterance such as "Send message" to the automated assistant to cause the message to be transmitted. In some implementations, the text field may be presented as an interface of a third-party messaging application, and thus, in response to receiving the spoken utterance "Send message", the automated assistant may communicate with the third-party messaging application to cause the message to be transmitted.
[0020] Alternatively or additionally, a user may choose to modify a message to change the content of a draft message before instructing an automated assistant via voice commands to send the message. For example, when the automated assistant has caused a text field of a messaging application to include the content of the spoken utterance "I am on my way now", the user may issue a subsequent spoken utterance such as "Delete the last word", "Replace the last word with 'in 15 minutes'", or "Replace 'now' with'soon'". In response, the automated assistant may interact with the messaging application to change the content of the draft message according to the subsequent spoken utterance from the user. In some implementations, the automated assistant may effect the change in content via one or more inputs to the operating system of the computing device on which the messaging application is executing. Alternatively or additionally, the automated assistant may effect the change in content by interfacing with the messaging application via an application programming interface (API) that allows the automated assistant to control the messaging application.
[0021] Additionally or alternatively, the automated assistant may effect the change in content via an API for interfacing with a keyboard application. The keyboard application may have the ability to interact with a third-party application such as a third-party messaging application to allow a user to edit content, perform searches, and / or perform any other functions of the third-party application that can be initialized at least via the separate keyboard application. The keyboard API may thus allow the automated assistant to also control the keyboard application to initiate the execution of such functions of the third-party application. Thus, in response to receiving a subsequent spoken utterance for editing the content of a draft message (e.g., "Delete the word 'now'."), the automated assistant may interact with the keyboard application via the keyboard API so that the word 'now' is deleted from the content of the draft message. Thereafter, the user may cause the message to be sent by providing the spoken utterance "Send message", which may cause the automated assistant to interface with the keyboard API so that a "send" command will be communicated to the messaging application - just as if the user had tapped the "send" button of the keyboard GUI.
[0022] In some implementations, if a user chooses to add additional content to a draft message, such as an image, an emoji, or other media, the user can provide a subsequent spoken utterance such as "Add 'thumbs up'". In response, the automated assistant can use the content of the subsequent spoken utterance, "thumbs up", to search for media data provided by the messaging application and / or the operating system in order to identify a suitable graphic to insert into the draft message. Alternatively or additionally, the user can have the automated assistant present search results for "GIFs" that can correspond to the content of the subsequent spoken utterance. For example, in response to receiving a subsequent spoken utterance such as "Show me 'thumbs up' GIFs", the automated assistant can cause the operating system and / or the messaging application to open the search interface of the keyboard GUI and cause a search query for "thumbs up" to be executed at the search interface. The user can then select the desired GIF by tapping a specific location on the keyboard GUI or by issuing another spoken utterance that describes the user's desired GIF. For example, the user can provide another spoken utterance such as "Third one" in order to have the automated assistant select the "third" GIF listed in the search results and incorporate the "third" GIF into the content of the draft message. Thereafter, the user can choose to send the message by issuing the spoken utterance "Send message".
[0023] The foregoing is provided as an overview of some implementations of the present disclosure. Further description of these and other implementations is provided below in more detail.
[0024] Other implementations may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform a method such as one or more of the methods described above and / or elsewhere herein. Other implementations may also include a system of one or more computers including one or more processors operable to execute the stored instructions to perform a method such as one or more of the methods described above and / or elsewhere herein.
[0025] It should be understood that all combinations of the foregoing concepts and additional concepts described in greater detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1A 、 Figure 1B and Figure 1C show views in which a user recites a reference to content to be incorporated into an input field with or without having to verbatim recite the content.
[0027] Figure 2A and Figure 2B shows a view in which a user recites a reference to image content to be incorporated into an input field without explicitly selecting the image content.
[0028] Figure 3 shows a system for providing an automated assistant that can optionally determine whether to incorporate a verbatim interpretation of a portion of spoken discourse and / or incorporate synonymous content into an input field.
[0029] Figure 4 shows a method for providing an automated assistant that can determine whether to merge one of multiple candidate interpretations into an input field.
[0030] Figure 5A and Figure 5B shows a method for providing an automated assistant that can determine whether to incorporate a verbatim interpretation and / or reference content into an input field selected by a user via an application interface.
[0031] Figure 6 is a block diagram of an example computer system. DETAILED DESCRIPTION
[0032] Figure 1A 、 Figure 1B 、 Figure 1C show views 100, 140, and 160, respectively, in which user 102 recites a reference to content to be incorporated into an input field with or without the need to verbatim recite the content. User 102 may initially interact with an application such as thermostat application 108 to establish a schedule 110 during which thermostat application 108 will operate according to a low power mode. For example, user 102 may interact with thermostat application 108 by using their hand 118 to select a specific input field of the schedule 110 interface of thermostat application 108. In response to user 102 selecting the input field, thermostat application 108, automated assistant 130, and / or the operating system of computing device 104 may optionally cause keyboard interface 112 to be presented at display panel 106 of computing device 104.
[0033] To provide content to be incorporated into the selected input field, user 102 may type content via keyboard interface 112 and / or provide spoken discourse that includes content and / or a reference to content. For example, as Figure 1AAs provided in, user 102 may provide a spoken utterance 116, such as "Assistant, what's the date tomorrow?". The audio data generated based on the spoken utterance may be processed at the data engine 126 of the computing device 104 to generate candidate content 122. The candidate content 122 may be generated based on auxiliary data 120, application data 124, and / or any other data accessible by the automated assistant 130. In some implementations, the application data 124 may indicate that the selected input field is formatted to receive a date. Thus, when the candidate content 122 is generated by the data engine 126, the data engine 126 may be biased towards selecting candidate content 122 formatted as a date relative to other candidate content 122.
[0034] Figure 1B FIG. 140 shows a view of a user interacting with a messaging application 146 to incorporate content into a selected input field (e.g., new message 148) of the messaging application 146. Additionally, the automated assistant 130 may further process the spoken utterance 142 from the user 102 to determine whether to generate content for the input field 150 that is the same as or different from the content (e.g., "3 / 14 / 19") generated in Figure 1A , even though the same spoken utterance (e.g., "what's the date tomorrow") is provided. For example, the audio data generated from the spoken utterance 142 may be converted to text, which may be processed in combination with the context data to determine the content suitable for incorporation into the input field 150. For example, one or more machine learning models may be used to process the text, and the one or more machine learning models are trained to provide candidate interpretations for the text based at least in part on the context in which the spoken utterance is provided. Thereafter, each candidate interpretation may be processed to identify the most suitable candidate interpretation in the given context.
[0035] For example, the context data generated based on the Figure 1A scenario may indicate that the user 102 is accessing a thermostat application 108 to provide a date for entering a schedule 110 for controlling the thermostat. Additionally, the context data may characterize the limitations of the field, such as the limitation that the field must receive at least a certain amount of numerical input representing a date. Thus, any candidate interpretation generated by the automated assistant 130 that includes numbers may be prioritized over candidate interpretations that do not include numbers. Additionally, the context data generated based on the Figure 1B scenario may indicate that the user 102 is accessing the messaging application 146 to provide a response message to a previously received message (e.g., "Irene: "What are the tickets for?""). The context data may characterize the state of the messaging application 146, the content presented at the interface of the messaging application 146 (e.g., the previous message from "Irene" and / or other previous messages), and / or any other information associated with the user's access to messages via the computing device 104.
[0036] In some implementations, a candidate interpretation of the spoken utterance 142 can be "the date tomorrow", and additional content can be generated based on the text of the spoken utterance 142. The candidate content 144 that can include the candidate interpretation and other content can be processed to identify the most suitable content to incorporate into the input field 150. For example, the input in the candidate content 144 can include "3 / 14 / 19", which can be the numerical representation of the day after the day when the user 102 accesses the messaging application 146. The candidate content 144 can be processed using one or more machine learning models, which can be different from the one or more machine learning models used to process the audio data representing the spoken utterance 142. Based on the processing of the candidate content 144, the candidate interpretation "the date tomorrow" can be prioritized over any other candidate content, and the candidate interpretation can be incorporated into the input field 150. Processing the spoken utterance in this way can conserve computational resources that might otherwise be wasted in cases where the user needs to more clearly specify the purpose of each part of the spoken utterance. As described above, the context information used in determining to provide the candidate interpretation "the date tomorrow" instead of the candidate content 144 "3 / 14 / 19" can include previous messages from "Irene". Note that in some implementations, if Irene's message were instead, for example, "What's the date again?", then the candidate content 144 "3 / 14 / 29" could be provided instead of the candidate interpretation "the date tomorrow".
[0037] In some implementations, as Figure 1C provided in the view 160 of, the natural language content of the spoken utterance 142 can be processed to determine whether the natural language content embodies an automated assistant command. For example, a natural language understanding engine can be used to process the natural language content to determine whether the intent corresponding to a particular assistant action is embodied in the natural language content. When the natural language content embodies the intent, the automated assistant can initiate the execution of the particular assistant action (e.g., recall the date to be "inserted" into the input field 150). However, when the intent is not embodied in the natural language content, one or more portions of the natural language content can be incorporated into the input field 150. In some implementations, the determination to incorporate the natural language content and / or alternatively the natural language content into the input field 150 can be based on one or more characteristics of the input field 150 (e.g., HTML, XML, text, and / or graphics within a threshold distance of the input field).
[0038] As Figure 1CAs shown, user 102 may receive a new message 148 at their computing device 104 and select a portion of the interface presenting the new message 148 to provide a response message to the sender (e.g., "Irene"). User 102 may cause the automated assistant 130 to incorporate the phrase "tomorrow's date" in their spoken utterance by omitting the natural language that would embody the assistant's intent (e.g., "respond...", "insert...", etc.), as Figure 1B shown. However, in order for the automated assistant 130 to incorporate alternative content (e.g., content that may be different from the verbatim natural language content), user 102 may incorporate the assistant's intent into the spoken utterance 162. For example, the spoken utterance 162 may include the command "insert tomorrow's date". Candidate content 164 may be generated in response to the spoken utterance 162, and the selection of the content to be incorporated into the input field 150 may be biased based on the context in which user 102 provides the spoken utterance 162. For example, the automated assistant 130 may determine that it is biased against inserting the verbatim content "insert tomorrow's date" into the input field 150, and instead is biased towards an alternative interpretation of the natural language content (e.g., "tomorrow's date") that is separate from the recognized intent (e.g., "insert"). As a result, appropriate additional content (e.g., March 15, 2019) may be identified and incorporated into the input field 150.
[0039] Figure 2A and Figure 2B show views 200 and 240, respectively, in which user 202 recites a reference to image content to be incorporated into an input field without having to explicitly select the image content. User 202 may initially interact with an application such as the messaging application 208 to send a new message 210 to a specific contact (e.g., Richard). For example, user 202 may interact with the messaging application 208 using their hand 218 to select a specific input field 246 of the graphical user interface of the messaging application 208. In response to user 202 selecting the input field, the messaging application 208, the automated assistant 230, and / or the operating system of the computing device 204 may optionally cause a keyboard interface 212 to be presented at the display panel 206 of the computing device 204. The keyboard interface 212 may include keyboard elements 232 that, when selected by the user, invoke the automated assistant 230 to assist the user in providing content to be incorporated into the selected input field 246. The automated assistant 230 may be an application separate from the keyboard application providing the keyboard interface 212. However, the automated assistant 230 may provide commands to the keyboard application based on the spoken utterance provided by user 202. The commands may be provided by the automated assistant 230 to the keyboard application via an API, interprocess communication, and / or any other technique for communicating between applications. In some implementations, the container of the keyboard application may be instantiated in memory while the container of the automated application is instantiated in memory.
[0040] To provide content to be incorporated into a selected input field 246, user 202 may type content via keyboard interface 212 and / or provide a spoken utterance that includes content and / or a reference to content. For example, as Figure 2A provided, user 202 may respond to a message (e.g., "Do you need anything for your house?"). The spoken utterance 216 may be "hammer emoji, nail emoji, and also, where can I book a truck?". Audio data generated based on the spoken utterance may be processed at data engine 226 of computing device 204 to generate candidate content 222. Candidate content 222 may be generated based on auxiliary data 220, application data 224, and / or any other data accessible by automated assistant 230. In some implementations, application data 224 may indicate that the selected input field is formatted to receive text and / or images.
[0041] Figure 2B A view 240 is shown of computing device 204 presenting response data 244 at messaging application 208 to incorporate content into a selected input field 246 of messaging application 208. For example, metadata may be stored at computing device 204 to characterize certain emojis. Thus, in response to receiving spoken utterance 216, computing device 204 may determine that the word "hammer" is stored associated with a particular emoji and the word "nail" is stored associated with a particular emoji. Then, each emoji may be incorporated into input field 246 as Figure 2B provided.
[0042] In addition, automated assistant 230 may further process spoken utterance 216 from user 202 to determine whether there are any other parts of spoken utterance 216 to which automated assistant 230 has not yet responded. For example, audio data generated from spoken utterance 216 may be converted to text, which may be parsed to identify separate parts of the text corresponding to separate intents of user 202. For example, one or more machine learning models may be used to process the text, the one or more machine learning models being trained to classify parts of natural language content that have been compiled together as a continuous input from user 202. Thereafter, a particular part of the text may be processed using another machine learning model (e.g., an image captioning model) corresponding to the classification of the text corresponding to that particular part, and another part of the text may be processed using a different machine learning model (e.g., a navigation search model) corresponding to another classification of the text corresponding to the other part of the text.
[0043] As an example, a first part of the text may be "hammer emoji, nail emoji", which when processed may generate candidate content 222 such asFigure 2B The file location of each emoji depicted in input field 246. Additionally, the second part of the text can be "Also, where can I book a truck", which when processed can generate additional candidate content 222, such as "Louisville Truck Company". While the automated assistant 230 incorporates candidate content 122 into input field 246, other candidate content 122 can subsequently be incorporated into the output 242 of the automated assistant 230. Processing spoken discourse in this manner can conserve computing resources that might otherwise be wasted in cases where the user needs to more clearly specify the purpose of each part of the spoken discourse.
[0044] Figure 3 A system 300 for an automated assistant is shown that provides for selectively determining whether to incorporate a verbatim interpretation of a portion of spoken discourse into an input field and / or incorporate synonymous content into the input field. The automated assistant 304 can operate as part of an assistant application provided at one or more computing devices, such as computing device 302 and / or server device. A user can interact with the automated assistant 304 via an assistant interface 320, which can be a microphone, camera, touch screen display, user interface, and / or any other device capable of providing an interface between the user and the application. For example, the user can initialize the automated assistant 304 by providing language input, text input, and / or graphical input to the assistant interface 320 such that the automated assistant 304 performs functions (e.g., provides data, controls peripheral devices, accesses agents, generates input and / or output, etc.). The computing device 302 can include a display device, which can be a display panel including a touch interface for receiving touch input and / or gestures for allowing the user to control the computing device 302 via the touch interface with an application 334. In some implementations, the computing device 302 can lack a display device, thereby providing an audible user interface output without providing a graphical user interface output. Additionally, the computing device 302 can provide a user interface, such as a microphone, for receiving spoken natural language input from the user. In some implementations, the computing device 302 can include a touch interface and can lack a camera, but can optionally include one or more other sensors.
[0045] The computing device 302 and / or other third-party client devices can communicate with the server device via a network (e.g., the Internet). Additionally, the computing device 302 and any other computing devices can communicate with each other via a local area network (LAN) such as a Wi-Fi network. The computing device 302 can offload computing tasks to the server device to conserve computing resources at the computing device 302. For example, the server device can host the automation assistant 304, and / or the computing device 302 can transmit inputs received at one or more assistant interfaces 320 to the server device. However, in some implementations, the automation assistant 304 can be hosted at the computing device 302 and various processes associated with the operation of the automation assistant can be performed at the computing device 302.
[0046] In various implementations, all or less than all aspects of the automation assistant 304 can be implemented on the computing device 302. In some of these implementations, aspects of the automation assistant 304 are implemented via the computing device 302 and can interface with the server device, which can implement other aspects of the automation assistant 304. The server device can optionally serve multiple users and their associated assistant applications via multiple threads. In implementations where all or less than all aspects of the automation assistant 304 are implemented via the computing device 302, the automation assistant 304 can be an application separate from the operating system of the computing device 302 (e.g., installed “on top” of the operating system) - or alternatively can be directly implemented by the operating system of the computing device 302 (e.g., considered an application that is part of but integrated with the operating system).
[0047] In some implementations, the automation assistant 304 can include an input processing engine 306, which can employ multiple different modules to process the input and / or output of the computing device 302 and / or the server device. For example, the input processing engine 306 can include a speech processing engine 308, which can process audio data received at the assistant interface 320 to identify the text embodied in the audio data. The audio data can be transmitted, for example, from the computing device 302 to the server device to conserve computing resources at the computing device 302. Additionally or alternatively, the audio data can be processed exclusively at the computing device 302.
[0048] The process for converting audio data to text can include a speech recognition algorithm, which can employ a neural network and / or a statistical model for identifying groups of audio data corresponding to words or phrases. The text converted from the audio data can be parsed by a data parsing engine 310 and made available as text data to an automated assistant 304 that can be used to generate and / or identify command phrases, intents, actions, slot values, and / or any other content specified by the user. In some implementations, the output data provided by the data parsing engine 310 can be provided to a parameter engine 312 to determine whether the user has provided input corresponding to a particular intent, action, and / or routine that can be performed by the automated assistant 304 and / or an application or agent accessible via the automated assistant 304. For example, assistant data 338 can be stored at the server device and / or the computing device 302 and can include data defining one or more actions that can be performed by the automated assistant 304 and the parameters necessary to perform those actions. The parameter engine 312 can generate one or more parameters for an intent, action, and / or slot value and provide the one or more parameters to an output generation engine 314. The output generation engine 314 can use the one or more parameters to communicate with an assistant interface 320 for providing output to the user and / or with one or more applications 334 for providing output to one or more applications 334.
[0049] In some implementations, the automated assistant 304 can be an application that can be installed "on top" of the operating system of the computing device 302 and / or can itself form part (or all) of the operating system of the computing device 302. The automated assistant application includes device-built-in speech recognition, device-built-in natural language understanding, and device-built-in implementation and / or access thereto. For example, device-built-in speech recognition can be performed using a device-built-in speech recognition module that processes (audio data detected by a microphone) using an end-to-end speech recognition machine learning model locally stored at the computing device 302. The device-built-in speech recognition generates recognition text for the spoken utterances (if any) present in the audio data. Additionally, for example, device-built-in natural language understanding (NLU) can be performed using a device-built-in NLU module that processes the recognition text generated using device-built-in speech recognition and optional context data to generate NLU data. The NLU data can include an intent corresponding to the spoken utterance and optionally parameters (e.g., slot values) for that intent.
[0050] The device - built implementation module can be used to perform device - built implementation. This module utilizes NLU data (from the device - built NLU) and optionally other local data to determine the actions to be taken to resolve the intent of the spoken utterance (and optionally the parameters for that intent). This can include determining local and / or remote responses (e.g., answers) to the spoken utterance, interactions with locally installed applications based on the spoken utterance, commands transmitted to Internet of Things (IoT) devices based on the spoken utterance (either directly or via a corresponding remote system), and / or other resolution actions performed based on the spoken utterance. Then, the device - built implementation can initiate the local and / or remote performance / execution of the determined actions to resolve the spoken utterance.
[0051] In various implementations, remote speech processing, remote NLU, and / or remote implementation can be at least selectively utilized. For example, the recognized text can be at least optionally transmitted to a remote automation assistant component for remote NLU and / or remote implementation. For example, the recognized text can be optionally transmitted for remote performance in parallel with the device - built performance, or in response to a failure of the device - built NLU and / or device - built implementation. However, device - built speech processing, device - built NLU, device - built implementation, and / or device - built execution can be prioritized at least due to the reduced latency they provide in resolving the spoken utterance (since there is no need for client - server round - trips to resolve the spoken utterance). Additionally, in the absence of or with limited network connectivity, device - built functions may be the only available functions.
[0052] In some implementations, the computing device 302 can include one or more applications 334, which can be provided by a third - party entity different from the entity providing the computing device 302 and / or the automation assistant 304. The application state engine 316 of the automation assistant 304 and / or the computing device 302 can access the application data 330 to determine one or more actions that can be performed by the one or more applications 334, and the state of each of the one or more applications 334. Additionally, the application data 330 and / or any other data (e.g., device data 332) can be accessed by the automation assistant 304 to generate context data 336, which can characterize the context in which a particular application 334 is being executed at the computing device 302 and / or the context in which a particular user is accessing the computing device 302 and / or accessing the application 334.
[0053] When one or more applications 334 are executing at computing device 302, device data 332 may characterize the current operating state of each application 334 executing at computing device 302. Additionally, application data 330 may characterize one or more features of the executing application 334, such as the content of one or more graphical user interfaces presented under the guidance of one or more applications 334. Alternatively or additionally, application data 330 may characterize action patterns that may be updated by the corresponding application and / or automated assistant 304 based on the current operating state of the corresponding application. Alternatively or additionally, one or more action patterns for one or more applications 334 may remain static, but may be accessed by application state engine 316 to determine appropriate actions initiated via automated assistant 304.
[0054] In some implementations, automated assistant 304 may include a field engine 324 for determining whether a particular input field selected by a user is associated with a particular characteristic and / or metadata. For example, a particular input field may be limited to numbers, letters, dates, years, currency, symbols, images, and / or any data that may be provided to an input field of an interface. When a user selects an input field and provides a spoken utterance, input processing engine 306 may process the audio data corresponding to the spoken utterance to generate natural language text characterizing the spoken utterance. Classification engine 322 of automated assistant 304 may process portions of the text to determine whether the text corresponds to a command directed to automated assistant 304 or other speech that may have been captured by computing device 302.
[0055] Text identified as being intended as a command for automated assistant 304 may be processed by field engine 324 to determine whether any portion of the text is suitable for incorporation into the input field selected by the user. When it is determined that a portion of the text is suitable for the input field, automated assistant 304 may incorporate that portion of the text into the input field. Additionally or alternatively, the text may be processed by candidate engine 318 to determine whether there are any alternative interpretations and / or reference content intended for incorporation into the input field. For example, candidate engine 318 may process one or more portions of the text of the spoken utterance with context data 336 to generate other suitable interpretations of the text and / or other suitable content (e.g., images, videos, audio, etc.). When a candidate interpretation is identified, the candidate interpretation may be processed by field engine 324 to determine whether the candidate interpretation is suitable for input into the input field. When the candidate interpretation is determined to be suitable for the input field and more relevant than any other candidate for input into the input field, automated assistant 304 may cause the candidate interpretation to be incorporated into the input field.
[0056] When a candidate interpretation is determined to be unsuitable for input into the input field, the text corresponding to the candidate interpretation can be further processed by the input processing engine 306. The input processing engine 306 can determine whether the candidate interpretation corresponds to a query that the automated assistant 304 can respond to. If the automated assistant 304 determines that the text corresponding to the candidate interpretation is a query that the automated assistant 304 can respond to, the automated assistant 304 can continue to respond to the query. In this way, users will not have to repeat their spoken utterances, even though the input field has been selected before the spoken utterance is provided. Additionally, this allows the automated assistant 304 to conserve computing resources by limiting how much audio data is buffered. For example, in this example, the user is able to have a smooth conversation with the automated assistant 304 without having to repeatedly provide a call phrase to initialize the automated assistant 304 each time, thus reducing the amount of audio data that occupies the memory of the computing device 302.
[0057] Figure 4 Method 400 for providing an automated assistant that can determine whether to incorporate a verbatim interpretation and / or reference content into an input field selected by a user via an application interface is shown. Method 400 can be performed by one or more computing devices, applications, and / or any other device or module that can be associated with the automated assistant. Method 400 can include an operation 402 of determining whether the user has selected an input field. When it is determined that the user has selected an input field, method 400 can proceed to an operation 404 of initializing the automated assistant and / or the keyboard interface. The automated assistant and / or the keyboard interface can be initialized when the user is accessing an application that includes the input field selected by the user.
[0058] Method 400 can proceed from operation 404 to an operation 406 of determining whether the user has provided a spoken utterance to the automated assistant. When it is determined that the user has provided a spoken utterance to the automated assistant, method 400 can continue to proceed to operation 410. When it is determined that the user has not provided a spoken utterance, method 400 can proceed from operation 406 to an operation 408 that can include determining whether the user has provided text input to the input field. When it is determined that the user has not provided text input to the input field, method 400 can return to operation 402. However, when it is determined that the user has provided text input to the input field, method 400 can proceed from operation 408 to an operation 416 of incorporating the text into the selected input field.
[0059] When it is determined that the user has provided a spoken utterance to the automated assistant, method 400 may proceed to operation 410, which may include generating candidate text strings that represent one or more portions of the spoken utterance. In some implementations, generating the candidate text strings may include generating text strings that are intended to be a verbatim representation of the spoken utterance. Method 400 may proceed from operation 410 to operation 412 for generating additional content based on the spoken utterance. The additional content may be generated based on the spoken utterance, the candidate text strings, and / or context data that represents the context in which the user provided the spoken utterance and / or historical data associated with the interaction between the user and the application.
[0060] Method 400 may proceed from operation 412 to operation 414, which may include processing the candidate text strings and / or the additional content to determine which to incorporate into the input field. In other words, operation 414 includes determining whether one or more portions of the candidate text strings and / or one or more portions of the additional content are to be incorporated into the selected input field. When determining whether to incorporate the candidate text strings or the additional content into the input field, method 400 may proceed from operation 414 to operation 416. Operation 416 may include incorporating content based on the spoken utterance into the selected input field. In other words, depending on whether the candidate text strings or the other content have been prioritized at operation 414, method 400 may incorporate the candidate text strings and / or the other content into the selected input field at operation 416.
[0061] Figure 5A and Figure 5B Methods 500 and 520 for providing an automated assistant that can determine whether to incorporate a verbatim interpretation and / or reference content into an input field selected by a user via an application interface are shown. Methods 500 and 520 may be performed by one or more computing devices, applications, and / or any other device or module that may be associated with the automated assistant. Method 500 may include an operation 502 of determining whether the user has selected an input field. An input field is one or more portions of an application interface that includes an area where the user specifies particular content for the input field. As an example, an input field may be a space on a web page where the user can specify their phone number. For example, a web page may include an input field such as "Phone number: ___-______" so that an entity that controls the web page can call the user regarding a service provided by the web page (e.g., a travel booking service). [[ID=~]]
[0062] When it is determined that an input field has been selected, method 500 may proceed from operation 502 to operation 504, which may include initializing a keyboard interface and / or an automated assistant. Otherwise, the application and / or the corresponding computing device may continue to monitor the input of the application. In some implementations, the keyboard interface and / or the automated assistant may be initialized in response to a user selecting an input field of the application interface. For example, an input field may be presented at a touch display panel of a computing device, and the user may select the input field by performing a touch gesture indicating the selection of the input field. Additionally or alternatively, the user may select the input field by providing a spoken utterance and / or providing input to a separate computing device.
[0063] Method 500 may proceed from operation 504 to operation 506, which may include determining whether the user has provided a spoken utterance to the automated assistant. After the user provides a selection of an input field, the user may optionally choose to provide a spoken utterance that is intended to help the automated assistant incorporate content into the input field and / or otherwise convey intent. For example, after the user selects a phone number input field on a web page, the user may provide a spoken utterance such as "three 5s, three 6s, 8..., assistant, what's the weather like tomorrow?". However, the user may alternatively choose not to provide a spoken utterance.
[0064] When the user chooses not to provide a spoken utterance, method 500 may proceed from operation 506 to operation 508, which may include determining whether the user has provided text input to the input field. When it is determined that the user has not provided text input to the input field, method 500 may return to operation 502. However, when it is determined that the user has provided text input to the input field, method 500 may proceed from operation 508 to operation 514. Operation 514 may include incorporating the text and / or content into the selected input field. Thus, if the user has provided text input (e.g., the user types "555-6668" on the keyboard). Or, when it is determined at operation 506 that the user has provided a spoken utterance after selecting the input field, method 500 may proceed from operation 506 to operation 510.
[0065] Operation 510 may include generating a candidate text string that represents one or more portions of the spoken utterance. For example, the spoken utterance provided by the user may be "Three 5s, three 6s, 8…, Assistant, what's the weather like tomorrow?" and a particular portion of the spoken utterance may be processed to provide a candidate text string such as "Three 5s, three 6s, 8". Method 500 may then proceed from operation 510 to operation 512 that may include determining whether the candidate text string is suitable for the selected input field. When it is determined that the candidate text string is suitable for the selected input field, method 500 may proceed to operation 514 to cause the candidate text to be incorporated into the selected input field. However, when it is determined that the candidate text string is not suitable for the selected input field, method 500 may continue from operation 512 via continuation element "A" to Figure 5B operation 518 of method 520 provided in
[0066] Method 520 may proceed from continuation element "A" to operation 518 that may include generating additional content based on the spoken utterance. The additional content may be generated based on the context data and / or any other information related to the situation in which the user selects the input field. For example, the application may include data that characterizes the characteristics of the selected input field, and the candidate text string may be processed based on the data to bias certain other content based on the candidate text string. For example, the other content may include other text interpretations such as "Three 5s, three 6s, 8" and "5556668". In addition, the data may characterize that the selected input field is specifically occupied by numbers and there is a requirement for at least 7 numbers. Therefore, the processing of the requirement data and the other content may result in a bias towards the text interpretation "5556668" compared to "Three 5s, three 6s, 8".
[0067] Method 520 may proceed from operation 518 to operation 522 that may include determining whether the additional content that has been generated is suitable for the selected input field. For example, the candidate content "5556668" may be determined to be suitable for the selected input field because the candidate content exclusively includes numbers, while the other candidate content "Three 5s, three 6s, 8" may be determined to be unsuitable because the other candidate content includes letters. When it is determined that no other content is suitable for the selected input field, method 500 may proceed from operation 522 to an optional operation 526 that may include prompting the user for further instructions regarding the selected input field. For example, the automated assistant may provide a prompt based on the context data and / or any other data associated with the selected input field. The automated assistant may, for example, present an audio output such as "This field is for numbers", which may be generated based on the determination that the other content is unsuitable because the other content includes letters. Thereafter, method 520 may proceed from operation 526 and / or operation 522 via continuation element "C" to Figure 5A operation 506 of method 500.
[0068] When it is determined that other content is suitable for the selected input field, method 520 may proceed from operation 522 to optional operation 524. Operation 524 may include updating the model based on the determination that other content is suitable for the selected input field. In this way, the automated assistant can adaptively learn whether the user is intending to facilitate the incorporation of content into the selected input field for certain spoken utterances and / or whether the user is instructing the automated assistant to take one or more other actions. Method 520 may proceed from operation 524 and / or operation 522 via continuation element "B" to operation 514 where the automated assistant may cause other content (e.g., other content biased with respect to any other generated content) to be incorporated into the selected input field (e.g., "Phone number: 555-6668").
[0069] Method 500 may proceed from operation 514 to operation 516 for determining whether other parts of the spoken utterance correspond to assistant operations. In other words, the automated assistant may determine whether any other specific part of the candidate text string (e.g., other than the part that is the basis for the other content) is provided to facilitate one or more operations to be performed by the automated assistant. When it is determined that the candidate text string and / or other parts of the spoken utterance do not facilitate assistant operations, method 500 may proceed to operation 502.
[0070] However, when another part of the candidate text string is determined to have been provided to facilitate automated assistant operations (e.g., "Assistant, what's the weather like tomorrow"), method 500 may proceed via method 500 and method 520 to operation 526 where the corresponding automated assistant operation is performed based on the other part of the candidate text string. In other words, since the other part of the candidate text string may not result in any other content suitable for the selected input field at operation 522, and since the other candidate text string refers to automated assistant operations at operation 526, the automated assistant operation may be performed. In this way, the user can rely on a more efficient voice method to simultaneously provide input to the interface to fill certain fields and instruct the automated assistant to take other actions.
[0071] Figure 6 is a block diagram of an example computer system 610. The computer system 610 generally includes at least one processor 614 that communicates with a plurality of peripheral devices via a bus subsystem 612. These peripheral devices may include a storage subsystem 624 (which includes, for example, a memory 625 and a file storage subsystem 626), a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow a user to interact with the computer system 610. The network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computer systems.
[0072] The user interface input device 622 can include a keyboard; a pointing device such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touchscreen incorporated into a display; an audio input device such as a voice recognition system, microphone; and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computer system 610 or a communication network.
[0073] The user interface output device 620 can include a display subsystem, printer, fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or other mechanisms for creating visual images. The display subsystem can also provide a non-visual display such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from the computer system 610 to a user or another machine or computer system.
[0074] The storage subsystem 624 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 624 can include logic for performing selected aspects of methods 400, 500, and / or 520, and / or implementing one or more of the system 300, computing device 104, computing device 204, and / or any other applications, devices, apparatuses, and / or modules discussed herein.
[0075] These software modules are typically executed by the processor 614 alone or in combination with other processors. The memory 625 used in the storage subsystem 624 can include multiple memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 for storing fixed instructions. The file storage subsystem 626 can provide persistent storage for program and data files and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical disk drive, or a removable media cartridge. Modules implementing the functionality of certain implementations can be stored in the storage subsystem 624 by the file storage subsystem 626 or in other machines accessible by the processor 614.
[0076] The bus subsystem 612 provides a mechanism for enabling the various components and subsystems of the computer system 610 to communicate with each other as expected. Although the bus subsystem 612 is schematically shown as a single bus, alternative implementations of the bus subsystem can use multiple buses.
[0077] The computer system 610 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 6 the description of the computer system 610 depicted in Figure 6 is only intended as a specific example for the purpose of illustrating some implementations. Many other configurations of the computer system 610 may have more or fewer components than the computer system depicted in
[0078] In situations where the systems described herein collect personal information about a user (or often referred to herein as a "participant") or where personal information may be used, the user can be provided with an opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social behavior or activities, occupation, user preferences, or the user's current geographical location), or to control whether and / or how content is received from a content server that may be more relevant to the user. Additionally, certain data can be processed in one or more ways before being stored or used so that personally identifiable information is removed. For example, a user's identity can be processed so that no personally identifiable information can be determined for that user, or the geographical location of a user from whom geographical location information is obtained can be generalized (such as to the city, postal code, or state level) so that the user's specific geographical location cannot be determined. Thus, the user can control how information about the user is collected and / or used.
[0079] Although several implementations have been described and illustrated herein, various other means and / or structures can be utilized for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each of these variations and / or modifications is considered to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the specific one or more applications for which the teachings are used. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific implementations described herein. Thus, it is to be understood that the foregoing implementations are presented only as examples, and that within the scope of the appended claims and their equivalents, implementations may be practiced in a different manner than specifically described and claimed. The implementations of the present disclosure relate to each and every individual feature, system, article, material, kit, and / or method described herein. Moreover, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
[0080] In some implementations, a method implemented by one or more processors is described as including operations such as determining that a selection of an input field is provided to a graphical user interface of an application presented at a computing device. The computing device may provide access to an automated assistant separate from the application and utilize one or more speech-to-text models stored at the computing device. The method may further include an operation of receiving a spoken utterance from a user after determining that the input field has been selected. The method may further include an operation of generating a candidate text string representing at least a portion of the spoken utterance provided by the user, wherein the candidate text string is generated using one or more speech-to-text models stored at the computing device. The method may further include an operation of determining, by the automated assistant and based on the candidate text string, whether to incorporate the candidate text string into the input field or to incorporate additional content into the input field. The method may further include an operation of causing the additional content to be provided as an input to the input field of the graphical user interface when a determination is made to incorporate the additional content into the input field, wherein the additional content is generated via execution of one or more automated assistant actions based on the candidate text string. The method may further include an operation of causing the candidate text string to be provided as an input to the input field of the graphical user interface when a different determination is made to generate additional content for incorporation into the input field.
[0081] In some implementations, determining whether to incorporate the candidate text string into the input field or to incorporate additional content into the input field includes: determining whether the input field is limited to a particular type of input content, and determining whether the candidate text string corresponds to the particular type of input content associated with the input field. In some implementations, the particular type of input content includes contact information of the user or another person, and causing the additional content to be generated via execution of the one or more automated assistant actions includes: determining that contact data accessible via the computing device includes the additional content and is stored in association with at least a portion of the candidate text string.
[0082] In some implementations, the additional content has no text characters derived from the user's primary language. In some implementations, the additional content includes at least one image. In some implementations, determining that a selection of the input field is provided to the graphical user interface includes: determining that a keyboard interface is being presented on the graphical user interface of the application. In some implementations, causing the candidate text string to be provided as an input to the input field of the graphical user interface includes: causing the candidate text string to be provided as an input to a keyboard application from the automated assistant, where the keyboard application provides the keyboard interface presented on the graphical user interface. In some implementations, the keyboard application is an application separate from the automated assistant. In some implementations, the additional content is generated without further using the one or more speech-to-text models and is different from the candidate text string.
[0083] In other implementations, a method implemented by one or more processors is set forth as including operations such as receiving spoken utterances from a user when an input field of a graphical user interface of an application is presented at a computing device, where the computing device provides access to an automated assistant separate from the application and utilizes one or more speech-to-text models stored at the computing device. In some implementations, the method may further include an operation of generating a candidate text string representing at least a portion of the spoken utterance provided by the user based on the spoken utterance, where the candidate text string is generated using the one or more speech-to-text models stored at the computing device. In some implementations, the method may further include an operation of determining, by the automated assistant and based on the candidate text string, whether any particular portion of the candidate text string is to be regarded as invoking a particular assistant operation, the particular assistant operation being different from an operation of incorporating a particular portion of the candidate text string into the input field and different from another operation of incorporating additional content into the input field.
[0084] In some implementations, the method may further include an operation of causing the automated assistant to perform the specific assistant operation at least based on a specific portion of the candidate text string when a determination is made to consider the specific portion of the candidate text string as invoking the assistant operation. In some implementations, the method may further include the following operation: when a different determination is made not to consider the specific portion of the candidate text string as not invoking the assistant operation: determining, by the automated assistant and based on the specific portion of the candidate text string, whether to incorporate the specific portion of the candidate text string into the input field or incorporate the additional content into the input field, and providing, based on the determination of whether to incorporate the specific portion of the candidate text string into the input field or incorporate the additional content into the input field, the specific portion of the candidate text string or the additional content as an input to the input field of the graphical user interface.
[0085] [[ID=**3**]]In some implementations, determining whether to incorporate the specific portion of the candidate text string into the input field or incorporate the additional content into the input field includes: determining whether the input field is limited to a specific type of input content, and determining whether the specific portion of the candidate text string corresponds to a specific type of input content associated with the input field. In some implementations, the additional content is generated without further using the one or more speech-to-text models and is different from the candidate text string. In some implementations, the additional content does not have text characters derived from the user's primary language. In some implementations, the additional content includes at least one image. In some implementations, causing the automated assistant to perform the assistant operation includes performing a web search based on the specific portion of the candidate text string.
[0086] In yet another implementation, a method implemented by one or more processors is described as including operations such as receiving spoken discourse from a user when accessing an application via a computing device, where the computing device provides access to an automated assistant separate from the application and utilizes one or more speech-to-text models. The method may further include an operation of generating a candidate text string representing at least a portion of the spoken discourse provided by the user based on the spoken discourse, where the candidate text string is generated using one or more speech-to-text models stored at the computing device. The method may further include an operation of determining, by the automated assistant and based on the candidate text string, whether to provide the candidate text string as an input to the application or to provide the additional content as an input to the application. The method may further include the following operations: when a determination is made to incorporate the additional content into the input field: generating the additional content based on the candidate text string and context data, where the context data characterizes the context in which the user provided the spoken discourse, and causing the additional content to be provided as an input to the application.
[0087] In some implementations, the method may further include an operation of causing the candidate text string to be provided as an input to the application when a different determination is made to provide the candidate text string as an input to the application. In some implementations, the method may further include bypassing presenting the candidate text string at a graphical user interface of the computing device when a determination is made to incorporate the additional content into the input field. In some implementations, the context data characterizes the content of the graphical user interface of the application. In some implementations, the context data characterizes one or more previous interactions between the user and the automated assistant. In some implementations, the context data characterizes formatting restrictions on content to be incorporated into the input field of the application.
[0088] In yet another implementation, a method implemented by one or more processors is described as including operations such as determining to provide a selection of a keyboard element to a graphical user interface of a keyboard application being presented at a computing device, where the computing device provides access to an automated assistant separate from the keyboard application and utilizes one or more speech-to-text models. The method may further include an operation of receiving a spoken utterance from a user after determining that the keyboard element has been selected, where the user is accessing a specific application including an input field when providing the spoken utterance. The method may further include an operation of generating a candidate text string representing at least a portion of the spoken utterance provided by the user based on the spoken utterance, where the candidate text string is generated using one or more speech-to-text models stored at the computing device. The method may further include an operation of determining by the automated assistant and based on the candidate text string whether to incorporate the candidate text string into the input field or to incorporate additional content into the input field. The method may further include the following operation: when a determination is made to incorporate the additional content into the input field: causing the additional content to be provided as an input to the input field of the graphical user interface, where the additional content is generated via the execution of one or more automated assistant actions based on the candidate text string.
[0089] In some implementations, the method may further include an operation of causing the candidate text string to be provided as an input to the input field of the graphical user interface when a different determination is made to generate additional content for incorporation into the input field. In some implementations, the additional content does not have text characters derived from the user's primary language. In some implementations, the additional content includes at least one image. In some implementations, the additional content is generated without further using the one or more speech-to-text models and is different from the candidate text string.
Claims
1. A method implemented by one or more processors, the method comprising: Determining that a selection of an input field of a graphical user interface of an application presented at a computing device is provided, wherein the computing device provides access to an automated assistant separate from the application and utilizes one or more speech-to-text models stored at the computing device; After determining that the input field is selected, receiving spoken discourse from a user; Generating a candidate text string representing at least a portion of the spoken discourse provided by the user based on the spoken discourse, wherein the candidate text string is generated using one or more speech-to-text models stored at the computing device; Determining by the automated assistant and based on the candidate text string whether to incorporate the candidate text string into the input field or whether to incorporate non-text visual content into the input field and replace the candidate text string, wherein determining by the automated assistant and based on the candidate text string whether to incorporate the candidate text string into the input field or whether to incorporate non-text visual content into the input field includes: Determining the non-text visual content based on processing the candidate text string; Identifying one or more non-text visual content characteristics of the non-text visual content based on the non-text visual content; and Determining whether to incorporate the candidate text string into the input field or whether to incorporate the non-text visual content into the input field based on comparing the one or more non-text visual content characteristics of the non-text visual content with one or more input field characteristics of the input field; When a determination is made to incorporate the non-text visual content into the input field: Causing the non-text visual content to be provided as an input to the input field of the graphical user interface, wherein the non-text visual content is determined via execution of one or more automated assistant actions based on the candidate text string; and When a different determination is made to incorporate the candidate text string into the input field: Causing the candidate text string to be provided as an input to the input field of the graphical user interface.
2. The method according to claim 1, wherein determining by the automated assistant and based on the candidate text string whether to incorporate the candidate text string into the input field or whether to incorporate the non-text visual content into the input field further includes: Determining whether the input field is limited to a specific type of input content, and Determining whether the candidate text string corresponds to the specific type of input content associated with the input field.
3. The method according to claim 1, wherein the non-text visual content does not have text characters derived from the user's primary language.
4. The method according to claim 1, wherein the non-text visual content includes at least one image.
5. The method according to claim 1, wherein determining that a selection of an input field of the graphical user interface is provided includes: Determining that a keyboard interface is being presented on the graphical user interface of the application.
6. The method according to claim 5, wherein causing the candidate text string to be provided as an input to the input field of the graphical user interface includes: causing the candidate text string to be provided as an input to a keyboard application from the automated assistant, wherein the keyboard application provides a keyboard interface presented on the graphical user interface.
7. The method according to claim 6, wherein the keyboard application is an application separate from the automated assistant.
8. The method according to claim 1, wherein the non-text visual content is determined without further using the one or more speech-to-text models and is different from the candidate text string.
9. The method according to claim 1, wherein determining whether to incorporate the candidate text string or the non-text visual content into the input field by the automated assistant and based on the candidate text string further includes: determining whether the candidate text string corresponds to any automated assistant commands.
10. The method according to claim 9, wherein determining whether the candidate text string corresponds to any automated assistant commands includes: using a natural language understanding (NLU) engine of the automated assistant to process the candidate text string, and determining whether an automated assistant intent is embodied in the candidate text string based on the processing using the NLU engine.
11. The method according to claim 1, wherein the non-text visual content is determined based on a portion of the candidate text string, and wherein, When a determination is made to incorporate the non-text visual content into the input field, the method further includes: causing an additional portion of the candidate text string to be provided as an additional input to the input field of the graphical user interface along with the non-text visual content.
12. The method according to claim 1, further including: determining one or more input field characteristics of the input field based on: HTML or XML tags associated with the input field, or text and / or graphics within a threshold distance of the input field.
13. The method according to any one of claims 1-12, wherein determining whether to incorporate the candidate text string or the non-text visual content into the input field by the automated assistant and based on the candidate text string further includes: determining whether one or more initial items of the candidate text string match one or more predefined items; and biasing the determination to incorporate the non-text visual content in response to determining that the one or more initial items match the one or more predefined items.
14. A method implemented by one or more processors, the method including: when an input field of a graphical user interface of an application is presented at a computing device, receiving spoken discourse from a user, wherein the computing device provides access to an automated assistant separate from the application and utilizes one or more speech-to-text models stored at the computing device; generating a candidate text string representing at least a portion of the spoken discourse provided by the user based on the spoken discourse, wherein the candidate text string is generated using one or more speech-to-text models stored at the computing device; Determine, by the automated assistant and based on the candidate text string, whether any particular part of the candidate text string is to be regarded as invoking a particular assistant operation, the particular assistant operation being different from the operation of incorporating a particular part of the candidate text string into the input field and different from another operation of incorporating additional content into the input field, wherein the additional content includes at least one image; When a determination is made that a particular part of the candidate text string is to be regarded as invoking the assistant operation: Cause the automated assistant to perform the particular assistant operation based at least on the particular part of the candidate text string; and When a different determination is made that a particular part of the candidate text string is not to be regarded as not invoking the assistant operation: Determine, by the automated assistant and based on the particular part of the candidate text string, whether to incorporate the particular part of the candidate text string into the input field or incorporate the additional content into the input field, and Based on the determination of whether to incorporate the particular part of the candidate text string into the input field or incorporate the additional content into the input field, cause the particular part of the candidate text string or the additional content to be provided as an input to the input field of the graphical user interface.
15. The method according to claim 14, wherein determining whether to incorporate the particular part of the candidate text string into the input field or incorporate the additional content into the input field includes: Determining whether the input field is limited to a particular type of input content, and Determining whether the particular part of the candidate text string corresponds to the particular type of input content associated with the input field.
16. The method according to claim 14, wherein the additional content is generated without further using the one or more speech-to-text models and is different from the candidate text string.
17. The method according to claim 14, wherein the additional content has no text characters derived from the user's primary language.
18. The method according to any one of claims 14-17, wherein causing the automated assistant to perform the assistant operation includes performing a web search based on the particular part of the candidate text string.
19. A method implemented by one or more processors, the method comprising: Receiving, from a user, a spoken utterance when accessing an application via a computing device, wherein the computing device provides access to an automated assistant separate from the application and utilizes one or more speech-to-text models; Generating a candidate text string representing at least a part of the spoken utterance provided by the user based on the spoken utterance, wherein the candidate text string is generated using one or more speech-to-text models stored at the computing device; Determining, by the automated assistant and based on the candidate text string, whether to provide the candidate text string as an input to the application or provide additional content as an input to the application; and When a determination is made to incorporate the additional content into the application: Generating the additional content based on the candidate text string and context data wherein the context data characterizes the context in which the user provided the spoken utterance, and causing the additional content to be provided as an input to the application; when making a different determination to provide the candidate text string as an input to the application: causing the candidate text string to be provided as an input to the application.
20. The method according to claim 19, further comprising: when making a determination to incorporate the additional content into the application: bypassing presenting the candidate text string at the graphical user interface of the computing device.
21. The method according to claim 19, wherein the context data characterizes the content of the graphical user interface of the application.
22. The method according to claim 19, wherein the context data characterizes one or more previous interactions between the user and the automated assistant.
23. The method according to any one of claims 19-22, wherein the context data characterizes formatting restrictions on content to be incorporated into an input field of the application.
24. A method implemented by one or more processors, the method comprising: determining that a selection of a keyboard element is provided to a graphical user interface of a keyboard application being presented at a computing device, wherein the computing device provides access to an automated assistant separate from the keyboard application and utilizes one or more speech-to-text models; after determining that the keyboard element is selected, receiving a spoken utterance from the user, wherein the user is accessing a specific application including an input field when the user provides the spoken utterance; generating a candidate text string representing at least a portion of the spoken utterance provided by the user based on the spoken utterance, wherein the candidate text string is generated using one or more speech-to-text models stored at the computing device; determining by the automated assistant and based on the candidate text string whether to incorporate the candidate text string into the input field or to incorporate additional content into the input field; and when making a determination to incorporate the additional content into the input field: causing the additional content to be provided as an input to the input field of the graphical user interface, wherein the additional content is generated via execution of one or more automated assistant actions based on the candidate text string.
25. The method according to claim 24, further comprising: when making a different determination to generate additional content for incorporation into the input field: causing the candidate text string to be provided as an input to the input field of the graphical user interface.
26. The method according to claim 24, wherein the additional content has no text characters derived from the user's primary language.
27. The method according to claim 24, wherein the additional content includes at least one image.
28. The method according to any one of claims 24-27, wherein the additional content is generated without further using the one or more speech-to-text models and is different from the candidate text string.
29. A computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 28.
30. A computer-readable storage medium comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 28.
31. A system comprising one or more processors for performing the method according to any one of claims 1 to 28.
Citation Information
Patent Citations
Voice interface ocx
US20100169092A1
Graphical keyboard application with integrated search
US20170308247A1
Speech-to-text conversion based on user interface state awareness
US20190214013A1