Arrange and / or clear speech-to-text content without requiring explicit user instructions.
By using speech processing and natural language understanding technologies, the automated assistant can automatically format text content based on the user's spoken words, solving the problem of users having to manually specify formatting operations, improving operational efficiency and reducing computing resource consumption.
Patent Information
- Application Number
- CN202180068837.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-06-03
- Filing Date
- 2021-12-10
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2041-12-10
AI Technical Summary
In existing technologies, when users perform voice-to-text operations using automated assistants, they need to manually specify the formatting of the text content, which leads to prolonged human-computer dialogue and wasted computing resources.
By using speech processing and natural language understanding technologies, an automated assistant can automatically determine the formatting intent of text content based on the user's spoken words and generate corresponding layout data to achieve automatic formatting of text content.
It reduces the need for users to explicitly specify text formatting, shortens human-computer interaction time, reduces computing resource consumption, and improves operational efficiency.
Smart Images

Figure CN116348844B_ABST
Abstract
Description
Background Technology
[0001] Humans can use interactive software applications to engage in human-computer dialogue; these applications are referred to herein as "automated assistants" (also known as "chatbots," "interactive personal assistants," "intelligent personal assistants," "personal voice assistants," "conversational agents," etc.). Automated assistants typically rely on a pipeline of components to interpret and respond to user input. For example, a speech processing engine can be used to process audio data capturing a user's spoken words and generate textual context, such as transcriptions of spoken words (i.e., sequences of terms and / or other lexical units). Furthermore, a natural language understanding (NLU) engine can be used to process the text content and generate NLU outputs, such as the user's intent when providing spoken words and slot values for parameters optionally associated with that intent.
[0002] In some cases, users may use an automation assistant and / or software applications accessible to the automation assistant (also referred to as "applications") to perform certain speech-to-text operations using a speech processing engine. For example, a user may use an automation assistant and / or an application to dictate text content on their behalf, and this text content may be incorporated into a document (e.g., a word processing document, email, text message, etc.). However, after the text content is incorporated into the document, the user typically must manually manipulate the text content using layout operations to format it with the desired layout (i.e., spacing, punctuation, capitalization, indentation, etc.). For example, the user may provide additional spoken commands to perform some of these layout operations, such as providing the spoken command "comma" to add a comma to the document, "newline" to start a new line in the document, "indent" to add indentation to the document, and / or providing similar layout operation commands via a separate keyboard or computer mouse interface.
[0003] Furthermore, in some of these cases, the user can provide verbal commands such as "clear" or "delete" to remove a portion of text already incorporated into the document. However, it may not be immediately clear which text the user wishes to remove from the document. Therefore, whenever the user provides one of these verbal commands, the application may remove a standard length of text. For example, whenever the user provides one of these verbal commands, the application may delete only a single letter or word, regardless of any context associated with the document or the spoken utterance containing the verbal command. Thus, when selecting specific text to be deleted, the user may again rely on other separate interfaces, such as a keyboard or computer mouse, rather than an audio interface. As a result, the computational resources of the device used to perform these voice-to-text operations when dictating text to be incorporated into the document may be wasted and / or the interaction with automation assistants and / or applications may be prolonged, based on the user having to provide these specific verbal commands to achieve the desired arrangement, based on subsequent manual manipulation of the text content, and / or based on the user switching between these separate interfaces. Summary of the Invention
[0004] This paper describes an implementation involving an automated assistant and / or application that can perform speech-to-text operations in response to spoken utterances received from a user. These operations involve organizing text content corresponding to the spoken utterances in a manner that the user may not explicitly specify. In this way, users can perform speech-to-text operations using the automated assistant and / or application without having to explicitly identify each operation that should be performed to render the desired text content. For example, a user can provide a spoken utterance to the automated assistant to perform a speech-to-text operation to facilitate drafting an email invitation. The spoken utterance can be provided while the user is accessing an email application and after the user has initialized the automated assistant to detect speech content that will be converted into text content to be incorporated into fields of the email application. For example, when the user has initialized an audio interface for communicating with the automated assistant, the user can provide a spoken utterance such as, “Hi Adam…Friendly follow up to send me those meeting notes. Take care, Ronald.” Although the user does not explicitly identify any formatting, punctuation, capitalization, and / or other guidelines for arranging text content as shown in spoken language, the automation assistant can generate content layout data that characterizes the layout of text content to be included in fields of an email application and guides the email application to arrange the text content accordingly.
[0005] In various implementations, in response to received spoken words, the automation assistant can use a speech processing engine to process the audio data corresponding to the spoken words to generate text content data representing the text content to be incorporated into fields of a document (e.g., a word processing document, email, text message, etc.) that will be included in the application. Furthermore, the automation assistant can determine the user's intent when drafting the document to generate content layout data representing the arrangement of fields within the document for incorporating the text content. The content layout data can identify one or more actions the application needs to perform to arrange the text content according to the layout and based on the user's spoken words.
[0006] In some versions of these implementations, a user's intent to draft a document can be determined based on spoken utterance. For example, one or more heuristic processes and / or one or more trained machine learning models (e.g., Natural Language Understanding (NLU) models) can be used to process audio data corresponding to spoken utterances and / or text content generated based on spoken utterances, and the user's intent can be determined based on the output generated using one or more heuristic processes and / or one or more trained machine learning models. For example, suppose a user provides the spoken utterance “Assistant, email Adam: Hi Adam…Friendly follow up to send me those meeting notes. Take care, Ronald”. In this case, the automated assistant can determine based on this processing that the user intends to use an email application to draft an email document, which may exhibit one or more first formatting features instructing the email document, even if the user does not explicitly identify these actions. One or more first formatting features may include, for example, a comma after a greeting in the text content (e.g., “HiAdam,”), one or more new paragraph lines or carriage returns after a comma, one or more periods (e.g., after “notes”), one or more new paragraph lines or carriage returns before the closing remarks (e.g., “Take care, Ronald”), a signature block after the closing remarks, etc. In some versions of these implementations, these first formatting features may be identified based on previous interactions between the user and the automated assistant via an email application and / or previous interactions between one or more other users and one or more other devices via the corresponding email application.
[0007] In contrast, suppose a user provides the spoken phrase, “Assistant, text Adam: Hi Adam…Friendly follow up to send me those meeting notes. Take care, Ronald.” In this case, the automated assistant can determine, based on this processing, that the user intends to use a text messaging application to draft a text message document. This text messaging application may embody one or more second formatting features that indicate the text message document, even if the user does not explicitly identify these actions, and these second formatting features are different from the one or more first formatting features that indicate the aforementioned email document. The one or more second formatting features may also include, for example, one or more punctuation features of the first formatting features, such as a comma after a greeting in the text content (e.g., “Hi Adam,”) and one or more periods (e.g., after “notes”). However, the second formatting features may additionally or alternatively include one or more second punctuation features that are different from the first formatting features. For example, a period could be provided after a greeting instead of a comma (e.g., “Hi Adam.”), and a period could be omitted after “notes,” replaced by a long dash or no punctuation at all. Furthermore, one or more second formatting features may omit one or more newline characters or carriage returns after a comma, one or more newline characters or carriage returns before a closing statement (e.g., “Take care, Ronald”), or a signature block after a closing statement, as text messaging documents are less formal. Similarly, in some versions of these implementations, these second formatting features can be identified based on previous interactions between the user and the automated assistant via the text messaging application and / or previous interactions between one or more other users and one or more other devices via the corresponding text messaging application. For example, previous interactions via the text messaging application might indicate that a greeting most often (or always) is followed by a period instead of a comma, and a period could be provided after the greeting based on such previous interactions (e.g., “Hi Adam.”).
[0008] In additional or alternative versions of these implementations, the user's intent in drafting the document can be determined based on the type of application in which the user initializes the audio interface for communicating with the automation assistant. For example, suppose a user provides the spoken phrase "Hi Adam…Friendly follow up to send me those meeting notes. Take care, Ronald" after initializing the audio interface from an email application. In this example, the automation assistant can generate one or more first formatting features instructing the aforementioned email document, even if the user does not explicitly identify these actions. In contrast, suppose a user provides the spoken phrase "Hi Adam…Friendly follow up to send me those meeting notes. Take care, Ronald" after initializing the audio interface from a text messaging application. In this example, the automation assistant can generate one or more second formatting features instructing the aforementioned text message document, even if the user does not explicitly identify these actions. In some versions of those implementations, the closing phrase can be automatically incorporated into the document based on the application type (e.g., when the application is an email application), even if the user does not explicitly identify it in the spoken phrase; in other versions of those implementations, the closing phrase can be omitted based on the application type (e.g., when the application is a text messaging application).
[0009] In some implementations, the application indicated in the user's speech and / or the user's initialization of the audio interface can be used, additionally or alternatively, to determine one or more formatting features. For example, suppose a user provides the spoken statement "Hi Catherine Jane wanted to let you know that swordfish is delayed two weeks" after initializing an automation assistant in a given application. In this first scenario, suppose data provided by and / or otherwise associated with the given application indicates that the user's contact is "Catherine Jane" and that the user lacks a "Catherine" contact (no "Jane") and also indicates that "swordfish" is a defined item being tracked via the application. In this first scenario, a comma can be inserted immediately after "Catherine Jane," followed by a carriage return, and "Swordfish" can be capitalized (i.e., the formatting would reveal that the message was sent to "Catherine Jane" and that the sending user is indicating that the Swordfish item will be delayed). In the second scenario, suppose data provided by and / or otherwise associated with a given application indicates that the user's contacts include "Catherine" (without "Jane") and "Jane" (without "Catherine"), and also lacks any indication that "swordfish" is a defined item or other defined entity of the application. In the second scenario, a comma can be inserted after "Catherine," followed by a carriage return, and "swordfish" will not be capitalized (i.e., the formatting will reveal that the message was sent to "Catherine," and that "Jane" has indicated that swordfish (e.g., the shipment of the fish) is delayed).
[0010] In various implementations, content placement data may be generated additionally or alternatively based on one or more other signals associated with spoken utterance. For example, content placement data may be generated additionally or alternatively based on one or more acoustic (or vocal) features of spoken utterance, determined by processing audio data capturing spoken utterance using one or more heuristic processes and / or one or more trained acoustic-based machine learning models. These acoustic features may include, for example, one or more prosodic attributes of spoken utterance, such as intonation, tone, stress, rhythm, beat, and / or pauses. For example, one or more pauses in spoken utterance (e.g., an ellipsis in the spoken utterance “Hi Adam…Friendly follow up to send me those meeting notes. Take care, Ronald”) may indicate that the user expects a comma after the greeting (e.g., “Hi Adam”) and / or that the user expects one or more new paragraph lines or carriage returns after the greeting. Furthermore, for example, the tone of spoken language can indicate the user's expectation of a professional tone, and therefore, a signature block can be inserted after the closing remarks (e.g., "Take care, Ronald").
[0011] In some implementations, even if the user may not specifically describe each operation to be performed, the modification of text content already incorporated into a document can be performed by the automation assistant in response to a user's request. For example, if a user wants to remove text content from a document, the user can verbally describe the text content to be removed. However, the user is more likely to simply provide verbal utterances such as "delete," "clear," or "backspace," which may have different meanings in different contexts. In response to receiving a command to remove some text content, the automation assistant can determine the amount of content to be removed from the text content based on the context of the text content and / or verbal utterances previously provided by the user. This determination may additionally or alternatively be based on one or more previous interactions between the user and the automation assistant, the user and another application, and / or one or more other users and / or one or more other computing devices via the respective application.
[0012] For example, based on previous spoken statements in which the user described individual characters of a proper noun (e.g., “C, H, A, N, C…”), the automation assistant can determine that when the user provides a subsequent “clear” command, that subsequent “clear” command is intended to remove only that single character (e.g., “C”) from the proper noun from which the user described a single character. Furthermore, when the user provides a further subsequent “clear” command, the automation assistant can determine that the further subsequent “clear” command is intended to remove only the previous single character (e.g., “N”) from the proper noun from which the user described a single character. In contrast, based on previous spoken statements by the user word-for-word (e.g., “We look forward to…”), the automation assistant can determine that when the user provides a subsequent “clear” command, that “clear” is intended to remove a single word recently incorporated into the application field (e.g., “to”) or a sequence of single words referring to a single entity (e.g., “soap opera”, “New York”). Furthermore, when a user provides a further follow-up "clear" command, the automation assistant can determine that the further follow-up "clear" command is intended to remove only the preceding single word (e.g., "forward") from the field incorporated into the application.
[0013] In some implementations, a "clear" command can be used to delete more than one word or a sequence of more than one word referring to a single entity. In some of these implementations, determining which more than one word or a sequence of more than one word referring to a single entity to be deleted can be based on the characteristics of the words to be deleted and / or the amount of time between completing the utterance that led to the deletion of the words and saying the "clear" command. The characteristics of the words to be deleted can include, for example, their speech recognition confidence scores, whether they are in a speech recognition vocabulary and / or an applied vocabulary, and / or whether they are typical for the user (and optionally the context). As an example, if the speech recognition confidence scores of the last three words do not meet a threshold (but the words preceding these last three words have speech recognition confidence scores that meet the threshold), and the "clear" command is provided within 30 milliseconds of completing the utterance that led to the transcription of these three words, then the "clear" command can delete all three words. On the other hand, if the speech recognition confidence score of the last word fails to meet the threshold (but the words preceding the last word have speech recognition confidence scores that meet the threshold) and / or if the "clear" command is provided 500 milliseconds after the utterance is completed, the "clear" command can only eliminate the last word.
[0014] In various implementations, successive commands indicating the text content to be removed from a document can progressively result in more text content being removed. For example, suppose the text content includes three paragraphs dictated by a user in word-for-word mode within an email document. In this example, a first instance of the "clear" command could result in the removal of the most recently incorporated word (e.g., the last word in the third paragraph). Furthermore, a second instance of the "clear" command could result in the removal of the most recently incorporated phrase or sentence (e.g., the last sentence in the third paragraph). Furthermore, a third instance of the "clear" command could result in the removal of the most recently incorporated group of sentences or paragraphs (e.g., the third paragraph itself). Therefore, for each successive instance of the command, the length of the content to be removed from the document can increase.
[0015] Various technical advantages can be achieved by using the techniques described herein. As a non-limiting example, the techniques described herein enable automated assistants and / or applications accessible to them to format text content to be included in documents generated based on user speech, without requiring the user to explicitly specify how the text content should be formatted. As a result, the duration of human-computer dialogue between the user and the automated assistant can be reduced, the amount of user input including commands specifying how the text content should be formatted can be reduced, the number of times the user switches between interfaces can be reduced, and / or the amount of manual editing of the user and document after the human-computer dialogue can be reduced, thereby saving computational resources at the client device used in performing the human-computer dialogue, and / or saving network resources in implementations where the automated assistant and / or application are at least partially executed by one or more remote systems.
[0016] The above description serves as an overview of some implementations of this disclosure. Further descriptions of those and other implementations are provided below in more detail. Attached Figure Description
[0017] Figure 1A , Figure 1B , Figure 1C and Figure 1D The image shows a view of the spoken words provided by the user to the application, which can then generate text content data and layout data based on those spoken words.
[0018] Figure 2 A system is shown that provides an automated assistant and / or application that can arrange text content to facilitate voice-to-text operations without requiring the user to explicitly identify the arrangement operation.
[0019] Figure 3A method is shown for operating applications and / or automated assistants to incorporate text content into text fields based on an arrangement that may not be explicitly identified from user input.
[0020] Figure 4 This is a block diagram of an exemplary computer system. Detailed Implementation
[0021] Figure 1A , Figure 1B , Figure 1C and Figure 1D Multiple views of spoken utterances provided by user 102 to an application are shown, which can generate text content data and arrangement data based on the spoken utterances. The text content data and arrangement data can be processed to incorporate natural language content into text field 114 based on the arrangement data. In some implementations, the application can be an automation assistant application and / or any other application that optionally utilizes the capabilities of an automation assistant to receive spoken input from a user. Alternatively or additionally, the application (e.g., a first application) can interact with an additional application (e.g., a second application) to generate content for the additional application's text field 114. For example, user 102 can provide spoken utterances 106 such as “Hi Adam...Where will we be meeting today?...Please let me know when you can...Thank you...William”. After user 102 has initialized the application’s voice input features by invoking gestures or buttons (e.g., software or hardware buttons) and / or provided the application with an invoking command to invoke an automation assistant (e.g., “Hey, Assistant.”), spoken words 106 may optionally be provided.
[0022] In response to receiving spoken words 106 at computing device 104, computing device 104 and / or an application may generate audio data 108 corresponding to the spoken words 106 (e.g., via the microphone of computing device 104). In some implementations, the audio data 108 may be processed (e.g., using one or more heuristics and / or one or more trained machine learning models (e.g., NLU models)) to determine one or more intentions and / or actions to be performed in response to the spoken words 106. For example, one or more intentions determined based on the processing of audio data 108 may provide a basis for generating layout data 110, which may characterize one or more formatted actions that should be performed in response to the spoken words 106. One or more intentions may be additionally or alternatively determined based on the application in which an automation assistant is invoked. The audio data 108 may also be processed (e.g., using references) to determine one or more intentions. Figure 2 The described speech processing engine 208) provides a basis for generating text content data 112, which can characterize the text content in text field 114 in response to spoken utterances that should be incorporated into the application or an additional application.
[0023] Using text content data 112, such as Figure 1B As shown in the view, the application can identify the text content 122 in text field 114 to be incorporated into an additional application (e.g., an email application). The text content 122 may correspond to natural language content containing spoken words 106 provided by the user. Using layout data 110, such as... Figure 1C As shown in the view, the application can recognize placement commands 142 (e.g., carriage return, ANSI code, ASCII code, ISO code, HTML, JavaScript, etc.) to provide the application, additional applications, and / or computing device 104 with respect to the text content 122 incorporated into the text field 114 for execution and / or implementation. Executing such placement commands 142 can cause the placed text content 162 to be rendered in the text field 114, such as... Figure 1D The view is shown.
[0024] In some implementations, audio data 108 and / or any other data can be used as the basis for performing certain operations and / or intentions in order to render the arranged text content 162 in text field 114. For example, as Figure 1AAs shown, the duration between segments of audio data 108 corresponding to text content 122 can be used as the basis for arranging command 142. The duration between segments of audio data 108 can be determined using, for example, an endpoint machine learning model trained to detect the start of spoken utterance 106, pauses within spoken utterance 106, and / or the end of spoken utterance 106. For example, a command string (e.g., two carriage returns) can be incorporated into text content 162 arranged between the two segments of the text content. In some implementations, the command string can be identified based on the duration of a speech interruption 116 between a first segment 118A of the spoken input (e.g., spoken utterance 106) and a second segment 118B of the spoken input. In additional or alternative implementations, the command string can be identified based on one or more determined intentions associated with spoken utterance 106. When the command string includes one or more “newline” commands, “new bullet point” commands, and / or “new listitem” commands, the vertical position of the text corresponding to the first part 118A of the verbal input may differ from the different vertical position of the second part 118B of the verbal input.
[0025] Alternatively or additionally, one or more commands for arranging text content can be identified based on: the appended application of text field 114, the type of application providing text field 114, the type of document being drafted, the status of the appended application and / or computing device 104, and / or any other information available to the application. For example, when the appended application is identified as an email application and / or a word processing application, arrangement data 110 can be generated such that a comma and one or more carriage returns can be inserted between the phrase “Thank you” and the name “William”. Alternatively or additionally, when the appended application is identified as a text messaging application, arrangement data 110 can be generated such that a comma and a space separate the phrase “Thank you” and the name “William”. The comma can be inserted in either case without the user 102 explicitly stating the use of punctuation. For example, even if user 102 may explicitly state "Thank you William" in their spoken words 106 without explicitly saying the word "comma" or other formatting commands, the application can still generate layout data 110 and / or text content data 112 to insert a comma and a carriage return between the phrases "Thank you" and "William".
[0026] In some implementations, the generation of placement data 110 can be based on heuristic processes and / or one or more trained machine learning models. For example, one or more trained machine learning models can be used to process one or more portions of text content to generate low-dimensional representations of the text content, such as embeddings or word2vec representations of one or more terms or phrases included in the text content. These low-dimensional representations can be mapped to an embedding space or another latent space. When a mapped embedding is determined to have a threshold distance in the embedding space or latent space from one or more other previously mapped embeddings corresponding to other terms and / or phrases, placement operations and / or formatting instructions corresponding to one or more other previously mapped embeddings of other terms and / or phrases can be identified as placement data 110 and / or formatting instructions for those one or more portions of the text content. Notably, these low-dimensional representations can be specifically tailored to the intent associated with the type of application to which the spoken utterance and / or text content is incorporated. For example, if the text content is to be incorporated into an email, the first arrangement data for a greeting can be determined as arrangement data 110 for a greeting, and if the text content is to be incorporated into a text message, the second arrangement data for the same greeting can be determined as arrangement data 110 for a greeting.
[0027] In some versions of these implementations, training data for training one or more trained machine learning models can be generated, along with how user 102, with user 102's permission, arranges certain text content for each corresponding application and text content. For example, each instance of the training data may include training instance input and training instance output. For each of these training instances, the training instance input may include, for example, instructions for one or more text fragments and / or applications in which one or more text fragments are incorporated, and the training instance output may include corresponding arrangement operations and / or formatting instructions for those one or more portions of the text content included in the training instance input. The training instance input can be applied as input across one or more machine learning models to generate predicted outputs indicating predicted arrangement operations and / or formatting instructions for those one or more portions of the text content included in the training instance input. The predicted arrangement operations and / or formatting instructions for those one or more portions of the text content can be compared with the training instance output to generate one or more losses, and one or more machine learning models can be updated based on one or more of the losses. Therefore, during inference, when generating layout data 110, one or more trained machine learning models can be used to weight certain layout operations and / or formatting operations on the text portion over other layout operations. In this way, when executing layout data 110 in response to spoken utterance 106, the historical interactions involving user 102 can serve as the basis for applying layout text content in text field 114. Training of the machine learning model can include, for example, on-device training and personalization of the machine learning model (without leaving the instance of training data on the corresponding device) and / or can include training via a joint learning framework, where gradients are generated locally on the client device based on instances of training data, and only the gradients are sent to a remote server (without leaving the instance of training data on the corresponding device) for training the machine learning model. Instances of training data can include those instances with training instance inputs of text fragments typed (e.g., via a virtual or hardware keyboard), and / or those instances with training instance inputs of text fragments transcribed from spoken utterances. Instances of training data can include, for example, those instances with training instance outputs based on manual corrections of formatting via voice or other means (e.g., via a virtual or hardware keyboard). The output of training instances based on manual correction can indicate to the corresponding user that the final format is acceptable.
[0028] Figure 2A system 200 is shown that provides an automation assistant 204 and / or an application. This system 200 can arrange text content to facilitate voice-to-text operations without requiring the user to explicitly identify the arrangement. The automation assistant 204 may operate as part of an assistant application provided on one or more computing devices, such as computing device 202 and / or server device. A user can interact with the automation assistant 204 via an assistant interface 220, which may be a microphone, camera, touchscreen display, user interface, and / or any other device capable of providing an interface between the user and the application. For example, a user can cause the automation assistant 204 to initiate one or more actions (e.g., providing data, controlling peripheral devices, accessing agents, generating inputs and / or outputs, etc.) by providing verbal, textual, and / or graphical input to the assistant interface 220.
[0029] Alternatively, the automation assistant 204 may be initialized using one or more trained machine learning models based on the processing of context data 236. Context data 236 may characterize one or more features of the environment in which the automation assistant 204 can access, and / or one or more features of a user predicted to want to interact with the automation assistant 204. The computing device 202 may include a display device, which may be a display panel including a touch interface for receiving touch input and / or gestures to allow the user to control applications 234 of the computing device 202 via the touch interface. In some implementations, the computing device 202 may not have a display device, thus providing audible user interface output without providing graphical user interface output. Furthermore, the computing device 202 may provide a user interface such as a microphone for receiving spoken natural language input from the user. In some implementations, the computing device 202 may include a touch interface and may not have a camera, but may optionally include one or more other sensors.
[0030] Computing device 202 and / or other client devices can communicate with server devices via a wide area network (WAN) such as the Internet. Furthermore, computing device 202 and any other computing devices can communicate with each other via a local area network (LAN) such as a Wi-Fi network. Computing device 202 can offload computing tasks to server devices to conserve computing resources at computing device 202. For example, server devices can host automation assistant 204, and / or computing device 202 can transmit input received at one or more assistant interfaces 220 to server devices. However, in some implementations, automation assistant 204 can be hosted at computing device 202, and various processes associated with automation assistant operations can be executed locally at computing device 202.
[0031] In various implementations, all or fewer aspects of the automation assistant 204 may be implemented on the computing device 202. In some of these implementations, aspects of the automation assistant 204 are implemented via the computing device 202 and may interface with a server device that can implement other aspects of the automation assistant 204. The server device may optionally serve multiple users and their associated assistant applications via multiple threads. In implementations where all or fewer aspects of the automation assistant 204 are implemented via the computing device 202, the automation assistant 204 may be an application separate from the operating system of the computing device 202 (e.g., installed "on top" of the operating system) – or alternatively, may be implemented directly by the operating system of the computing device 202 (e.g., considered an application of the operating system, but integrated with it).
[0032] In some implementations, the automation assistant 204 may include an input processing engine 206, which may use multiple different modules to process the input and / or output of the computing device 202 and / or the server device. For example, the input processing engine 206 may include a speech processing engine 208, which can process audio data received at the assistant interface 220 to recognize the text content embodied in the audio data. The audio data may be transmitted from, for example, the computing device 202 to the server device to conserve computing resources at the computing device 202, and the text content may be received from the server device. Alternatively, the audio data may be specifically processed at the computing device 202 to recognize the text content.
[0033] The process of converting audio data into text content may include processing the audio data using a speech recognition algorithm that may use neural networks and / or statistical models to identify audio data sets corresponding to words or phrases. The text converted from the audio data may be parsed by a data profiling engine 210 and is available to the automation assistant 204 as text data that can be used to generate and / or identify command phrases, intents, actions, slot values, and / or any other user-specified content. In some implementations, the output data generated by the data profiling engine 210 may be provided to a parameter engine 212 to determine whether the user has provided input corresponding to a specific intent, action, and / or routine that can be performed by the automation assistant 204 and / or by an application or agent accessible via the automation assistant 204. For example, assistant data 238 may be stored at a server device and / or computing device 202 and may include data defining one or more actions that can be performed by the automation assistant 204, as well as parameters necessary to perform these actions. The parameter engine 212 may generate one or more parameters for intents, actions, and / or slot values and provide these parameters to an output generation engine 214. The output generation engine 214 can use one or more parameters to communicate with the assistant interface 220 to provide output to be presented to the user, and / or communicate with one or more applications 234 to provide output to be presented to the user via one or more applications 234.
[0034] In some implementations, the automation assistant 204 may be an application that can be installed "on top" of the operating system of the computing device 202 and / or may form part (or all) of the operating system of the computing device 202. The automation assistant application includes and / or has access to on-device speech recognition, on-device NLU, and on-device execution. For example, on-device speech recognition can be performed using an on-device speech recognition module that processes audio data (detected by a microphone) using an end-to-end speech recognition machine learning model locally stored at the computing device 202. On-device speech recognition generates recognized text content for spoken utterances (if any) present in the audio data. Furthermore, on-device NLU can be performed, for example, using an on-device NLU module (or data profiling engine 210) that processes the recognized text generated using on-device speech recognition, along with optional context data, to generate NLU data.
[0035] NLU data may include an intent corresponding to a spoken utterance and optional parameters (e.g., slot values) for that intent. On-device fulfillment can be performed using an on-device fulfillment module that uses NLU data (from on-device NLU) and optionally other local data to determine the action to be taken to resolve the intent of the spoken utterance (and optionally the parameters of that intent). This may include determining local and / or remote responses to the spoken utterance (e.g., answers), interactions with locally installed applications performed based on the spoken utterance, commands transmitted to Internet of Things (IoT) devices (directly or via corresponding remote systems) based on the spoken utterance, and / or other parsed actions performed based on the spoken utterance. On-device fulfillment can then initiate local and / or remote execution / enforcement of the determined actions to resolve the spoken utterance.
[0036] In various implementations, remote speech processing, remote NLU, and / or remote execution can be utilized at least selectively. For example, identified text content can be selectively transmitted to a remote automation assistant component for remote NLU and / or remote execution. For instance, identified text content can optionally be transmitted for remote execution in parallel with on-device execution, or for remote execution in response to failures of on-device NLU and / or on-device execution. However, on-device speech processing, on-device NLU, on-device execution, and / or on-device execution can be prioritized, at least due to the reduced latency they provide when parsing spoken utterance (since no client-server round trip is required to parse spoken utterance). Furthermore, in the absence of or with limited network connectivity, on-device functionality may be the only available functionality.
[0037] In some implementations, computing device 202 may include one or more applications 234, which may be provided by a third-party entity different from the entity providing computing device 202 and / or automation assistant 204. The application state engine of automation assistant 204 and / or computing device 202 may access application data 230 to determine one or more actions that can be performed by the one or more applications 234, and the state of each of the one or more applications 234 and / or the state of the corresponding device associated with computing device 202. The device state engine of automation assistant 204 and / or computing device 202 may access device data 232 to determine one or more actions that can be performed by computing device 202 and / or one or more devices associated with computing device 202. In addition, application data 230 and / or any other data (e.g., device data 232) can be accessed by the automation assistant 204 to generate context data 236, which can characterize the context in which a particular application 234 and / or device is being executed, and / or the context in which a particular user is accessing computing device 202, accessing application 234 and / or any other device or module.
[0038] When one or more applications 234 are executing at computing device 202, device data 232 can characterize the current operating state of each application 234 executing at computing device 202. Furthermore, application data 230 can characterize one or more features of the executing application 234, such as the content of one or more graphical user interfaces rendered under the guidance of one or more applications 234. Alternatively or additionally, application data 230 can characterize action patterns, which can be updated by the respective application and / or automation assistant 204 based on the current operating state of the respective application. Alternatively or additionally, one or more action patterns of one or more applications 234 can remain static, but can be accessed by the application state engine to determine appropriate actions to be initiated via automation assistant 204.
[0039] The computing device 202 may also include an assistant invocation engine 222, which can use one or more trained machine learning models to process application data 230, device data 232, context data 236, and / or any other data accessible to the computing device 202. The assistant invocation engine 222 can process this data to determine whether to wait for the user to explicitly utter the invocation phrase to invoke the automation assistant 204, or to interpret the data as an intent to invoke the automation assistant—instead of requiring the user to explicitly utter the invocation phrase. For example, instances of training data can be used to train one or more trained machine learning models based on scenarios where the user is in various operating states across multiple devices and / or applications.
[0040] Instances of training data can be generated to capture training data characterizing scenarios where the user invokes the automation assistant and other scenarios where the user does not invoke the automation assistant. For example, when a user is accessing a specific word processing application, text messaging application, email application, social media application, task application, calendar application, reminder application, etc., the user can invoke their automation assistant to perform voice-to-text operations. Therefore, when the user opens one or more of these applications, one or more trained machine learning models can assist in indicating that the user intends to invoke their automation assistant. In this way, the user can provide spoken words to the automation assistant to perform voice-to-text operations without having to provide an explicit invocation command. When training one or more trained machine learning models based on these instances of training data, the assistant invocation engine 222 can enable the automation assistant 204 to detect or restrict the detection of spoken invocation phrases from the user based on features of the scenario and / or environment of the computing device 202. Alternatively, the assistant invocation engine 222 can enable the automation assistant 204 to detect or restrict the detection of one or more assistant commands from the user based on features of the scenario and / or environment of the computing device 202.
[0041] In some implementations, system 200 may include an application detection engine 216, which may optionally be used to detect the type of application that a user may be using to perform voice-to-text operations at computing device 202. Automation assistant 204 may use data available to system 200 to detect the type of application in order to determine the text layout that should be implemented for voice-to-text operations. For example, when the type of application is determined to be an email application, system 200's text layout engine 226 may generate carriage return and tab data for arranging text in alphabetical format. Alternatively or additionally, when the type of application is determined to be a task application, system 200's text layout engine 226 may generate emphasis data for arranging text in an emphasis list format.
[0042] In some implementations, system 200 may include a text content engine 218 and a text placement engine 226 for determining how text content should be placed in one or more fields of one or more applications in response to spoken utterances from a user. In some implementations, vocalization characteristics of the spoken utterances may be identified to determine how the text content should be placed within the application's fields. In additional or alternative implementations, intent associated with the spoken utterances may be utilized to determine how the text content should be placed within the application's fields. The intent associated with the spoken utterances may be determined based on processing of the spoken utterances (e.g., using the data profiling engine 210 as described above) and / or based on the type of application being used when the automation assistant is invoked. Alternatively or additionally, vocalization characteristics of the spoken utterances may be identified for use by text elimination engine 224 to determine the amount of text content to be eliminated from the application's fields.
[0043] In some implementations, the text content engine 218 and the text placement engine 226 can identify text content and placement commands for use in fields of the application based on vocalization features. For example, the intonation characteristics of a particular word or phrase spoken by the user can indicate how that particular word or phrase should be placed within the field. In some implementations, the amount of time between the letters spoken (e.g., determined using an endpoint machine learning model described with reference to Figure 1) can indicate that the user wants to pronounce the acronym. Therefore, even if the user has not explicitly identified the punctuation or other placement of the acronym, the automation assistant 204 can use the text placement engine 226 to identify certain punctuation and / or placement data based on intonation characteristics. Subsequently, if the user provides a command to delete a certain amount of text content, the text elimination engine 224 can identify the amount of text to be eliminated based at least on these vocalization features of one or more previously provided spoken utterances. For example, a user who provides the command "clear" after uttering the acronym can cause the automation assistant 204 to eliminate individual characters, punctuation marks, and / or spaces from the field. However, if the user provides the command "clear" after saying a word (e.g., "Thank…"), the automation assistant 204 can choose the word to be cleared instead of a single character.
[0044] In some implementations, a user can provide the automation assistant 204 with a series of identical or similar commands to remove text content from fields in the application. In response, the automation assistant 204 can remove text content segments of varying lengths for each corresponding command. For example, when a user drafts a letter by providing spoken words to the automation assistant 204, the user can provide the "delete" command via spoken words to remove text content from the letter. In response, the text removal engine 224 can identify the length of a first segment of text to be removed from the field and remove that segment. If the user provides another instance of the delete command, the text removal engine 224 can identify a second segment of text to be removed from the field, and this second segment may be longer than the first segment. Furthermore, if the user provides yet another instance of the delete command via spoken words, the text removal engine 224 can identify a third segment of text to be removed from the field, and this third segment may be longer than the first and second segments. In some implementations, the third segment may be longer than the combined lengths of the first and second segments.
[0045] Figure 3 A method 300 is illustrated for operating an application and / or an automation assistant to incorporate text content into a text field based on an arrangement that may not be explicitly identified from user input. Method 300 can be performed by one or more applications, devices, and / or any other means or modules capable of performing speech-to-text operations. Method 300 may include operation 302 determining whether a request to incorporate text content into the text field has been received. This request may be embodied in spoken words provided by the user to an application and / or automation assistant accessible via a computing device. In some implementations, the request may be provided while the user is accessing an additional application rendering the text field on a display interface of the computing device. In other implementations, the request may be provided without the user accessing any additional application. In other words, the request may be initiated to provide text content to be included in the text field (e.g., initiating an email, text message, to-do list, etc.). When it is determined that the user has provided a request to incorporate text content into the text field, method 300 may proceed from operation 302 to operation 304.
[0046] For example, an add-on app could be a note-taking app that users frequently use to create shopping lists, to-do lists, and other reminders. When the add-on app is open on a computing device, the user can invoke it to perform voice-to-text by providing a call phrase such as “Assistant.” Following this call phrase, the user can provide the items to be listed in the text field, such as “Three gallons of water, a roll of aluminum foil, and salt in a box.” As another example, the user could simply provide the spoken phrase “start a new to-do list,” which, when detected, causes the automation assistant to create a new list. Furthermore, when the user adds a new entry to a to-do list, each entry can begin on a separate line or at a separate emphasis mark (and optionally, the user does not need to explicitly provide a spoken command to do so, such as “next”).
[0047] Operation 304 may include generating text content data for incorporating the text content into a field of the application. The text content may be generated based on processing audio data corresponding to spoken utterances from the user (e.g., using...). Figure 2 (Speech processing engine 208). Text content data can represent natural language content to be incorporated into fields of the application. In some implementations, one or more trained machine learning models can be used to perform audio data processing to identify words and / or phrases that should be grouped together within fields of the application and / or omitted from fields of the application. Once the text content has been identified from the text content data, method 300 can proceed from operation 304 to operation 306.
[0048] Operation 306 may include determining whether available data in the application provides a basis for arranging text content within a field of the attached application. When data providing a basis for arranging the text content is available, method 300 may proceed from operation 306 to operation 308. Otherwise, method 300 may proceed from operation 306 to operation 310. Operation 310 may include incorporating the text content into a text field of the attached application, and the operation may proceed to operation 314. Operation 314 is described below. Operation 308 may include generating content arrangement data for the text content. The content arrangement data may characterize one or more operations and / or commands to be performed to arrange text content within a text field according to the user's predicted intent.
[0049] For example, in some implementations, data available to the application can indicate that the application is a note-taking application that the user has previously used to create lists. Based on this determination, the application and / or automation assistant can process audio data and / or text content data to identify the location of placement data inserted within the text content. For example, audio data and / or text content data can be processed to identify the first part of the text content that should be separated from the second part of the text content by one or more placement operations (e.g., one or more carriage returns). In some implementations, one or more heuristic processes and / or one or more trained machine learning models can be used to process audio data and / or text content to identify one or more placement operations and / or where the placement operation is performed within the text content.
[0050] For example, training data can be used to train one or more trained machine learning models that characterize one or more other documents generated using a note-taking app and / or one or more other apps that can be used for note-taking. In this way, apps and / or automated assistants can be trained to identify the arrangement of text content in different documents based on their corresponding text content, the type of app, the type of user (with prior permission from the user), the time the document was created, the location where the document was created, and / or any other information that may provide a basis for arranging the text content in a certain way. For example, users typically access their "to-do" lists at home in the morning, so voice-to-text operations performed at the user's home in the morning can be incorporated into the arrangement data used to create the "list," in contrast to the arrangement of creating formal letters and / or log entries.
[0051] Alternatively or additionally, users may typically use their automation assistants to generate log entries on their way home from get off work in the evening. Therefore, when a user is identified as being on their way home from get off work and requests the automation assistant to perform a voice-to-text operation, the automation assistant can identify one or more placement operations. These placement operations can be identified based on previous instances of log entries generated by the user and / or one or more other users, to be incorporated into the text content of the log entries (e.g., indentation, newlines, date, signature, heading, font, color, size, etc.). For example, one or more trained machine learning models can be used to process one or more portions of the text content and / or contextual data to generate embeddings that can be mapped to a latent space. When the mapped embeddings are determined to be within a threshold distance from one or more other embeddings corresponding to placement operations and / or formatting instructions, the placement operations and / or formatting instructions can be performed relative to those one or more portions of the text content.
[0052] Method 300 can proceed from operation 308 to operation 312, whereby operation 312 may include arranging data according to content so that text content is incorporated into the text field. In some implementations, when text content is incorporated into the text field, method 300 may proceed to optional operation 314 to determine whether the user has provided a request to remove a certain amount of text content from the text field. Optional operation 314 can be considered optional because its execution can be based on receiving one or more specific verbal utterances from the user. If it is determined that the user has not provided such a request, method 300 may return to operation 302. Otherwise, method 300 may proceed to optional operation 316 to remove a specific amount of text content from the text field. Similar to optional operation 314, optional operation 316 can be considered optional because its execution can be based on receiving one or more specific verbal utterances from the user in optional operation 314.
[0053] In some implementations, the specific amount of text content to be eliminated can be based on one or more prior inputs from the user to the application and / or automation assistant. Alternatively or additionally, the amount of text content to be eliminated can be based on the content of the text field, the additional application providing the text field, the application type corresponding to the additional application, the arrangement of the content within the text field, and / or any other information that may be associated with the text content. For example, a request to eliminate a certain amount of text content can be embodied in additional spoken words such as "Clear".
[0054] In response to receiving additional spoken words, the automation assistant may determine that one or more recent interactions between the user and the automation assistant include the user instructing the automation assistant to incorporate a list of highlighted items into the text field. Based on this determination, the automation assistant may, in response to the additional spoken words, remove a single list item or a number of list items from the list of highlighted items. Alternatively or additionally, the automation assistant may determine that one or more recent interactions between the user and the automation assistant include the user instructing the automation assistant to incorporate acronyms into the text field. Based on this determination, the automation assistant may, in response to additional spoken words from the user, remove a single character or a number of characters from the acronym.
[0055] Figure 4This is a block diagram of an exemplary computer system 410. Computer system 410 typically includes at least one processor 414 that communicates with a plurality of peripheral devices via a bus subsystem 412. These peripheral devices may include a storage subsystem 424 (including, for example, memory 425 and file storage subsystem 426), a user interface output device 420, a user interface input device 422, and a network interface subsystem 416. The input and output devices allow users to interact with computer system 410. The network interface subsystem 416 provides an interface to an external network and is coupled to corresponding interface devices in other computer systems.
[0056] User interface input device 422 may include a keyboard, pointing devices (such as a mouse, trackball, touchpad, or graphics tablet), scanner, touchscreen integrated into a display, audio input devices (such as a voice recognition system, microphone, and / or other types of input devices). Generally, the term "input device" is intended to include all possible types of devices and methods for inputting information into computer system 410 or a communication network.
[0057] User interface output device 420 may include a display subsystem, a printer, a fax machine, or a non-visual display (such as an audio output device). The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating visual images. The display subsystem may also provide a non-visual display, such as via an audio output device. Generally, the term "output device" is intended to encompass all possible types of devices and methods for outputting information from computer system 410 to a user or to another machine or computer system.
[0058] Storage subsystem 424 stores the functional programming and data structures of some or all of the modules described herein. For example, storage subsystem 424 may include selected aspects of performing method 300 and / or implementing the logic of one or more of system 200, computing device 104, automation assistant and / or any other application, device, apparatus and / or module discussed herein.
[0059] These software modules are typically executed by processor 414 alone or in conjunction with other processors. The memory 425 used in storage subsystem 424 may include multiple memories, including main random access memory (RAM) 430 for storing instructions and data during program execution and read-only memory (ROM) 432 for storing fixed instructions. File storage subsystem 426 can provide persistent storage for program and data files and may include hard disk drives, floppy disk drives, and associated removable media, CD-ROM drives, optical drives, or removable media cartridges. Modules implementing certain functionalities may be stored by file storage subsystem 426 within storage subsystem 424 or in other machines accessible to processor 414.
[0060] Bus subsystem 412 provides a mechanism for enabling the various components and subsystems of computer system 410 to communicate with each other as intended. Although bus subsystem 412 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0061] Computer systems 410 can be of different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing systems or computing devices. Due to the constantly evolving nature of computers and networks, Figure 4 The description of the computer system 410 depicted herein is intended only as a specific example for illustrating some implementation methods. Many other configurations of the computer system 410 may have... Figure 4 The computer system described in the document has more or fewer components.
[0062] In situations where the system described herein collects or can utilize personal information about a user (or, as often referred to herein, a “participant”), the user may be given the opportunity to control whether the program or feature collects user information (e.g., information about the user’s social networks, social actions or activities, occupation, user preferences, or the user’s current geolocation), or to control whether and / or how content that may be more relevant to the user is received from the content server. Furthermore, certain data may be processed in one or more ways before storage or use to remove personally identifiable information. For example, a user’s identity may be processed to the point that personally identifiable information about the user cannot be determined, or the user’s geolocation may be generalized (e.g., to the city, zip code, or state level) where geolocation information is obtained, making it impossible to determine the user’s specific geolocation. Therefore, the user can control how information about them is collected and / or used.
[0063] While several implementations have been described and illustrated herein, various other components and / or structures may be utilized for performing functions and / or obtaining results and / or one or more advantages described herein, and each of such variations and / or modifications is considered to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be illustrative, and actual parameters, dimensions, materials, and / or configurations will depend on the specific application using the teachings. Those skilled in the art will recognize, or be able to determine, many equivalents of the specific implementations described herein using only conventional experimentation. It will therefore be understood that the foregoing implementations are presented by way of example only, and that implementations may be practiced in ways different from those specifically described and claimed within the scope of the appended claims and their equivalents. Implementations of this disclosure relate to each other feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods that are not contradictory is included within the scope of this disclosure.
[0064] In some implementations, a method implemented by one or more processors is provided, the method including receiving, at a computing device, a spoken utterance from a user directed to a first application. The spoken utterance corresponds to a request for the first application to perform a speech-to-text operation to incorporate text into a field of a second application different from the first application. The method also includes generating text content data based on the spoken utterance, the text content data representing text content to be incorporated into the field of the second application. The method further includes generating content layout data based on an intent associated with the spoken utterance, the content layout data representing the arrangement of a first portion of the text content relative to a second portion of the text content within the field of the second application. The method also includes, in response to the spoken utterance and based on the text content data and the content layout data, incorporating the text content into the field of the second application according to the arrangement.
[0065] These and other implementations of the techniques disclosed herein may include one or more of the following features.
[0066] In some implementations, generating content layout data includes determining the duration between a first spoken portion and a second spoken portion of the spoken utterance, and determining the layout of the first portion of text content relative to the second portion of text content based on that duration. In some versions of those implementations, this layout includes the vertical position of the first portion of text content relative to the second portion of text content within a field of the second application. In some versions of these versions, the arrangement such that the text content is incorporated into the field of the second application includes determining the vertical position of the incorporated first portion of text content relative to the second portion of text content within the field of the second application. As an example, the vertical position of the incorporated first portion could include incorporating a carriage return after the first portion of text content into the field of the second application, and incorporating the second portion of text content after the carriage return.
[0067] In some implementations, generating text content data involves identifying one or more punctuation marks to include in the text content of a field to be incorporated into a second application. In some of these implementations, spoken utterances do not explicitly identify punctuation marks to be incorporated into a field of the second application.
[0068] In some implementations, the text content data represents the natural language content embodied in spoken discourse, and the content layout data represents a formatting command that, when executed by a second application, causes the second application to separate the first part of the text content from the second part of the text content within the fields of the second application.
[0069] In some implementations, the method further includes receiving additional spoken words from a user at a computing device, directed to a first application. The additional spoken words correspond to an additional request directed to the first application to perform additional speech-to-text operations to incorporate additional text content into a field of a second application. In some versions of these implementations, the method further includes: in response to the additional spoken words, causing the second application to incorporate the additional text content into a field of the second application; and in response to the additional spoken words, causing the second application to perform one or more formatting operations that modify the additional text content relative to another arrangement of text content within the field. The one or more formatting operations are not explicitly identified by the user via the additional spoken words. In some versions of these implementations, the first application is an automation assistant and the second application is a word processing application, and optionally, the method further includes identifying one or more formatting operations based on the additional spoken words and the text content incorporated into the field of the second application.
[0070] In some implementations, a method implemented by one or more processors is provided, the method including receiving a first spoken utterance at a computing device, the first spoken utterance corresponding to a request for a first application to perform a speech-to-text operation for a user. The method also includes rendering text content within a field of a second application based on the first spoken utterance. The text content includes natural language content of the first spoken utterance. The method further includes receiving a second spoken utterance from the user, the second spoken utterance corresponding to an additional request from the first application to remove a portion of the text content from the field of the second application, but the second spoken utterance does not explicitly identify the portion of the text content to be removed. The method also includes, in response to the second spoken utterance, determining the amount of content to be removed from the text content rendered within the field of the second application. The method also includes, in response to the second spoken utterance, causing the amount of content to be removed from the text content rendered within the field of the second application.
[0071] These and other implementations of the techniques disclosed herein may include one or more of the following features.
[0072] In some implementations, the amount of content to be eliminated is based on the vocal characteristics exhibited by the user when delivering at least a portion of the first spoken utterance. In some versions of those implementations, the vocal characteristics include the duration between discrete segments of the first spoken utterance, and the discrete segments of the first spoken utterance describe different corresponding parts of the text content. In some additional or alternative versions of these implementations, the vocal characteristics include intonation characteristics embodied in the first spoken utterance, and the discrete segments of the first spoken utterance describe different corresponding parts of the text content.
[0073] In some implementations, determining the amount of content to be removed from the text content includes: determining that the vocal characteristics of the first spoken utterance include the explicit pronunciation of individual natural language characters, and determining the number of individual natural language characters to be removed from the text content rendered within a field of the second application. The text content rendered within a field of the second application includes individual natural language characters, and the amount of content to be removed from the text content corresponds to the number of individual natural language characters.
[0074] In some implementations, determining the amount of content to be removed from the text content includes identifying the length of a first segment of the text content. In those implementations, the first segment of the text is a first portion of the text content that was recently incorporated into a field of the application. In some versions of these implementations, the method further includes receiving an additional instance of a second spoken utterance at a computing device for removing the additional portion of the text content from the field of the application and determining the amount of additional content to be removed from the text content. The additional content amount includes a second segment of the text content that has a longer length than the first segment of the text content. In some versions of these implementations, the method further includes receiving an additional instance of a second spoken utterance at a computing device for removing an additional portion of the text content from the field of the application and determining an additional amount of content to be removed from the text content. The additional content amount includes a third segment of the text content with a discrete length that is longer than both the first and second segments of the text content. As an example, the first segment of the text content may consist of words, the second segment may consist of sentences, and the third segment may consist of paragraphs containing multiple sentences.
[0075] In some implementations, a method implemented by one or more processors is provided, the method including receiving, at a computing device, a spoken utterance from a user directed to a first application. The spoken utterance corresponds to a request for the first application to perform a speech-to-text operation—incorporating text into a field of a second application different from the first application. The method also includes generating text content data based on the spoken utterance, the text content data representing text content to be incorporated into the field of the second application. The method further includes generating content layout data based on the application type of the second application, the content layout data representing the layout of a first portion of the text content relative to a second portion of the text content within the field of the second application. The method also includes, in response to the spoken utterance and based on the text content data and the content layout data, incorporating the text content into the field of the second application according to the layout.
[0076] These and other implementations of the techniques disclosed herein may include one or more of the following features.
[0077] In some implementations, content placement data is further generated based on one or more prior interactions, which involve the user providing additional text content to an application type corresponding to the second application.
[0078] Other implementations may include a non-transitory computer-readable storage medium containing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform methods such as those described above and / or elsewhere herein. Other implementations may include a system of one or more computers including one or more processors operable to execute the stored instructions to perform methods such as those described above and / or elsewhere herein.
[0079] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are considered part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered part of the subject matter disclosed herein.
Claims
1. A method for arranging text content, the method comprising: Receive verbal commands from the user directing them to the first application at the computing device. The spoken words correspond to a request for the first application to perform a speech-to-text operation to incorporate the text into a field of the second application. Text content data is generated based on the spoken utterance, and this text content data represents the text content to be incorporated into the field of the second application. The second application differs from the first application; Content layout data is generated based on the intent associated with the spoken utterance, the content layout data representing the arrangement of a first portion of the text content relative to a second portion of the text content within the field of the second application; Based on the text content data and the content layout data, in response to the spoken utterance, the text content is incorporated into the field of the second application according to the layout; The computing device receives additional spoken words from the user pointing to the first application. Wherein, the additional spoken utterance corresponds to an additional request for the first application to perform additional speech-to-text operations to incorporate additional text content into the field of the second application; In response to the additional spoken utterance, the second application incorporates the additional text content into the field of the second application; and In response to the additional spoken utterance, the second application performs one or more formatting operations that modify another arrangement of the additional text content within the field relative to the original text content. In this case, one or more formatting operations are not explicitly identified by the user via the additional verbal utterances.
2. The method according to claim 1, wherein, Generating the content layout data includes: Determine the duration between the first spoken portion and the second spoken portion of the spoken utterance. The arrangement of the first part of the text content relative to the second part of the text content is based on the duration.
3. The method according to claim 2, wherein, The arrangement includes the vertical position of a first portion of the text content relative to a second portion of the text content in the field of the second application.
4. The method according to claim 3, wherein, The arrangement such that the text content is incorporated into the field of the second application includes: At the field of the second application, the first portion of the text content is positioned vertically relative to the second portion of the text content.
5. The method according to claim 4, wherein, The vertical position incorporated into the first part includes: After the first part of the text content, the carriage return data is incorporated into the field of the second application, and The second part of the text content is incorporated after the carriage return data.
6. The method according to claim 1, wherein, Generating the text content data includes: Identify one or more punctuation marks to include in the text content of the field to be incorporated into the second application. The spoken utterances do not explicitly identify punctuation marks to be incorporated into the field of the second application.
7. The method according to any of the preceding claims, in, The text content data represents the natural language content embodied in the spoken discourse, and The content layout data represents a formatting command, which, when executed by the second application, causes the second application to separate the first part of the text content from the second part of the text content within the field of the second application.
8. The method according to claim 1, wherein, The first application is an automation assistant and the second application is a word processing application, and the method further includes: The one or more formatting operations are identified based on the additional spoken words and the text content incorporated into the fields of the second application.
9. A method for rendering text content, the method comprising: Receive a first spoken utterance at the computing device, the first spoken utterance corresponding to a request for a first application to perform a voice-to-text operation for a user; The text content is rendered within a field of the second application based on the first spoken utterance. The text content includes the natural language content of the first spoken utterance; The user receives a second spoken utterance, which corresponds to an additional request for the first application to remove a portion of the text content from the field of the second application. The additional request did not explicitly identify the text content of the portion to be removed; In response to the second spoken utterance, determine the amount of content to be removed from the text content rendered within the field of the second application, wherein determining the amount of content to be removed from the text content includes: The length of the first segment of the text content is identified. The first fragment of the text is the first part of the text content that was recently incorporated into the field of the second application; In response to the second spoken utterance, the content volume is removed from the text content rendered within the field of the second application; Receive additional instances of the second spoken utterance at the computing device for removing additional portions of the text content from the field of the second application, and Determine the amount of additional content to be removed from the text content. The additional content includes a second segment of the text content, the second segment having a longer length than the first segment of the text content; and In response to the additional instance of the second spoken utterance, the additional content is removed from the text content rendered within the field of the second application.
10. The method according to claim 9, wherein, The amount of content to be eliminated is based on the vocal characteristics exhibited by the user when providing at least a portion of the first spoken utterance.
11. The method according to claim 10, in, The vocal characteristics include the duration between the separate portions of the first spoken utterance, and The separate portions of the first spoken utterance describe different corresponding parts of the text content.
12. The method according to claim 10, in, The vocal characteristics include the intonation features manifested in the first spoken utterance.
13. The method according to any one of claims 9 to 12, wherein, Determining the amount of content to be removed from the text content includes: Determining the vocal features of the first spoken utterance includes the explicit pronunciation of individual natural language characters. Wherein, the text content rendered within the field of the second application includes the individual natural language characters; and Determine the number of individual natural language characters to be removed from the text content rendered within the field of the second application. The amount of content to be removed from the text content corresponds to the number of individual natural language characters.
14. The method of claim 9, further comprising: Receive another instance of the second spoken utterance at the computing device for use in removing further portions of the text content from the field of the application, and Determine the amount of additional content to be removed from the text content. The additional content includes a third segment of the text content with a discrete length, the discrete length being longer than the first segment of the text content and longer than the second segment of the text content.
15. The method according to claim 14, wherein, The first segment of the text content includes words, the second segment of the text content includes sentences, and the third segment of the text content includes paragraphs containing multiple sentences.
16. A system comprising: At least one processor; as well as A memory for storing instructions, which, when executed, cause the at least one processor to perform an operation corresponding to any one of claims 1 to 15.
17. A non-transitory computer-readable storage medium storing instructions, which, when executed, cause at least one processor to perform an operation corresponding to any one of claims 1 to 15.
Citation Information
Patent Citations
Training punctuation models
US9135231B1
Method for entering digit sequences by voice command
WO1989004035A1