Document creation and editing via automated assistant interaction
By leveraging semantic annotations and machine learning models through automated assistants, users can edit and comment on documents through verbal interaction, solving the resource waste and synchronization delay problems in editing and reviewing documents in existing technologies, and achieving efficient and flexible document processing.
Patent Information
- Application Number
- CN202080101564.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-09-08
- Filing Date
- 2020-12-14
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2040-12-14
AI Technical Summary
Existing technologies require a large amount of user input and device resources when editing and reviewing content-rich documents, and it is difficult to achieve real-time synchronization of multiple users editing and commenting at the same time, resulting in delays and waste of resources.
By leveraging semantic annotations and machine learning models, automated assistants allow users to edit, comment, and share documents through verbal and off-device interactions. Automated assistants process user verbal utterances to perform document-related tasks, reducing reliance on graphical user interfaces.
It achieves high efficiency in document creation and review, reduces device power consumption and user input, supports multi-user real-time editing and commenting, and improves the flexibility and synchronization of document processing.
Smart Images

Figure CN115769220B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to document creation and editing. More particularly, the present disclosure relates to document creation and editing via an automated assistant that allows users to create, edit, and / or share documents without directly interacting with a document editing application. Background Art
[0002] A human may engage in a human-computer conversation with an interactive software application, which is referred to herein as an "automated assistant" (also referred to as a "digital agent," "chatbot," "interactive personal assistant," "intelligent personal assistant," "conversational agent," etc.). For example, a human (who may be referred to as a "user" when they interact with an automated assistant) may provide commands and / or requests using spoken natural language input (i.e., utterances) and / or by providing textual (e.g., typed) natural language input, which in some cases may be converted to text and then processed.
[0003] In some instances, automated assistants can be used to perform discrete actions, such as opening music applications, adjusting settings for smart home devices, and many other tasks. However, editing content-rich documents (e.g., articles to be published) typically still retains a desktop environment with a dedicated monitor and common peripherals such as a keyboard and mouse. Although many tablet-style devices have enabled other means for editing documents, such as via a touch screen interface, it may be necessary to dedicate all of their dexterity to each editing session in order to edit content-rich documents. For example, adding a paragraph of text to a specific document stored on a cloud drive may require the user to: access the specific document via the foreground application of the tablet device, scroll to the specific paragraph to be edited, and manually type to edit the paragraph. This can require a large amount of user input and a large amount of client device resources to process the input in order to render the specific document over an extended duration, etc. Additionally, any other tasks performed via the tablet device may be delayed because the user cannot engage with any other applications during this time.
[0004] In addition, various document editing applications that exist as cloud applications can allow multiple different users to edit a document simultaneously through a desktop-style interface. Such applications can allow multiple reviewers to review a document simultaneously via a desktop-style interface. However, as long as editing and commenting are restricted to certain application interfaces, the review cycle may be excessively delayed. For example, a reviewer may receive an email notification on their phone that another user has added a comment to a document. Unfortunately, the user may not be able to fully process the comment until they have access to a desktop computer or other device with an appropriate graphical user interface. Furthermore, as a result, the user may not be informed of the substance of the comment and therefore cannot prepare a response to the comment before accessing the comment. These restrictions can cause various users to check for document comment updates from interfaces that may not enable editing. This can lead to unnecessary consumption of computing resources such as power and processing bandwidth. Summary of the Invention
[0005] Embodiments described herein relate to an automated assistant that can operate as a modality for performing various document-related actions for content-rich documents. A content-rich document can refer to any data set incorporated into a single document. The data set can include, but is not limited to, any combination of multiple different sections, topics, subtopics, styles, cells in a spreadsheet, slides in a presentation, graphics, and / or features that can be incorporated into a document. The automated assistant can operate to allow a user to edit, comment on, and / or share an existing document, or create a new document, through one or more interactions between the automated assistant and the user. In other words, the user does not need to be viewing a document editing program in the foreground of a graphical user interface (GUI) in order to perform such operations. Instead, the automated assistant can, for example, allow the user to perform various document-related tasks through verbal interaction and / or any other type of interaction—optionally without requiring the user to view the document when providing the verbal interaction. Such document-related tasks can be implemented by allowing the automated assistant to generate semantic annotations of corresponding portions of a single document that the user can request the automated assistant to access and / or modify. For example, a document-related task to be performed (e.g., to determine which portion of the document the document-related task should be performed on) can be determined based on processing at least a portion of the user's spoken utterance in view of the semantic annotation of the document. Referencing semantic annotations in this manner can simplify document creation and / or document review, which might otherwise require prolonged graphical rendering of the document and / or direct user interaction with a document editing application, such as one accessible via a desktop computing device. Furthermore, document review time and device power consumption can be reduced when a document reviewer is able to quickly review a content-rich document via any device that provides access to an automated assistant. Such devices may include, but are not limited to, watches, cell phones, tablet computers, home assistant devices, and / or any other computing device that can provide access to an automated assistant.
[0006] As an example, a user may be a researcher working with a team of researchers to review an electronic document to be submitted for publication. During the review process, each researcher may be working on a schedule that does not allow for much time to sit in front of a computing device to review edits and / or comments on the document. To edit and / or review comments on the document, the user may rely on an automated assistant that is accessible through the "ecosystem" of the user's device. For example, while the researcher is reviewing the document, the document application providing access to the document may send and receive certain data associated with the document to and from the automated assistant.
[0007] In some instances, the document may be a spreadsheet, and when a user provides a spoken utterance such as “Assistant, add a column to my latest 'research' document and add a comment saying 'Could someone add this month's data to this column?'” specific edits to the document may be caused. The user may provide such a spoken utterance to an interface of their watch, which may provide access to the automated assistant but may not include a native document editing application for editing the spreadsheet. In response to receiving the spoken utterance, the automated assistant may process audio data corresponding to the spoken utterance and determine one or more actions to perform.
[0008] The processing of the audio data may involve utilizing one or more trained machine learning models, heuristic processes, semantic analysis, and / or any other process that may be employed when processing a spoken utterance from a user. As a result of the processing, the automated assistant may initiate execution of one or more actions specified by the user via the spoken utterance. In some embodiments, the automated assistant may utilize an application programming interface (API) to cause a particular document application to perform one or more actions. For example, in response to the spoken utterance described above, an automated assistant accessible via the user's watch may generate one or more functions that are executed in response to the spoken utterance from the user. For example, the automated assistant may cause execution of one or more functions to determine where an additional column should be added to a particular document. The function used to determine where to place the additional column may be total_columns(most_recent('research')), which may identify the total number of non-blank columns included in the most recently accessed document that has a semantic annotation with the term "research." In some embodiments, the "total_columns" function may be identified based on one or more previous interactions between the user and the document editing application and / or the automated assistant. Alternatively or additionally, the “total_columns” function may be identified using one or more trained machine learning models, which may be used to rank and / or score the one or more functions to be executed in order to identify additional information to use when responding to the user.
[0009] For example, the total_columns(most_recent('research')) function may return a value of "16," which may be used by the automated assistant to generate another function to be executed for adding a column (e.g., "16+1") to a specific document. For example, the automated assistant may initiate execution of functions such as: action:new_column((16+1), most_recent('research')) and action:comment(column(16+1), "Could someone add this month's data to this column?", most_recent('research')). In this example, the command "most_recent('research')," when executed, may result in identification of one or more documents that have recently been accessed by the user and include a semantic annotation with the term "research." When the function most_recent('research') results in a specific document that the user is referring to, the "new_column" and "comment" functions may be executed to edit the specific document based on the user's spoken utterance.
[0010] Execution of the above function can cause other instances of the automated assistant to notify each researcher and / or each user with permission to view the spreadsheet of changes to the spreadsheet. For example, another user can receive a notification from the automated assistant indicating that the user edited the spreadsheet and incorporated a comment (e.g., in some instances, the user can edit via the desktop computer GUI without invoking the automated assistant). The automated assistant can generate the notification by comparing a previous version of the spreadsheet with the current version of the spreadsheet (e.g., using an API call to a document application) and / or by processing a spoken utterance from the user. The automated assistant can audibly provide the notification via a push notification rendered in the foreground of the cell phone's GUI. The push notification can include content such as "The spreadsheet has been edited by Mary to include a new column and a new comment." In response to receiving the notification, the other user can view the new column in the spreadsheet via their cell phone. Alternatively or additionally, the other user can also edit the spreadsheet via their automated assistant by providing another spoken utterance—without having to open a specific document application in the foreground of the cell phone's GUI.
[0011] As an example, based on a push notification provided via the automated assistant, another user may provide an additional spoken utterance such as “Assistant, what does the new comment say?” However, because the other user may not have explicitly identified the document to be accessed, the automated assistant may infer the identity of the document based on various semantic annotations and / or other contextual data. For example, the automated assistant may identify one or more documents recently accessed and / or modified by the other user and determine whether any recently accessed documents have the characteristics described in the spoken utterance. For example, each of the one or more documents may include a semantic annotation that characterizes a sub-portion of the respective document. Based on this analysis, the automated assistant may identify a spreadsheet as being subject to the additional spoken utterance because the spreadsheet includes a recently added comment (i.e., “new comment”). Furthermore, based on the additional spoken utterances, the automated assistant can access the content of the most recently added comment and audibly render the content of the added comment for other users without graphically rendering the entire (or any one) spreadsheet (e.g., "Mary's new comment recites: 'Could someone add this month's data to this column?'").
[0012] In some embodiments, other users can use the automated assistant to supplement a spreadsheet with data from a separate document—without directly interacting with the document editing application's interface. For example, as part of a backend process, and with prior permission from the other user, the automated assistant can analyze documents stored in association with the other user (e.g., associated with the user's account, such as that utilized by the automated assistant or linked to the user's automated assistant account) so that the automated assistant has a semantic understanding of those documents. This semantic understanding can be embodied in semantic annotations that can be stored as metadata associated with each document. In this way, users will be able to edit and review documents via their automated assistant using assistant commands that are specific to the document's semantic understanding rather than being part of the document.
[0013] As an example, and based on the automated assistant audibly rendering the added comments, other users may provide spoken utterances to their cell phones to cause the automated assistant to edit a spreadsheet using data from a separate document. The spoken utterance may be "Assistant, please fill in that column using data from this month's sensor data spreadsheet." In response to the spoken utterance, the automated assistant and / or other associated applications may process the spoken utterance and / or one or more documents to identify one or more actions to be performed. For example, the automated assistant may use one or more trained machine learning models to determine actions that are synonymous with the verb "fill." As a result, the automated assistant may identify an "insert()" function for the document application. In some embodiments, to identify the data that other users are referring to within the document, the automated assistant may identify a list of recently created documents and filter out any documents that were not created "this month." As a result, the automated assistant may be left with a reduced list of documents that the automated assistant may select to be subjected to the "insert()" function.
[0014] In some embodiments, even though other users provided a descriptor of the source document (e.g., "this month's 'sensor data' spreadsheet") rather than the actual name (e.g., "August TL-9000"), the automated assistant may still identify the correct source document from the reduced list of "this month's" documents. For example, the automated assistant may employ one or more heuristics and / or one or more trained machine learning models to identify the source document and / or the data within the source document that "fills" into the spreadsheet based on the comments. In some embodiments, semantic annotation data may already be stored in association with the source document and may indicate that the device identifier "TL-9000," listed as the name of a column in the source document, is synonymous with the type of temperature "sensor." This information about temperature sensors may be described in search results from an internet search or other knowledge base lookup performed by the automated assistant using document data (e.g., the document title "August TL-9000"), thereby providing a correspondence between the column names and the automated assistant's request using "sensor data."
[0015] When the automated assistant has identified the specific source document and data column that the other user is referring to in the spoken utterance, the automated assistant can use that data column to perform an "insert()" function. For example, the automated assistant can generate a command such as "insert(column("August_TL-9000",11),column("Research_Document",17)", where "11" refers to the column of the sensor data spreadsheet that includes "this month's" data, and where "17" refers to the "new" column previously added by the user.
[0016] In some embodiments, instances of document-related data can be shared between instances of an automated assistant to enable each instance of the automated assistant to more accurately edit a document according to user instructions. For example, because a user initially caused their automated assistant to add a new column identified as column "17," the identifier "17" for "new column" can be shared with another instance of the automated assistant invoked by another user. Alternatively or additionally, data representing the user-assistant interaction can be stored in association with a document and / or a portion of a document to enable each instance of the automated assistant to more accurately execute user instructions. For example, a semantic annotation stored in association with a spreadsheet can characterize column "17" as "this month's data per Mary." In this way, any other instance of the automated assistant that receives a command associated with "Mary" or "this month's data" can refer to column "17" because of the correlation between the command and the semantic annotation associated with column "17." This allows the automated assistant to more efficiently execute automated assistant requests for certain documents as the user continues to interact with the automated assistant to edit certain documents.
[0017] When the automated assistant has caused the "insert()" function to be executed and the spreadsheet is modified to include the additional "sensor" data, the automated assistant can generate a notification that will be pushed to each researcher. In some embodiments, the push notification generated based on the modification can represent the edits made by the other user via the automated assistant. For example, the push notification can be rendered for the other researchers as a GUI element, such as a callout bubble with a graphical rendering of the portion of the spreadsheet modified by the other user. The GUI element can be generated by the automated assistant using an API or other interface associated with the document application to generate graphical data representing the portion of the spreadsheet modified by the other user.
[0018] In some embodiments, when the other user is speaking to the automated assistant, the automated assistant can smoothly transition between speaking and interpreting commands. For example, instead of having the automated assistant copy data from a separate document to a spreadsheet, the other user can speak a combination of: (i) instructions to be executed and (ii) text to be merged into the spreadsheet. For example, the other user can be referring to a printed data set and provide spoken utterances such as "Assistant, in the new column of the research document, add '39 degrees' to the first cell, '42 degrees' to the second cell, and '40 degrees' to the third cell." In one embodiment, the automated assistant may select a cell (Assistant, add "39 degrees" to the first cell, "42 degrees" to the second cell, and "40 degrees" to the third cell in a new column for the research document) in response to the spoken utterance from the other user. The automated assistant may identify the spreadsheet that was recently modified to include the new column and select one or more cells in the new column. The automated assistant may determine that the spoken utterance from the other user includes a certain amount of dictation and select portions of the text data transcribed from the spoken utterance to be incorporated into specific cells. Alternatively or additionally, the format of each cell in the new column may be modified to correspond to a "degree" value so that the new column will reflect the units of the data to be added to the new column and as specified in the spoken utterance (e.g., degrees Celsius). The automated assistant may input each numerical value according to the spoken utterance based at least on text-to-speech processing and natural language understanding of the entire spoken utterance. In some embodiments, the spoken utterance may be performed without requiring the other user to access a document application that provides access to a GUI for editing the spreadsheet, but the modifications may be performed through verbal interaction with the automated assistant.
[0019] The above description is provided as an overview of some embodiments of the present disclosure. These embodiments and further descriptions of other embodiments will be described in more detail below.
[0020] Other embodiments may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform a method (such as one or more methods described above and / or elsewhere herein). Other embodiments may include a system of one or more computers including one or more processors operable to execute stored instructions to perform a method, such as one or more methods described above and / or elsewhere herein.
[0021] It should be understood that all combinations of the above concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1A 、 Figure 1B and Figure 1C Illustrated is a view of one or more users interacting with an automated assistant to create and edit a document without having to directly edit the document via a GUI interface.
[0023] Figure 2A 、 Figure 2B 、 Figure 2C and Figure 2D Illustration of a view of a user creating and editing a document using an automated assistant.
[0024] Figure 3 Illustrated is a system for providing an automated assistant that can edit, share, and / or create various types of documents in response to user input.
[0025] Figure 4 Illustrated is a method for enabling an automated assistant to interact with a document application to edit a document without requiring a user to directly interact with the document application's interface.
[0026] Figure 5 is a block diagram of an example computer system. DETAILED DESCRIPTION
[0027] Figure 1A 、 Figure 1B and Figure 1C Views 100, 120, and 140 illustrate one or more users interacting with an automated assistant to create and edit documents without having to edit the document directly via a GUI interface. In this way, users are not limited to a display interface when creating and editing documents, but can rely on an automated assistant that can be accessed from a variety of different interfaces. For example, Figure 1AAs shown, the first user 102 may be jogging outside when they happen to have an idea for a specific report they want to generate. The first user 102 may request the automated assistant to create a report by providing a first spoken utterance 106 such as "Assistant, create a report from my report template and share it with Howard." The first spoken utterance 106 may be received at an interface of a client computing device 104, which may be a wearable computing device. The client computing device 104 may provide access to an instance of the automated assistant, which may interface with a document application for creating, sharing, and / or editing documents.
[0028] In response to receiving the first spoken utterance 106, the automated assistant may initiate execution of one or more functions to cause the document application to create a new document from the "report template". In some embodiments, the automated assistant may use natural language understanding to identify and / or generate one or more functions for execution. In some embodiments, one or more trained machine learning models may be used when processing the first spoken utterance 106 and / or generating one or more functions to be executed by the document application. With prior permission from the user 102, such processing may include accessing data that has recently been accessed by the document application. For example, historical data associated with the document application may indicate that the user 102 has previously identified another user with the name "Howard" when providing editing permissions to certain documents created via the document application. In this manner, the automated assistant may call a previously executed function, but swap one or more slot values of that function in order to satisfy any request embodied in the first spoken utterance 106.
[0029] As a result of executing one or more functions identified by the automated assistant, and in response to the first spoken utterance 106, the document application may create a new report and provide editing permission to another user named "Howard." For example, the user "Howard" may be identified in a contact list stored in association with the user 102. Furthermore, the user 102 may provide a second spoken utterance 108 such as, "In the intro section, add a paragraph and a comment," and a third spoken utterance 110 such as, "The comment should say: 'This is where you should discuss the results.'" In response to receiving the second spoken utterance 108 and the third spoken utterance 110, the automated assistant may generate one or more additional functions to be executed by the document application. For example, in some implementations, because the user 102 only uses the automated assistant to perform document-related tasks, the automated assistant may select one or more automatic speech recognition (ASR) technologies suitable for understanding document-related queries.
[0030] In some embodiments, the selected ASR technology may employ a specifically trained machine learning model that is trained using document-related data and / or based on data associated with previous interactions between the user and the document application. In some additional or alternative embodiments, the selected ASR technology may bias the specifically trained machine learning model (e.g., a general model) toward the recognition of terms that are frequently encountered in document-related queries. One or more of a variety of biasing techniques may be utilized to bias speech recognition toward a particular term or terms. As an example, a language model used in some ASR technologies may include weights for terms, where each weight reflects the degree of bias for a corresponding term. As another example, biasing toward terms may be achieved simply by including the terms in the language model used in the ASR. As yet another example, decoding graphs optionally utilized in the ASR may be decoded while biasing toward certain terms. As another example, biasing may be used to generate one or more additional hypotheses in addition to the initial hypothesis (or multiple initial hypotheses) generated by the ASR model, and these additional hypotheses may be considered candidate transcriptions. For example, additional hypotheses may be generated and / or selected based on the inclusion of biased terms. In these and other ways, any spoken utterance provided by user 102 in the context of first spoken utterance 106 may be more accurately interpreted by the automated assistant.
[0031] In response to the second spoken utterance 108, the automated assistant may generate a function that causes the document application to create a new paragraph in the portion of the new report just created in response to the first spoken utterance 106. In some embodiments, the newly created report document may be associated with one or more semantic annotations that each include one or more corresponding semantic interpretations for a particular portion of the report document. This may be in part because the report document is created from a template, which may include existing semantic annotations. However, in some embodiments, when the automated assistant is requested to perform an operation associated with a particular document, the automated assistant may generate and / or access semantic annotations associated with the particular document. Such semantic annotations may allow the automated assistant to associate certain input from the user with certain portions of one or more documents accessible to the user. For example, the report document created by the user 102 may include a semantic annotation that characterizes a paragraph in the second page of the report document as an "introduction" (e.g., <paragraph-3>} == ["introduction," "beginning," "opening"]). Thus, because "intro" as mentioned in the second spoken utterance 108 is synonymous with "introduction" as described by the semantic annotation, the automated assistant can generate a function that causes the document application to create a new paragraph in the second page of the report document.
[0032] In addition, as requested by user 102 via second spoken utterance 108, the automated assistant can generate another function that associates a comment with the new paragraph created on the second page. The automated assistant can also generate another function based on third spoken utterance 110 to include the text "This is where you should discuss theresults" in the comment. As each generated function can be created and executed, the automated assistant can optionally provide output 112 that represents the progress of the request from user 102. For example, output 112 can be audibly rendered via client computing device 104 and include natural language content such as "Ok, I've created the report and added the paragraph and the comment. I've also shared the report with Howard."
[0033] Figure 1B 1 shows a view 120 of a second user 126 responding to an automated assistant that has provided a notification 128 indicating that the first user 102 has shared a document with the second user 126. In some implementations, the second user 126 may have a client computing device 124 that provides access to another instance of the automated assistant. The automated assistant associated with the second user 126 may be provided by the same entity or a different entity that provides access to the automated assistant. Figure 1A In some implementations, separate instances of the automated assistant can communicate via an API and / or other interface for communicating between applications.
[0034] The notification 128 provided by the automated assistant via the client computing device 124 may include natural language content, such as "Katherine has shared a document with you." The natural language content of the notification 128 may optionally be audibly rendered to the second user 126 and / or a graphical notification 138 may be rendered at the GUI 136 of the client computing device 124. For example, in some implementations, when the first user 102 invokes the automated assistant to perform an action associated with a document accessible to the second user 126, the automated assistant may cause the graphical notification 138 to be rendered at the GUI 136. The graphical notification 138 may include rendering a portion 134 of a report created by the first user 102, and the rendered portion 134 may include an annotation 132 directed to the second user 126. In some implementations, to generate the graphical notification 138, the automated assistant associated with the first user 102 and / or the second user 126 may request, via an API call or other request, that the document application provide certain data related to the second user 126. For example, the automated assistant may request that the document application provide a sub-portion of the document as a whole in a form similar to the GUI that the document application would generate.
[0035] The document application can optionally provide a graphical rendering of the specific portion 134 of the document corresponding to the comment 132, thereby allowing the second user 126 to visualize the context of the comment 132. In some embodiments, an automated assistant associated with the second user 126 can request that the document application provide a rendering of the specific portion 134 of the document based on the type of interface available via the client computing device 124. In some embodiments, this request can be fulfilled when the document application provides an image file, text data, audio data, video data, and / or any combination of data representing the specific portion 134 of the document. In this way, the second user 126 can receive an audible rendering of the comment 132 and / or any sub-portion of the document associated with the comment 132 when a display interface is not currently available to the second user 126.
[0036] When the second user 126 has confirmed the comment 132 from the first user 102, the second user 126 can provide a corresponding request to their respective automated assistants. For example, the second user 126 can provide a spoken utterance 130 such as "Assistant, please add the following statement to that new paragraph:'The results confirm our earliest predictions.'". In response, the automated assistant can generate one or more functions that, when executed, cause the document application to modify the portion of the report corresponding to the new paragraph to include the statement from the second user 126. In some embodiments, the one or more functions can be more efficiently generated based on one or more previous interactions between the first user 102 and the automated assistant and / or the second user 126 and the automated assistant. For example, the automated assistant can access data characterizing one or more interactions that have occurred in which the report document was the subject of the one or more interactions. Such data can be used by the automated assistant to identify slot values for one or more functions to be executed in order to fulfill the request from the user. For example, in response to spoken utterance 130, the automated assistant can determine whether any document accessible to second user 126 has been recently edited to include a new paragraph. When a report document is identified as having been recently edited to include a new paragraph, the automated assistant can invoke a document application to edit the new paragraph based on spoken utterance 130 from second user 126.
[0037] In some implementations, based on the edits made by the second user 126, the automated assistant can generate one or more additional semantic annotations that characterize one or more sub-portions of the entire report document. For example, the additional semantic annotations can be Figure 1B The edits made in are characterized as "new paragraph by Howard; results". Thereafter, the content of this semantic annotation can be used when selecting whether the report document and / or a sub-part of the report document is the subject of another spoken utterance from the first user 102 or the second user 126. As an example, in Figure 1C In FIG. 1 , a first user 102 may provide spoken utterance 142 to an automated assistant accessible via a client computing device 146 as a standalone speaker device. The spoken utterance 142 may be, for example, “Assistant, were anymore edits made by Howard?” The spoken utterance 142 may be provided by the first user 102 in a manner that includes information about the automated assistant. Figure 1A and Figure 1B The time point after the period of interaction described is provided.
[0038] In response to receiving spoken utterance 142, the automated assistant may determine that first user 102 is referring to the contact "Howard" and identify a report document. The automated assistant may identify the report document based on the fact that first user 102 identified "Howard" in an interaction with the automated assistant when first user 102 requested the automated assistant to create the report document. This may enable the automated assistant to rank and / or otherwise prioritize the report document relative to other documents accessible to the automated assistant in order to fulfill one or more requests in spoken utterance 142. Once the automated assistant has identified the report document, it may determine that first user 102 is requesting the automated assistant to identify any recent changes made by the contact "Howard." In response, the automated assistant may generate one or more functions for causing the document application to provide information about certain edits based on the author of the identified edits. For example, the one or more functions may include recent_edits("report document", "Howard", most_recent()), which, when executed, may return a summary of one or more recent modifications made to the identified document (e.g., "report document"). For example, the summary may include a semantic understanding of edits made to the report document and / or an indication of one or more types of edits made to the identified document (eg, additions of text).
[0039] As a result, and in response to the spoken utterance 142, the automated assistant may provide output 148, such as "Howard added text to the new paragraph." In some implementations, the automated assistant may generate semantic annotations for the document as it is being edited. In this manner, subsequent automated assistant input related to the document may be more easily fulfilled while mitigating delays that may occur during document recognition. For example, when the first user 102 provides another spoken utterance 150, such as "Read the conclusion," the automated assistant may recognize the semantic annotation characterizing a portion of a report document as having concluding language—even though the report document does not yet have the word "conclusion" in the content of the report document. Thereafter, the automated assistant may provide audible output 152 via the client computing device 146 and / or visual output at the television computing device 144. For example, the client computing device 146 may render the audible output "Sure…," followed by the automated assistant audibly rendering the paragraph of the report document corresponding to the concluding semantic annotation.
[0040] Figure 2A 、 Figure 2B 、 Figure 2C and Figure 2D Views 200, 220, 250, and 260 illustrate user 202 creating and editing a document using an automated assistant. User 202 can initiate the creation of a document, such as a spreadsheet, by first invoking the automated assistant to determine whether certain document sources are available to the automated assistant. For example, user 202 can provide a first spoken utterance 206, such as "Do I have any notes related to solar cells?" to an interface of a computing device in a vehicle 208, which can provide access to the automated assistant. In response, the automated assistant can perform one or more searches for documents that include and / or are associated with the term "solar cells." When the automated assistant identifies multiple different documents related to the term identified by user 202, the automated assistant can provide output 204, such as "Sure."
[0041] When user 202 confirms output 204 from the automated assistant, user 202 may provide a second spoken utterance 214 such as, "Could you consolidate those notes into a spreadsheet and read it to me?" In response to receiving second spoken utterance 214, the automated assistant may generate one or more functions to be executed by the document application to cause the document application to merge the identified documents into a single document. When the merged single document is created by the document application, the automated assistant may access the merged document to fulfill the subsequent user request for the automated assistant to read the merged document to user 202. For example, the automated assistant may respond with another output 216 such as, "Ok..." and then read the merged document (e.g., a new spreadsheet) to user 202.
[0042] In some implementations, the automated assistant may store semantic annotations in association with the merged document in order for the automated assistant to perform further operations on the merged document. Figure 2B As provided, the automated assistant can generate a request for a semantic annotation to be associated with a merged document (e.g., spreadsheet 232). The request can be executed at a vehicle computing device and / or a remote computing device 222 (such as a remote server device). In some embodiments, one or more techniques for generating semantic annotations for specific sub-portions of a document can be employed. For example, one or more trained machine learning models and / or one or more heuristic methods can be utilized to generate semantic annotations for spreadsheet 232. The data processed to generate a particular semantic annotation can include: the content of spreadsheet 232, interaction data representing interactions between user 202 and the automated assistant, the document on which spreadsheet 232 is based, and / or any other data source that can be associated with spreadsheet 232 and / or the automated assistant. For example, each of the corresponding semantic annotations (224, 226, 228, and 230) can be generated based on the content of spreadsheet 232 and / or the various design documents (Design_1 238, Design_2 240, Design_3 242, and Design 244) used to create each corresponding row of spreadsheet 232. Alternatively or additionally, semantic annotations 234 may be generated for the entire electronic form 232 to provide a semantic understanding that the automated assistant may refer to when attempting to fulfill subsequent automated assistant requests.
[0043] For example, Figure 2C , the user 202 may provide another spoken utterance 252, such as "Assistant, anytime wattage is mentioned, add a comment." In response, the automated assistant may generate one or more functions to be executed by the document application to fulfill the request from the user 202. When the document application executes the one or more functions, the document application and / or the automated assistant may identify instances of the term "wattage" in the spreadsheet 232 and associate a corresponding comment with each instance of the term wattage. Upon completion, the document application may optionally invoke an API call to the automated assistant to cause the automated assistant to provide an indication 254 (e.g., "Sure") that the requested action has been performed. Alternatively or additionally, the indication 254 may be generated using a summary of recent edits, such as: the number of edits performed across the document, a summary of the changes made, a graphical indication of the latest version of the document, and / or any other information that may characterize the changes to the document.
[0044] In some implementations, user 202 may provide spoken utterance 256 to cause the automated assistant to notify another user of changes that have been made to spreadsheet 232. For example, spoken utterance 256 may be "Also, could you tag William in each comment and ask him to confirm the wattage amounts?" In response, the automated assistant may generate one or more functions that, when executed, cause the document application to modify the comments corresponding to the term "wattage" and also provide a message about each comment to the contact (e.g., "William"). To perform the above operations on the intended document, the automated assistant may identify one or more documents recently accessed by user 202 and / or the automated assistant. This may allow the automated assistant to identify spreadsheet 232 that may have been recently modified to include comments and / or certain semantic annotations.
[0045] For example, each respective semantic annotation (224, 226, 228, and 230) can be stored in association with a particular row in spreadsheet 232, and each respective semantic annotation can include one or more terms synonymous with the unit of measurement "Watts." For example, each respective semantic annotation can identify terms such as "wattage," "watts," "power," and / or any other terms synonymous with "Watts." Thus, in response to spoken utterance 256, the automated assistant can identify spreadsheet 232 as being most associated with "wattage amounts" and cause each wattage-related comment in spreadsheet 232 to include a request for "William" to confirm any "wattage amounts." In response, the automated assistant can provide output 258, such as "Ok, I've tagged William in each comment and asked him to confirm."
[0046] In some implementations, a second user 276 (e.g., William) can interact with an instance of the automated assistant to further edit the spreadsheet 232. For example, the automated assistant accessible to the second user 276 can provide output 262 via the client computing device 272. The output 262 can include natural language content, such as "You have been tagged in comments within a spreadsheet." In response, the second user 276 can provide a first spoken utterance 264, such as "What did they say?" The automated assistant can process the first spoken utterance 264 and determine that the second user 276 is referring to a comment within the spreadsheet 232, and then access the text of the comment. The automated assistant can then cause the client computing device 272 or another computing device 270 (e.g., a television) to render output representing the text of the comment. For example, the automated assistant may cause the client computing device 272 to render another audible output 266, such as "Confirm the wattage amounts in the spreadsheet," and also provide an indication that the spreadsheet 232 will be rendered at a nearby display interface 274 (e.g., "I will display the spreadsheet for you"). The automated assistant may then cause the display interface 274 to render a sub-portion of the entire spreadsheet 232. The second user 276 may continue editing the spreadsheet 232 via the document application by providing a second spoken utterance 268, such as "Reply to each comment by saying: These all appear correct." In response, the automated assistant may generate one or more functions that, when executed, cause the document application to edit each spreadsheet comment directed to the second user 276 (e.g., William). In this way, the second user 276 is able to review and edit the document without having to use their limbs to manually control certain peripherals of the dedicated document editing device.
[0047] Figure 3 The diagram illustrates a system 300 for providing an automated assistant 304 that can edit, share, and / or create various types of documents in response to user input. The automated assistant 304 can operate as part of an assistant application provided at one or more computing devices, such as a computing device 302 and / or a server device. A user can interact with the automated assistant 304 via an assistant interface 320, which can be a microphone, a camera, a touch screen display, a user interface, and / or any other device capable of providing an interface between a user and an application. For example, a user can initialize the automated assistant 304 by providing spoken, textual, gestural, and / or graphical input to the assistant interface 320, causing the automated assistant 304 to initiate one or more actions (e.g., provide data, control peripheral devices, access an agent, generate input and / or output, etc.). Alternatively, the automated assistant 304 can be initialized based on processing context data 336 using one or more trained machine learning models. The context data 336 can characterize one or more characteristics of the environment in which the automated assistant 304 is accessible, and / or one or more characteristics of a user predicted to be intended to interact with the automated assistant 304. The computing device 302 may include a display device, which may be a display panel, including a touch interface for receiving touch input and / or gestures to allow a user to control an application 334 of the computing device 302 via the touch interface. In some embodiments, the computing device 302 may lack a display device, thereby providing an audible user interface output rather than a graphical user interface output. In addition, the computing device 302 may provide a user interface such as a microphone for receiving oral natural language input from the user. In some embodiments, the computing device 302 may include a touch interface and may lack a camera, but may optionally include one or more other sensors.
[0048] Computing device 302 and / or other third-party client devices can communicate with the server device via a network such as the Internet. In addition, computing device 302 and any other computing devices can communicate with each other via a local area network (LAN) such as a Wi-Fi network. Computing device 302 can offload computing tasks to the server device to conserve computing resources at computing device 302. For example, the server device can host automated assistant 304, and / or computing device 302 can transmit input received at one or more assistant interfaces 320 to the server device. However, in some embodiments, automated assistant 304 can be hosted at computing device 302, and various processes that may be associated with automated assistant operations can be performed at computing device 302.
[0049] In various embodiments, all or less than all aspects of automated assistant 304 can be implemented on computing device 302. In some of these embodiments, aspects of automated assistant 304 are implemented via computing device 302 and can be interfaced with a server device that can implement other aspects of automated assistant 304. The server device can optionally serve multiple users and their associated assistant applications via multiple threads. In embodiments where all or less than all aspects of automated assistant 304 are implemented via computing device 302, automated assistant 304 can be an application separate from the operating system of computing device 302 (e.g., installed "on top" of the operating system)—or can alternatively be implemented directly by the operating system of computing device 302 (e.g., considered an application of the operating system but integrated with the operating system).
[0050] In some embodiments, automated assistant 304 may include an input processing engine 306 that may employ a plurality of different modules to process input and / or output from computing device 302 and / or a server device. For example, input processing engine 306 may include a speech processing engine 308 that may process audio data received at assistant interface 320 to identify text embodied in the audio data. The audio data may be transmitted from, for example, computing device 302 to a server device in order to conserve computing resources at computing device 302. Additionally or alternatively, the audio data may be processed exclusively at computing device 302.
[0051] The process for converting audio data into text can include a speech recognition algorithm that can employ a neural network and / or a statistical model for identifying audio data sets corresponding to words or phrases. The text converted from the audio data can be parsed by a data parsing engine 310 and made available to the automated assistant 304 as text data, which can be used to generate and / or identify command phrases, intents, actions, slot values, and / or any other content specified by the user. In some embodiments, the output data provided by the data parsing engine 310 can be provided to a parameter engine 312 to determine whether the user has provided input corresponding to a specific intent, action, and / or a routine that can be executed by the automated assistant 304 and / or an application or agent that can be accessed via the automated assistant 304. For example, assistant data 338 can be stored at a server device and / or computing device 302 and can include data defining one or more actions that can be performed by the automated assistant 304, as well as parameters necessary to perform these actions. The parameter engine 312 can generate one or more parameters for the intent, action, and / or slot value, and provide the one or more parameters to the output generation engine 314. Output generation engine 314 may use one or more parameters to communicate with assistant interface 320 for providing output to a user and / or with one or more applications 334 for providing output to one or more applications 334 .
[0052] In some embodiments, automated assistant 304 can be an application that can be installed "on top of" the operating system of computing device 302 and / or can itself form part of (or all of) the operating system of computing device 302. The automated assistant application includes on-device speech recognition, on-device natural language understanding, and on-device implementation and / or has access to on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition can be performed using an on-device speech recognition module that processes audio data (detected by a microphone) using an end-to-end speech recognition machine learning model stored locally at computing device 302. On-device speech recognition generates recognized text for spoken utterances (if any) present in the audio data. Additionally, on-device natural language understanding (NLU) can be performed, for example, using an on-device NLU module that processes recognized text generated using on-device speech recognition and, optionally, processes context data to generate NLU data.
[0053] The NLU data may include an intent corresponding to the spoken utterance and, optionally, parameters for the intent (e.g., slot values). An on-device fulfillment module that utilizes the NLU data (from the on-device NLU) and, optionally, other local data may be used to perform on-device fulfillment to determine actions to be taken to resolve the intent of the spoken utterance (and, optionally, parameters for the intent). This may include determining local and / or remote responses (e.g., answers) to the spoken utterance, interactions with locally installed applications that execute based on the spoken utterance, commands transmitted to an Internet of Things (IoT) device based on the spoken utterance (directly or via a corresponding remote system), and / or other resolution actions performed based on the spoken utterance. The on-device fulfillment may then initiate local and / or remote execution / implementation of the determined actions to resolve the spoken utterance.
[0054] In various embodiments, remote speech processing, remote NLU, and / or remote fulfillment can be at least selectively utilized. For example, the recognized text can be at least selectively transmitted to a remote automated assistant component for remote NLU and / or remote fulfillment. For example, the recognized text can be optionally transmitted for remote execution in parallel with on-device execution, or transmitted in response to a failure of the on-device NLU and / or on-device fulfillment. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution can be prioritized at least due to the reduced latency they provide when resolving spoken utterances (since no client-server round trip is required to resolve spoken utterances). In addition, the device functionality can be the only functionality available in the absence of network connectivity or with limited network connectivity.
[0055] In some implementations, computing device 302 can include one or more applications 334, which can be provided by a third-party entity distinct from the entity providing computing device 302 and / or automated assistant 304. An application state engine of automated assistant 304 and / or computing device 302 can access application data 330 to determine one or more actions that can be performed by one or more applications 334, as well as a state of each of one or more applications 334 and / or a state of a corresponding device associated with computing device 302. A device state engine of automated assistant 304 and / or computing device 302 can access device data 332 to determine one or more actions that can be performed by computing device 302 and / or one or more devices associated with computing device 302. Additionally, application data 330 and / or any other data (e.g., device data 332) can be accessed by automated assistant 304 to generate context data 336, which can characterize the context in which a particular application 334 and / or device is executing and / or the context in which a particular user is accessing computing device 302, accessing application 334, and / or any other device or module.
[0056] When one or more applications 334 are executed at computing device 302, device data 332 can represent the current operating state of each application 334 executing at computing device 302. In addition, application data 330 can represent one or more characteristics of the executing applications 334, such as the content of one or more graphical user interfaces rendered at the command of one or more applications 334. Alternatively or additionally, application data 330 can represent action plans that can be updated by the corresponding applications and / or by automated assistant 304 based on the current operating state of the corresponding applications. Alternatively or additionally, the one or more action plans for one or more applications 334 can remain static but can be accessed by the application state engine to determine the appropriate action to initiate via automated assistant 304.
[0057] Computing device 302 may also include an assistant invocation engine 322 that can use one or more trained machine learning models to process application data 330, device data 332, context data 336, and / or any other data accessible to computing device 302. Assistant invocation engine 322 can process this data to determine whether to wait for the user to explicitly speak an invocation phrase to invoke automated assistant 304, or to consider the data to indicate the user's intent to invoke the automated assistant—without requiring the user to explicitly speak an invocation phrase. For example, one or more trained machine learning models can be trained using instances of training data based on scenarios in which the user is in an environment where multiple devices and / or applications are exhibiting various operating states. The instances of training data can be generated to capture training data that characterizes contexts in which the user invokes the automated assistant and other contexts in which the user does not invoke the automated assistant. When training one or more trained machine learning models based on these instances of training data, assistant invocation engine 322 can cause automated assistant 304 to detect or limit detection of spoken invocation phrases from the user based on characteristics of the context and / or environment. Additionally or alternatively, assistant invocation engine 322 can cause automated assistant 304 to detect or limit detection of one or more assistant commands from the user based on characteristics of the context and / or environment. In some implementations, assistant invocation engine 322 can be disabled or limited based on computing device 302 detecting an assistant that is suppressing output from another computing device. In this manner, when computing device 302 is detecting an assistant that is suppressing output, automated assistant 304 will not be invoked based on contextual data 336—otherwise, if no assistant that is suppressing output is detected, automated assistant 304 will be invoked.
[0058] In some embodiments, system 300 may include a document recognition engine 316 that can identify one or more documents that a user may be requesting automated assistant 304 to access, modify, edit, and / or share. Document recognition engine 316 may be used when processing natural language content of user input, such as spoken utterances. Based on this processing, document recognition engine 316 may determine scores and / or probabilities for one or more documents accessible to automated assistant 304. The particular document with the highest score and / or highest probability may then be identified as the document to which the user is referring. In some instances, when two or more documents have similar scores and / or probabilities, the user may be prompted to clarify the document to which they are referring, and the prompt may optionally include characteristics of the two or more documents (e.g., title, content, collaborators, recent edits, etc., from the document with a particular score). In some embodiments, factors influencing whether a particular document is recognized by document recognition engine 316 may include the context of the user input, previous user input, whether the user input identifies another user, the user's timeline, whether the content of the user input is similar to the content of one or more semantic annotations, and / or any other factors that may be associated with a document.
[0059] In some embodiments, system 300 may include a semantic annotation engine 318 that can be used to generate and / or identify semantic annotations for one or more documents. For example, when automated assistant 304 receives an indication that a person is sharing a document with an authenticated user of automated assistant 304, semantic annotation engine 318 can be employed to generate a semantic annotation for the document. The semantic annotation can include natural language content and / or other data that provides an interpretation of one or more sub-portions of the entire document. In this manner, automated assistant 304 can rely on semantic annotations for a variety of different documents when determining that a particular document is the subject of a request from a user.
[0060] In some embodiments, semantic annotations can be generated for specific documents based on how one or more users refer to the specific document. For example, a specific document can include semantic annotations along with other content in the document body. While the document's content and semantic annotations can include various descriptive language, the document's content may not include any terms that users tend to use when referring to the document. For example, a user may refer to a specific spreadsheet as a "home maintenance" spreadsheet, even though the specific spreadsheet does not include the terms "home" or "maintenance." However, based on automated assistant 304 receiving user input referring to a "home maintenance" spreadsheet and automated assistant 304 identifying the "home maintenance" spreadsheet using document recognition engine 316, semantic annotation engine 318 can generate a semantic annotation. The semantic annotation can be incorporated into the term "home maintenance" and can be stored in association with the "home maintenance" spreadsheet. In this way, automated assistant 304 can conserve processing bandwidth when identifying documents that users are likely to refer to. Additionally, this can allow for more efficient document editing via automated assistant 304, as automated assistant 304 can adapt to the dynamic perspectives that users may have regarding certain documents.
[0061] In some embodiments, system 300 may include a document action engine 326 that can identify one or more actions to be performed on a particular document. Document action engine 326 can identify the one or more actions to be performed based on user input, past interaction data, application data 330, device data 332, context data 336, and / or any other data that can be stored in association with a document. In some embodiments, one or more semantic annotations associated with one or more documents can be used to identify the specific action that the user is requesting automated assistant 304 to perform. Alternatively, or in addition, document action engine 326 can identify the one or more actions to be performed on a particular document based on the content of the particular document and / or the content of one or more other documents.
[0062] For example, a user may provide input such as "Assistant, add a new 'Date' row to my finance spreadsheet." In response, the document action engine 326 may determine that the user is associated with a "finance" spreadsheet that includes various dates listed under columns, and may then determine that appropriate actions to perform include executing a new_row() function and insert_date(). In this way, a new row may be added to the "finance" spreadsheet, and a current date entry may be added to the new row. In some embodiments, the selection of a function to perform may be based on processing the user input and / or other contextual data using one or more trained machine learning models and / or one or more heuristic processes. For example, a particular trained machine learning model may be trained using training data that is based on instances where another user requested their corresponding automated assistant to perform a particular operation, but then the other user manually performed the particular operation. Thus, the training data may be derived from crowdsourcing techniques used to teach automated assistants to accurately respond to requests from various users who directed their automated assistants to perform operations associated with particular documents.
[0063] In some implementations, system 300 may include a document preview engine 324 that can process data associated with one or more documents to allow automated assistant 304 to provide an appropriate preview of a particular document. For example, a first user may have the automated assistant edit a particular document and then share the particular document with a second user. In response, the instance of the automated assistant associated with the second user may utilize document preview engine 324 to render a preview of the particular document—without having the document application occupy the entire display interface of the computing device. For example, the automated assistant may utilize an API to retrieve graphical preview data from the document application and then render a graphical notification for the second user based on the graphical preview data. Alternatively or additionally, the automated assistant may include functionality for rendering a portion of a document that has been edited by the user, without necessarily rendering the entire document. For example, a user may provide a spoken utterance such as, "Assistant, what is the latest slide added to my architecture presentation?" In response, automated assistant 304 may invoke document recognition engine 316 to identify the "architecture" presentation document and also invoke document preview engine 324 to capture a preview (e.g., image and / or text) of the slide recently added to the document. Automated assistant 304 may then cause the display interface of computing device 302 to render a graphic of the recently added slide without causing the entire presentation application to be loaded into the memory of computing device 302. Thereafter, the user may view the slide preview and provide automated assistant 304 with another command for editing the slide and / or adding a comment to the slide (e.g., “Assistant, add a comment to this slide and tag William in the comment”).
[0064] Figure 4 A method 400 is illustrated for enabling an automated assistant to interact with a document application to edit a document without requiring a user to directly interact with an interface dedicated to the document application. Method 400 can be performed by one or more computing devices, applications, and / or any other apparatus or module that can be associated with an automated assistant. Method 400 may include an operation 402 for determining whether automated assistant input has been received by the automated assistant. The automated assistant input can be a spoken word, an input gesture, a text input, and / or any other input that can be used to control the automated assistant. When the automated assistant input is received, method 400 can proceed from operation 402 to operation 404. Otherwise, the automated assistant can proceed to operation 416 for causing one or more other functions to be performed in response to the automated assistant input.
[0065] Operation 404 may include determining whether the automated assistant input relates to a particular document, such as a document accessible via a document application. In some embodiments, the document application may be an application installed on the client computing device and / or accessible via a browser or other web application. Alternatively or additionally, the document application may be provided by the same or a different entity than the entity providing the automated assistant. When the automated assistant determines that the automated assistant input relates to a document, method 400 may proceed to operation 406. Otherwise, the automated assistant may perform one or more operations to facilitate the fulfillment of any request embodied in the automated assistant input and return to operation 402.
[0066] Operation 406 may include identifying a specific document that the user is requesting to modify. In some embodiments, interaction data representing one or more previous interactions between the user and the automated assistant may be used to identify the specific document. Alternatively or additionally, the automated assistant may access application data associated with one or more document-related applications to identify one or more documents that may be relevant to the automated assistant input. For example, the automated assistant may generate one or more functions that, when executed by the corresponding document application, cause the document application to return a list of recently accessed documents. The automated assistant may optionally, and with prior permission from the user, access one or more of the listed documents to determine whether a specific document in the one or more listed documents is the document that the user is currently referring to. The automated assistant and / or another application may generate and / or identify semantic annotations associated with each of the listed documents and may use the semantic annotations to determine whether the automated assistant input is relevant to the content of the specific listed document. For example, when a term included in a specific semantic annotation for a specific document is identical or synonymous with a term included in the automated assistant input, the specific document may be prioritized over other less relevant documents when selecting a document to be subject to the automated assistant input.
[0067] In some embodiments, method 400 may proceed from operation 406 to operation 408, which may include determining whether any semantic annotations have been stored in association with the particular document. When semantic annotations are not stored in association with the particular document, method 400 may proceed to operation 410. However, when semantic annotations are stored in association with the particular document, method 400 may proceed to operation 412. Operation 412 may include identifying one or more functions to be performed based on the automated assistant input and / or the semantic annotations. For example, when the automated assistant input refers to a subsection of a particular document, the semantic annotation corresponding to the subsection may include terms that are synonymous with the automated assistant input. For example, when the automated assistant input includes a request to add a comment to a "statistical data" section of a particular document, but the particular document does not have a subsection explicitly labeled "statistical data," the automated assistant may identify statistical terms in one or more semantic annotations. Terms such as "average" and "distribution" can be included in the specific semantic annotation for the specific subsection, thereby providing the automated assistant with a correlation between the specific subsection and the automated assistant input. As a result, the automated assistant can generate a function that points to the specific subsection of the specific document. The function can include a slot value or other parameter that identifies a portion of text in the specific topic via a word count, line reference, paragraph number, page number, and / or any other identifier that can be used to identify the subsection of the document. From operation 412, method 400 can proceed to operation 416, which can include causing execution of one or more functions in response to the automated assistant input.
[0068] When no semantic annotations are currently stored in association with a particular document, method 400 may proceed from operation 408 to operation 410. Operation 410 may include identifying one or more functions (i.e., actions) to be performed based on the automated assistant input and / or the content of the particular document. Method 400 may include an optional operation 414 of generating one or more semantic annotations based on the automated assistant input and / or the content of the particular document. For example, when a user uses one or more terms to refer to a particular sub-portion of a particular document, the automated assistant may generate a semantic annotation comprising the one or more terms. The generated semantic annotation may then be associated with the particular document and, in particular, associated with the particular sub-portion of the particular document, stored as metadata. Method 400 may proceed from operation 410 and / or operation 414 to operation 416, in which one or more functions are performed in response to the automated assistant input from the user.
[0069] Figure 5 5 is a block diagram 500 of an example computer system 510. The computer system 510 generally includes at least one processor 514 that communicates with a number of peripheral devices via a bus subsystem 512. These peripheral devices may include a storage subsystem 524, including, for example, memory 525 and a file storage subsystem 526; an interface output device 520; a user interface input device 522; and a network interface subsystem 516. The input and output devices allow a user to interact with the computer system 510. The network interface subsystem 516 provides an interface to an external network and couples to corresponding interface devices in other computer systems.
[0070] User interface input devices 522 may include a keyboard; a pointing device such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touch screen incorporated into a display; an audio input device such as a voice recognition system; a microphone; and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computer system 510 or a communication network.
[0071] User interface output device 520 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanisms for creating visual images. The display subsystem may also provide non-visual display such as via an audio output device. Typically, the use of the term "output device" is intended to include all possible types of devices and the mode of outputting information from computer system 510 to a user or another machine or computer system.
[0072] The storage subsystem 524 stores programming and data constructs that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 524 may include logic to perform selected aspects of the method 400 and / or implement one or more of the system 300, client computing device 104, client computing device 124, client computing device 146, vehicle 208, computing device 222, client computing device 264, and / or any other applications, devices, apparatuses, and / or modules discussed herein.
[0073] These software modules are typically executed by the processor 514 alone or in combination with other processors. The memory 525 used in the storage subsystem 524 may include multiple memories, including a main random access memory (RAM) 530 for storing instructions and data during program execution and a read-only memory (ROM) 532 for storing fixed instructions. The file storage subsystem 526 may provide persistent storage for program and data files and may include a hard drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular embodiment may be stored by the file storage subsystem 526 in the storage subsystem 524 or in another machine accessible to the processor 514.
[0074] The bus subsystem 512 provides a mechanism for the various components and subsystems of the computer system 510 to communicate with each other as intended. Although the bus subsystem 512 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.
[0075] Computer system 510 can be of various types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 5 The description of the computer system 510 depicted in FIG is intended only as a specific example for the purpose of illustrating some embodiments. Many other configurations of the computer system 510 are possible with more Figure 5 The computer system depicted in FIG.
[0076] Where the systems described herein collect personal information about users (or "participants" as they are often referred to herein) or may use personal information, users may be provided with the opportunity to control whether a program or feature collects information about the user (e.g., information about the user's social network, social behavior or activities, occupation, the user's preferences, or the user's current geographic location), or to control whether and / or how content is received from content servers that may be more relevant to the user. Additionally, certain data may be processed in one or more ways before being stored or used so that personally identifiable information is removed. For example, a user's identity may be treated so that no personally identifiable information can be determined for that user, or the user's geographic location for which geographic location information is obtained may be generalized (such as to the city, zip code, or state level) so that the user's specific geographic location cannot be determined. Thus, users may have control over how information about the user is collected and / or used.
[0077] Although several embodiments have been described and illustrated herein, various other means and / or structures for performing the functions and / or obtaining the results and / or one or more of the advantages described herein may be utilized, and each of these variations and / or modifications is considered within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or configurations will depend on the specific application or applications in which the teachings are used. Those skilled in the art will recognize or be able to ascertain, using no more than routine experimentation, many equivalents to the specific embodiments described herein. Therefore, it should be understood that the foregoing embodiments are presented only as examples, and that, within the scope of the appended claims and their equivalents, the embodiments may be practiced otherwise than as specifically described and claimed. Embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods, provided such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the scope of the present disclosure.
[0078] In some embodiments, a method implemented by one or more processors is described as including operations such as: receiving user input directed to an automated assistant from a user at an automated assistant interface of a computing device, wherein the user input includes a request for the automated assistant to access or modify a document. The method may further include the following operations: in response to receiving the user input, identifying a specific document that the user is requesting to access or modify, wherein the specific document is stored at the computing device or another computing device, wherein an explicit dictation of a corresponding name of the specific document is omitted from the user input, and wherein identifying the specific document includes processing data including natural language content of the user input and content of each of a plurality of different documents accessible via the computing device. The method may further include the following operations: determining one or more actions to be performed on the specific document, wherein determining the one or more actions is based on the user input and one or more semantic annotations for the specific document stored in association with the specific document, and wherein each of the one or more semantic annotations includes a semantic interpretation of a corresponding sub-portion within the entirety of the specific document. The method may further include the following operations: causing the one or more actions to be performed to access or modify the specific document based on the user input.
[0079] In some embodiments, the particular document was not originally created by a user, and the particular document was created using a document application different from an automated assistant. In some embodiments, the data further includes additional semantic annotations, and each of the additional semantic annotations includes another semantic interpretation of another corresponding sub-portion of a corresponding additional document of the plurality of different documents. In some embodiments, the automated assistant interface of the computing device includes a microphone, and the user input is received when a document editing program for editing the particular document is not present in the foreground of a graphical user interface of the computing device. In some embodiments, the one or more semantic annotations include a specific semantic annotation, the specific semantic annotation includes a semantic interpretation of a document comment created by an additional user, and the one or more actions include causing a notification to be provided to the additional user via another interface of a separate computing device associated with the additional user.
[0080] In some embodiments, the method may further include the following operations: prior to receiving user input at the automated assistant interface of the computing device: receiving another user input, the other user input comprising another request for the automated assistant to render a description of supplemental content added to the particular document by another user. In some embodiments, the request provided via the user input directs the automated assistant to access or modify the supplemental content added to the particular document by another user. In some embodiments, causing the one or more actions to be performed includes: performing speech-to-text processing to convert a portion of the user input into text data, and causing the text data to be incorporated into a portion of the particular document corresponding to a particular semantic annotation of the one or more semantic annotations. In some embodiments, the method may further include generating the one or more semantic annotations using a trained machine learning model, the trained machine learning model being trained using training data based on previous user interactions between the user and other portions of various different documents.
[0081] In other embodiments, a method implemented by one or more processors is described as including operations such as: receiving a request corresponding to a spoken utterance from a user at an automated assistant interface of a computing device; wherein the computing device provides access to the automated assistant. The method may also include the following operations: identifying natural language content from a portion of a specific document based on the request, wherein the portion of the specific document was not present in the foreground of a graphical user interface of the computing device when the user provided the spoken utterance. The method may also include the following operations: determining one or more specific actions that the user is requesting the automated assistant to perform based on the natural language content from the portion of the specific document. The method may also include the following operations: causing execution of an action of the one or more actions to be initialized based on the request.
[0082] In some embodiments, causing the execution of an action includes causing the automated assistant to audibly render the natural language content from the portion of the specific document. In some embodiments, the method may further include the following operations: after initiating execution of an action in the one or more actions: receiving an additional request corresponding to an additional spoken utterance from the user at the automated assistant interface of the computing device, and determining, based on the additional spoken utterance and the natural language content from the portion of the specific document, that the user is requesting the automated assistant to edit the portion of the specific document. In some embodiments, the method may further include the following operations: after initiating execution of an action in the one or more actions: receiving an additional request corresponding to an additional spoken utterance from the user at the automated assistant interface of the computing device, and determining, based on the additional spoken utterance and the natural language content from the portion of the specific document, that the user is requesting the automated assistant to communicate with another user. In some embodiments, the other user adds the natural language content to the specific document before the user provides the additional request.
[0083] In yet another embodiment, a method implemented by one or more processors is described as receiving a request from an automated assistant at an application, wherein the request is provided by the automated assistant in response to a first user providing user input to the automated assistant via a first computing device, and wherein the automated assistant is responsive to natural language input provided by the first user to an interface of the first computing device. The method may also include the following operations: modifying a document by the application in response to receiving the request from the automated assistant, wherein the document is editable by the first user via the first computing device and editable by a second user via a second computing device different from the first computing device. The method may also include the following operations: generating notification data indicating that the document has been modified by the first user based on modifying the document. The method may also include the following operations: using the notification data to cause an additional automated assistant to render a notification for the second user via the second computing device, wherein the additional automated assistant is responsive to other natural language input provided by the second user to a separate interface of the second computing device.
[0084] In some embodiments, the request includes a description of a sub-portion of the document, and modifying the document includes comparing the description to a plurality of different semantic annotations stored associated with the document. In some embodiments, comparing the description to the plurality of different semantic annotations stored associated with the document includes assigning a similarity score to each of the plurality of different semantic annotations, wherein a particular similarity score for a respective semantic annotation indicates a degree of similarity between the respective semantic annotation and the description. In some embodiments, causing the additional automated assistant to render the notification for the second user includes causing a graphical rendering of the sub-portion of the document to be rendered at the separate interface of the second computing device. In some embodiments, the application is provided by an entity that is distinct from one or more other entities that provide the automated assistant and the additional automated assistant. In some embodiments, the automated assistant and the additional automated assistant communicate with the application via an application programming interface.
Claims
1. A method implemented by one or more processors, the method comprising: receiving user input directed to the automated assistant from a user at an automated assistant interface of the computing device, wherein the user input comprises a request for the automated assistant to access or modify a document; In response to receiving the user input, identifying the document that the user is requesting to access or modify, wherein the document is stored at the computing device or another computing device, wherein an explicit verbalization of the corresponding name of the document is omitted from the user input, and Wherein, identifying the document comprises: processing data, the data comprising natural language content input by the user and content of each of a plurality of different documents accessible via the computing device; determining one or more actions to be performed on the document, wherein determining the one or more actions is based on the user input and one or more semantic annotations of the document stored in association with the document, and wherein each of the one or more semantic annotations comprises a semantic interpretation of a corresponding sub-portion of the entirety of the document; causing the one or more actions to be performed to access or modify the document based on the user input; and Prior to receiving the user input at the automated assistant interface of the computing device: Another user input is received, the other user input comprising another request for the automated assistant to render a description of supplemental content added to the document by another user.
2. The method according to claim 1, in, the document was not originally created by the user, and The document is created using a document application different from the automated assistant.
3. The method according to claim 1, wherein The data further comprises additional semantic annotations, and each of the additional semantic annotations comprises another semantic interpretation of another respective sub-portion of a respective additional document of the plurality of different documents.
4. The method according to claim 1, in, The automated assistant interface of the computing device includes a microphone, and The user input is received when a document editing program used to edit the document is not present in the foreground of a graphical user interface of the computing device.
5. The method according to claim 1, in, The one or more semantic annotations include a semantic annotation containing a semantic interpretation of a comment on the document created by an additional user, and Wherein the one or more actions include causing a notification to be provided to the additional user via another interface of a separate computing device associated with the additional user.
6. The method according to claim 1, wherein The request provided via the user input directs the automated assistant to access or modify the supplemental content added to the document by the other user.
7. The method according to claim 1, wherein Causing the one or more actions to be performed includes: performing speech-to-text processing to convert said portion of the user input into text data, and The text data is caused to be incorporated into a portion of the document corresponding to a semantic annotation of the one or more semantic annotations.
8. The method according to any one of claims 1 to 7, further comprising: The one or more semantic annotations are generated using a trained machine learning model, the trained machine learning model being trained using training data based on previous user interactions between the user and other portions of various different documents.
9. A method implemented by one or more processors, the method comprising: receiving, at an automated assistant interface of a computing device, a request corresponding to a spoken utterance from a user; wherein the computing device provides access to an automated assistant; identifying natural language content from a portion of the document based on the request, wherein, when the user provides the spoken utterance, the portion of the document is not present in the foreground of a graphical user interface of the computing device; determining one or more actions that the user is requesting the automated assistant to perform based on the natural language content from the portion of the document; causing performance of an action of the one or more actions to be initiated based on the request; and After initiating execution of an action in the one or more actions: receiving, at the automated assistant interface of the computing device, additional requests corresponding to additional spoken utterances from the user, and A determination is made based on the additional spoken utterance and the natural language content from the portion of the document that the user is requesting the automated assistant to edit the portion of the document.
10. The method according to claim 9, wherein: The execution of the action includes: The automated assistant is caused to audibly render natural language content from the portion of the document.
11. The method according to any one of claims 9 to 10, further comprising: After initiating execution of an action in the one or more actions: receiving, at the automated assistant interface of the computing device, additional requests corresponding to additional spoken utterances from the user, and A determination is made based on the additional spoken utterance and the natural language content from the portion of the document that the user is requesting the automated assistant to communicate with another user.
12. The method according to claim 11, wherein Before the user provides the append request, the other user adds the natural language content to the document.
13. A method implemented by one or more processors, the method comprising: Receive requests from automated assistants at the application, wherein the request is provided by the automated assistant in response to a first user providing user input to the automated assistant via a first computing device, and wherein the automated assistant is responsive to natural language input provided by the first user to an interface of the first computing device; modifying the document by the application in response to receiving the request from the automated assistant, wherein the document is editable by the first user via the first computing device and editable by a second user via a second computing device different from the first computing device; generating notification data indicating that the document has been modified by the first user based on modifying the document; and using the notification data to cause an additional automated assistant to render a notification for the second user via the second computing device, The additional automated assistant is responsive to other natural language input provided by the second user to a separate interface of the second computing device.
14. The method according to claim 13, wherein The request includes a description of a sub-portion of the document, and modifying the document includes: The description is compared to a plurality of different semantic annotations stored in association with the document.
15. The method according to claim 14, wherein Comparing the description to a plurality of different semantic annotations stored in association with the document includes: assigning a similarity score to each of the plurality of different semantic annotations; The corresponding similarity score of the corresponding semantic annotation indicates the degree of similarity between the corresponding semantic annotation and the description.
16. The method according to claim 13, wherein: Causing the additional automated assistant to render the notification for the second user includes: A graphical rendering of the sub-portion of the document is caused to be rendered at the separate interface of the second computing device.
17. The method according to any one of claims 13 to 16, wherein The application is provided by an entity that is distinct from one or more other entities that provide the automated assistant and the additional automated assistant.
18. The method according to claim 17, wherein The automated assistant and the additional automated assistant communicate with the application via an application programming interface.
19. A method implemented by one or more processors, the method comprising: receiving, at an automated assistant interface of a computing device, a request corresponding to a spoken utterance from a user; wherein the computing device provides access to an automated assistant; identifying natural language content from a portion of the document based on the request, wherein, when the user provides the spoken utterance, the portion of the document is not present in the foreground of a graphical user interface of the computing device; determining one or more actions that the user is requesting the automated assistant to perform based on the natural language content from the portion of the document; causing performance of an action of the one or more actions to be initiated based on the request; and After initiating execution of an action in the one or more actions: receiving, at the automated assistant interface of the computing device, additional requests corresponding to additional spoken utterances from the user, and A determination is made based on the additional spoken utterance and the natural language content from the portion of the document that the user is requesting the automated assistant to communicate with another user.
20. The method according to claim 19, wherein The execution of the action includes: The automated assistant is caused to audibly render natural language content from the portion of the document.
21. A system comprising one or more processors and a memory operatively coupled to the one or more processors, wherein: The memory stores instructions which, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 20.
22. At least one non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1 to 20.
Citation Information
Patent Citations
Methods for searching with semantic similarity scores in one or more ontologies
US20110040766A1
Personal Assistant for Task Utilization
US20110314375A1
System and management of semantic indicators during document presentations
US20200159838A1
Providing access to user-controlled resources by automated assistants
US20200184156A1