Compose Assistant Manager for Applications

JP2026530461APending Publication Date: 2026-09-08GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026512384
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-25
Filing Date
2024-08-26
Publication Date
2026-09-08

Smart Images

  • Figure 2026530461000001_ABST
    Figure 2026530461000001_ABST
Patent Text Reader

Abstract

An application may receive prompts from the user related to inputting into text fields of digital content displayed on the user's device. The application may generate contextual data about the digital content. The application may provide prompts and contextual data to a generative language model. The application may receive responses generated by the generative language model and provide these responses as suggestions for input into text fields.
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 578,816, filed on August 25, 2023, the entire disclosure of which is incorporated herein by reference. [[Background Art]]

[0002] Some web pages include text boxes that obtain text input from users. Examples include web pages that allow users to leave reviews of products, services, locations, etc., web pages that allow users to leave comments or reply to comments, web pages that allow users to post messages (e.g., web pages of social media websites), and / or web pages including surveys, etc. Users may use generative language models to help draft input content for web pages. However, users may need to specify terms relatively concretely when drafting prompts, and / or may need to repeatedly iterate the language model multiple times to generate a desired review. Furthermore, obtaining context data from web content for use in generative language models may cause one or more technical problems related to security. [[Summary of Invention]]

[0003] This disclosure relates to a compose assistant manager for an application (e.g., a browser application) that integrates a generative model (e.g., a language model) for drafting content as input to a text field of digital content (e.g., a web page), providing one or more technical benefits such as maintaining the security of the application content (e.g., a web page) and / or reducing the amount of computing resources (e.g., memory, CPU) consumed to generate the generated content and insert it (e.g., directly) into the text field of the digital content. The compose assistant manager may provide the user with reduced overhead when creating prompts and may adapt the generated output to the context of the digital content. The compose assistant manager may generate one or more context signals (also called context data) about the digital content (e.g., a web page), and the compose assistant manager may send the text data (e.g., also called prompts or user-provided prompts) and content signals received from the user to the generative language model, which returns a model response that can be inserted directly into the text field. In other words, the Compose Assistant Manager helps users enter text into text fields provided by the computer system, using contextual information about the web page, which may be technical information, specifically content from the web page.

[0004] In some embodiments, the techniques described herein include a method for receiving text data from a user relating to input into a text field of digital content displayed on a user device; generating contextual data relating to the digital content; providing the text data and contextual data to a generative language model; receiving a response generated by the generative language model; and providing the response as a suggestion for input into the text field.

[0005] In some embodiments, the technology described herein relates to an apparatus comprising at least one processor and a non-temporary computer-readable medium storing executable instructions causing the at least one processor to perform an operation, wherein the operation includes receiving text data from a user relating to input into a text field of digital content displayed on a user device, generating contextual data relating to the digital content, providing the text data and contextual data to a generative language model, receiving a response generated by the generative language model, and providing the response as a suggestion for input into the text field.

[0006] In some embodiments, the technology described herein relates to a non-temporary computer-readable medium storing executable instructions causing at least one processor to perform an operation, wherein the operation includes receiving text data from a user relating to input into a text field of digital content displayed on a user device; generating contextual data relating to the digital content; providing the text data and contextual data to a generating language model; receiving a response generated by the generating language model; and providing the response as a suggestion for input into the text field.

[0007] Details of one or more embodiments are described in the accompanying drawings and the description below. Other features will become apparent from the description and drawings. [Brief explanation of the drawing]

[0008] [Figure 1A] An exemplary callout affordance for invoking the Compose Assistant Manager is shown in one aspect. [Figure 1B] An exemplary callout affordance for invoking the Compose Assistant Manager is shown in one aspect. [Figure 1C]One embodiment of a compose assistant interface for receiving prompts is shown. [Figure 1D] One embodiment of a composition assistant interface for displaying a model response is shown. [Figure 1E] An example of a text field into which a model response is input is shown in one aspect. [Figure 1F] One embodiment describes a system having a compose assistant manager for a browser application that integrates a language model for drafting content as input to a text field on a web page. [Figure 1G] An example of a context signal for generating a model response, according to one embodiment, is shown. [Figure 1H] An example of a trigger engine in one embodiment is shown. [Figure 1I] An example of a web page with embedded resources, according to one embodiment, is shown. [Figure 2] An example of a composition assistant interface according to one embodiment is shown. [Figure 3] Examples of other forms of compose assistant interfaces are shown. [Figure 4A] This illustrates a composition assistant interface rendered on a social media webpage, according to one embodiment. [Figure 4B] This illustrates a composition assistant interface rendered on a social media webpage, according to one embodiment. [Figure 4C] This illustrates a composition assistant interface rendered on a social media webpage, according to one embodiment. [Figure 5A] This document illustrates various aspects of a composition assistant interface. [Figure 5B] This document illustrates various aspects of a composition assistant interface. [Figure 5C]This document illustrates various aspects of a composition assistant interface. [Figure 5D] This document illustrates various aspects of a composition assistant interface. [Figure 5E] This document illustrates various aspects of a composition assistant interface. [Figure 5F] This document illustrates various aspects of a composition assistant interface. [Figure 6] Examples of other forms of compose assistant interfaces are shown. [Figure 7] Examples of other forms of compose assistant interfaces are shown. [Figure 8] This figure shows a computing system and server components for implementing the concept described herein, according to one embodiment. [Figure 9] A flowchart shows an exemplary process for providing a compose assistant manager according to one embodiment. [Figure 10] This flowchart shows an exemplary process for providing a compose assistant manager in another embodiment. [Figure 11] This flowchart shows an exemplary process for providing a compose assistant manager in another embodiment. [Modes for carrying out the invention]

[0009] The present disclosure relates to a compose assistant manager for an application (e.g., a browser application) that integrates a generative model (e.g., a language model) for generating text field content and inserting (e.g., directly inserting) the content into a text field, which can provide one or more technical advantages of maintaining web page security and / or reducing the amount of computing resources (e.g., memory, CPU) for generating generated content and inserting the generated content into one or more text fields of web content. The compose assistant manager can assist a user with leaving a review, commenting on an article, responding to a survey, drafting a social media post, filling out a customer complaint, and / or responding to a chatbot, among others.

[0010] In some examples, a user can explicitly invoke the compose assistant manager. For example, by right-clicking a text field of a web page and selecting a menu option (e.g., an option for supporting text composition), the user can cause a compose assistant interface to be displayed to receive a prompt for the generative model. In some examples, the text field can be any type of input field configured to receive text from a user (e.g., via a keyboard, voice, a touch screen, or the like), and text received by the user is input into the text field. In some examples, the text field is a free-form text field. In some examples, the text field is a structured text field. In some examples, the text field is a multi-line text field. In some examples, the text field is a single-line text field. In some examples, the text field can be input with data received via a microphone (e.g., a voice assistant).

[0011] A user can provide a prompt to a composition assistance interface (for example, "write a five-star review for this product"). For example, the composition assistance interface includes an input field that allows a user to draft a natural language description of the prompt, for example, the type of content generated by a generative model. In response to submission of the prompt, the composition assistance manager may transmit the prompt and one or more context signals (also referred to as context data) relating to the underlying web page. The context data may include information about the subject matter of the web page. In response to the prompt and the context data, the generative model can generate and return a context-appropriate response, which can be inserted directly into the text box of the web page.

[0012] The composition assistance manager provides a technical solution for generating context data (for example, one or more context signals) relating to the underlying web page, and the context signals help create a context-appropriate response using the generative model. In some examples, the context data includes a resource locator, a page title, page content, a document object model (DOM) representation, and / or an accessibility content structure (for example, an accessibility tree). In some examples, the composition assistance manager can retrieve first page content of a web page having a text field, can retrieve second page content of one or more embedded web pages, and can include the first page content and the second page content in the context signal. In some examples, the web page includes one or more inline frames (for example, iframes). An iframe is a Hypertext Markup Language (HTML) element that embeds another HTML document within the current page. Retrieving page content from embedded web pages may cause one or more technical challenges with respect to maintaining security.

[0013] However, the Compose Assistant Manager overcomes technical challenges by performing context extraction of context signals, not only by requesting the internal text of a specified host, but also by requesting the internal text of local same-origin iframes (e.g., all local same-origin iframes). A same-origin iframe can be an iframe that shares the same origin as the main webpage (e.g., a frame embedded within a webpage). The origin of a webpage is determined by its protocol, hostname, and port number. In some cases, the embedded iframe resides on the same server or domain as the main webpage. In some cases, the embedded iframe may have the same protocol, hostname, and / or port number as the main webpage. Internal text can refer to visible text content within an HTML element and text from one or more child elements of an HTML element. The returned internal text includes the combined internal text of the iframe (e.g., all iframes). The Compose Assistant Manager retrieves the internal text of the webpage, and the internal text of the webpage is combined with the internal text of the iframe (e.g., the embedded webpage) whenever an iframe is detected. The Compose Assistant Manager provides context signals and user-provided prompts to the generation model. The generation model generates a model response and returns it to the Compose Assistant Manager, which can then directly insert the model response into an input text box (for example, with or without user prompts).

[0014] The Compose Assistant Manager may display model responses in text fields. The Compose Assistant interface may include one or more UI elements that allow the user to adjust the model response (e.g., more formal, less formal, augmented, shortened, etc.), which causes the generated model to regenerate the model response. In some examples, the user may manually edit the model response. The Compose Assistant Manager may include an insert control, which, when selected, inserts the model response into a text field on a web page. For example, in response to a selection of the insert control, the Compose Assistant Manager transfers the text from an input field in the Compose Assistant interface to a text field on a web page.

[0015] In some examples, the compose assistant interface may provide the user with one or more suggestion prompts, which the user can select and / or edit. In other words, before the user begins drafting a prompt in the input field of the compose assistant interface, the compose assistant interface may provide selectable suggestion prompts, and when a suggestion prompt is selected, the suggestion prompt is entered into the input field of the compose assistant interface. These suggestion prompts may be based on context signals obtained from a web page. For example, before the user submits a prompt, the compose assistant manager generates and provides a prompt suggestion request with one or more context signals to the generative model, which returns one or more suggestion prompts to be displayed in the compose assistant interface. In some examples, the suggestion prompts are selectable elements in the compose assistant interface. In some examples, in response to the selection of a suggestion prompt, the compose assistant manager may send the selected (suggested) prompt and context signals to the generative model.

[0016] In some cases, the Compose Assistant Manager may selectively trigger the display of a callout affordance, and the user can interact with the callout affordance to invoke the Compose Assistant interface. For example, instead of the user directly invoking the Compose Assistant Manager (e.g., by selecting a menu item associated with the Compose Assistant Manager), the Compose Assistant Manager may selectively display a callout affordance that informs the user about the Compose Assistant Manager, which can help draft the content of a text field. The callout affordance can be a UI object displayed on a web page located near the text field. In response to a user selection on the callout affordance (or a control on the callout affordance), the Compose Assistant Manager may display the Compose Assistant interface to allow the user to submit a prompt to the generative model for creating the content of the text field.

[0017] The Compose Assistant Manager may decide whether and / or when to display callout affordances (or, in some examples, the Compose Assistant interface). The Compose Assistant Manager may include a heuristic and / or machine learning (ML) model that receives one or more signals and decides whether or not to render callout affordances on the web page based on those signals. In some examples, signals may include signals about text fields on the web page, signals about page content, and / or signals about previous use of the Compose Assistant interface on the web page. In some examples, previous use signals may include one or more signals about whether a user has previously used the Compose Assistant interface (and / or has previously denied permission for the Compose Assistant interface), and / or whether another user has previously used the Compose Assistant on that particular text field.

[0018] In some examples, the generative model is a machine learning (ML) model. In some examples, the generative model is a pre-trained large-scale language model (LLM). In some examples, the generative model is a specially trained language model. The generative model can generate high-quality responses to text fields. In some examples, the generative model can be trained to generate responses to specific categories (types) of text fields. The generative model uses contextual signals from a web page to generate the content of a text field. The generative model can use contextual signals from a web page to determine the categories associated with a text field (e.g., the categories the text field represents). In some examples, a generative model specially trained to generate responses to specific categories of text fields may be smaller (e.g., in terms of required CPU and memory) and faster computationally (e.g., generating responses in a short time, such as 5 or 10 seconds) than a general-purpose large-scale language model, and may produce more relevant and high-quality responses that satisfy expectations for the categories of the text field. Such relevant and appropriate responses minimize user interaction for generating responses and provide a well-guided human-machine process for generating content.

[0019] Figures 1A to 1I show a system 100 having a Compose Assistant Manager 110 of a browser application 108 that assists the user in generating content for one or more text fields 136 of a web page 134. The Compose Assistant Manager 110 can initiate a generation model 152 to generate a model response 124 for the text fields 136 of the web page 134, and then insert the model response 124 into the text fields 136 (e.g., by direct input). For example, a user may interact with the Compose Assistant Manager 110 to help draft the content for the text fields 136 of the web page 134. In some examples, the web page 134 may be called digital content. The term digital content may include web content, and in some examples, may include non-web content.

[0020] A browser application 108 executable by the user device 102 may render a web page 134 on the display 126, as shown in Figures 1A and 1F. While the example in Figure 1A depicts a web page for writing a review, the web page 134 can be any type of web page 134. Furthermore, the techniques described herein are not limited to the browser application 108 and may apply to any application capable of rendering web content, or in some examples, non-web content. The web page 134 includes a text field 136 configured to receive text input from the user. In some examples, the text field 136 includes a free-form input field. A free-form input field includes an input field that receives unrestricted input from the user. In some examples, the text field 136 includes an input field that receives structured data. In some examples, the text field 136 includes a multi-line input field. In some examples, the text field 136 includes a single-line input field.

[0021] To access the functionality of the Compose Assistant Manager 110, the Compose Assistant Manager 110 includes a trigger engine 112 configured to render a callout affordance 138 on the display 126 of the user device 102. The callout affordance 138 may be a user interface (UI) element, an object, a menu item, or a control that identifies the Compose Assistant Manager 110. In some examples, the callout affordance 138 may be directly accessed by the user using one or more controls provided by the browser application 108. For example, as shown in Figure 1B, the trigger engine 112 may render the callout affordance 138 as a menu item 138b (e.g., "Support for Writing") from menu 111. In some examples, the user may right-click on a text field 136 on a web page 134, and the browser application 108 may display a menu 111 (e.g., a right-click menu) near the text field 136, as shown in Figure 1B. Menu 111 may include menu item 138b, which, when selected, renders the compose assistant interface 128 as shown in Figures 1A and 1D.

[0022] In some examples, the trigger engine 112 may selectively trigger the display of a callout affordance 138. For example, as shown in Figure 1A, the trigger engine 112 may display the callout affordance 138 as a selectable UI object 138a. In some examples, the trigger engine 112 may detect a user interaction with the text field 136 (for example, the user focuses on the text field 136, such as by placing the cursor over the text field 136), and in response to the detected interaction, the trigger engine 112 may render the selectable UI object 138a. The user selection of the selectable UI object 138a causes the compose assistant manager 110 to render the compose assistant interface 128, as shown in Figures 1C and 1F.

[0023] In some examples, the trigger engine 112 may decide whether and / or when to display a callout affordance 138 (or, in some examples, the compose assistant interface 128). In some examples, the trigger engine 112 may detect a trigger event for displaying the compose assistant interface 128 based on one or more signals 180. In some examples, as shown in Figure 1H, the trigger engine 112 includes a machine learning (ML) model 114 configured to receive a signal 180 and compute a prediction 188 on whether to display the callout affordance 138 (e.g., the selectable UI object 138a in Figure 1B). In some examples, the trigger engine 112 uses one or more heuristics with the signal 180(or) to aggressively render the callout affordance 138 (e.g., the selectable UI object 138a in Figure 1B). In some examples, the trigger engine 112 uses a combination of heuristics and ML predictions to decide whether to display the callout affordance 138.

[0024] In some examples, signal 180 includes a text field signal 182 (for example, a signal about text field 136 on web page 134), a content signal 184 (for example, a signal about page content), and / or a previous usage signal 186 (for example, a signal about previous usage of Compose Assistant Manager 110). In some examples, the previous usage signal 186 may include one or more signals indicating whether a user has previously used Compose Assistant Manager 110 (and / or has previously not allowed Compose Assistant Manager 110), and / or whether another user has previously used Compose Assistant on that particular text field 136 or web page 134.

[0025] Heuristics may include the results of existing autofill functionality. For example, a browser application 108 may include an autofill functionality for a text field 136 that already uses several heuristics to identify a target text field that is important to its purpose. A heuristic for proactively triggering a callout affordance 138 may be applied when the autofill functionality does not trigger a suggestion (for example, when the autofill functionality does not determine a text field 136 that has the appropriate focus for an autofill suggestion). A heuristic may include the fact that the web page 134 is in a supported language. A heuristic may include the fact that the compose assistant manager 110 is not suppressed for reasons that are supported (for example, the functionality is disabled by the user, or the web page 134 or website (domain) is considered a policy violation). A heuristic may include the fact that the use of the compose assistant manager 110 does not conflict with other browser functionality. A heuristic may include the fact that the text field 136 is not related to corporate or work productivity documents (e.g., word processing documents, slide decks, etc.). The heuristic may include the fact that text field 136 is not a prompt input box for a large language model (e.g., a text box designed to provide prompts (queries) sent to a large language model). The heuristic may also consider past user history (e.g., stored locally on the user's device) based on user permissions. For example, if a user is using Compose Assistant Manager 110 on a review website but ignores callout affordances 138 on social media sites, the heuristic may allow trigger engine 112 to render callout affordances 138 for text fields related to product / service reviews but not for web pages related to social media.

[0026] The trigger engine 112 may actively render the callout affordance 138 using one or more heuristics in any combination. In some examples, the trigger engine 112 may actively render the callout affordance 138 using one or more heuristics in any combination in response to detecting a user interaction with the text field 136 (e.g., a focus applied to the text field 136). In some examples, the trigger engine 112 may actively render the callout affordance 138 using one or more heuristics in any combination without detecting a user interaction with the text field 136 (e.g., without a focus applied to the text field 136). In some examples, the trigger engine 112 may render the callout affordance 138 in response to the amount of text data entered by the user into the text field 136 reaching a threshold level. When selected, the callout affordance 138 is configured to render the compose assistant interface 128 to the text field 136, and the compose assistant interface 128 has an input field 130 configured to receive a prompt 118 from the user.

[0027] In some examples, as described above, the trigger engine 112 may include (or communicate with) an ML model 114 to generate a prediction 188 about whether to render the callout affordance 138. If the prediction 188 includes the probability that the user is likely to use the Compose Assistant Manager 110, the trigger engine 112 may render the callout affordance 138. In some examples, the ML model 114 may be trained with one or more (or any combination thereof) of the heuristics described herein to determine whether and when to trigger the callout affordance 138. For example, if the probability is high (meets a first threshold), the trigger engine 112 may trigger the callout affordance 138 (e.g., when the text field 136 receives focus). If the probability is neither high nor low (does not meet the first threshold but meets the second threshold), the trigger engine 112 may trigger the callout affordance 138 when the user types some characters or words into the text field 136 and then stops.

[0028] Referring to Figure 1F, the Compose Assistant Manager 110 includes a Prompt Manager 116. The Prompt Manager 116 generates a context signal 120 about the web page 134. The context signal 120 may be called context data. The context data contains information about the subject of the web page 134. In some examples, the Prompt Manager 116 generates a context signal 120 in response to the Compose Assistant Manager 110 being invoked (e.g., when a callout affordance 138 is selected and / or when the Compose Assistant Interface 128 is rendered). In some examples, the Prompt Manager 116 generates a context signal 120 after the callout affordance 138 is rendered (e.g., UI object 138a) and before the callout affordance 138 is selected. In some examples, the Prompt Manager 116 generates a context signal 120 in response to the selection of a generate control 131 on the Compose Assistant Interface 128.

[0029] As shown in Figure 1G, the context signal 120 may include the page title 172 of the web page 134, the page content 170 associated with the web page 134, and / or the resource locator 176 of the web page 134. In some examples, the context signal 120 includes a DOM representation 178. In some examples, the context signal 120 includes an accessible content structure 174. The accessible context structure 174 may be called an accessible tree.

[0030] The prompt manager 116 provides a technical solution to generate context signals 120 about the underlying web page 134, which then use the generative model 152 to create a context-appropriate response. The prompt manager 116 performs context extraction to extract the page content of the web page 134 in a way that preserves the security of the web page 134.

[0031] In some examples, as shown in Figure 1I, page content 170 includes page content 170-1 of a web page 134 (e.g., a first web page) and page content 170a of one or more embedded resources 139 embedded in the structure of the web page 134. For example, the prompt manager 116 can retrieve page content 170-1 of a web page 134 having a text field 136, and page content 170a of one or more embedded resources 139 (e.g., web pages). For example, web page 134 may have resource 139-1 (e.g., a second web page) and resource 139-2 (e.g., a third web page) embedded in it. The prompt manager 116 can retrieve page content 170-1 of web page 134, page content 170-2 of resource 139-1, and page content 170-3 of resource 139-2. In some examples, page content 170-1, page content 170-2, or page content 170-3 may be referred to as internal text. Extracting page content from embedded resources 139 (e.g., web pages) can introduce one or more technical challenges, such as security risks.

[0032] In other words, webpage 134 contains one or more inline frames (e.g., iframes) (e.g., hypertext markup language (HTML) elements that embed other HTML documents (e.g., resource 139-1 or resource 139-2) within the current page (e.g., webpage 134)). The prompt manager 116 performs context extraction of the context signal 120, overcoming technical challenges by requesting the internal text of the specified host (e.g., webpage 134) and the internal text of local same-origin iframes (e.g., all local same-origin iframes). A same-origin iframe may be an iframe that shares the same origin as the main webpage (e.g., webpage 134). The origin of a webpage is determined by its protocol, hostname, and port number. In some examples, the embedded iframe resides on the same server or domain as the main webpage. In some examples, the embedded iframe may have the same protocol, hostname, and / or port number as the main webpage. Internal text can refer to visible text content within an HTML element, and text from one or more child elements of an HTML element. The returned internal text includes the combined internal text of an iframe (e.g., all suitable iframes). The prompt manager 116 retrieves the internal text of the web page 134, and the internal text of the web page is combined with the internal text of the iframe (e.g., embedded resource 139) whenever an iframe is detected.

[0033] Referring to Figure 1F, in some examples, the prompt manager 116 includes an ML model 122. The ML model 122 may receive context signals 120 as input, such as the page title 172 of the web page 134, the page content 170 associated with the web page 134, the resource locator 176 of the web page 134, the DOM representation 178, and the accessible content structure 174. The ML model 122 may use the context signals 120 to generate context data (or use the first context data (e.g., a larger set of content data) to generate a second context data (e.g., a smaller set of content data)), and the context data generated by the ML model 122 will include a subset of information smaller than the context signals 120, and the context data will be provided to the generative model 152. In some examples, the ML model 122 selects a subset of information contained in the context signals 120, and this subset is provided to the generative model 152. In some examples, the ML model 122 generates a summary of the context signal 120, which is then provided to the generative model 152. By using the ML model 122 to generate or select a portion of the context signal 120, a smaller set of information may be provided to the generative model 152, which can result in one or more technical advantages, such as a reduction in the computational cost of the inference computation by the generative model 152. In other words, the token size of the prompts provided to the generative model 152 can be reduced, which reduces the computational cost of generating the model response 124.

[0034] Referring to Figures 1C and 1F, the compose assistant interface 128 includes an input field 130 configured to receive a prompt 118 from the user. In some examples, the prompt 118 is called text data, e.g., data entered by the user. The user can provide the compose assistant interface 128 with a prompt 118 (e.g., "Write a 5-star review about this product"). The prompt 118 can be a natural language description of the type of content generated by the generative model 152. For example, the user can type the prompt 118 or provide a voice command to insert the prompt 118 into the input field 130.

[0035] Referring to Figure 1C, the compose assistant interface 128 may include a generation control 131. In response to a user selection in the generation control 131, the prompt manager 116 may send a prompt 118 and a context signal 120 to the generation model 152. In some examples, the generation control 131 may remain inactive until the user provides text in the input field 130. Thus, after the user enters text in the input field 130, the generation control 131 may become active (and selectable). In some examples, the compose assistant interface 128 may include an option (e.g., a three-dot menu) to enable or disable the compose assistant manager 110.

[0036] In response to prompt 118 and context signal(s) 120, the generating model 152 may generate a model response 124. As shown in Figure 1D, the prompt manager 116 may receive the model response 124 from the generating model 152 and display the model response 124 on interface 133 of the compose assistant interface 128.

[0037] As shown in Figure 1D, the compose assistant interface 128 may include one or more UI elements that allow the user to adjust the model response 124 (e.g., more formal, less formal, extended, shortened, etc.), thereby causing the generated model 152 to regenerate the model response 124. In some examples, the user may manually edit the model response 124. Referring to Figure 1D, the compose assistant manager 110 may include an insert control 141, which, when selected, inserts the model response 124 into a text field 136 of the web page 134, as shown in Figures 1E and 1F. For example, in response to a selection of the insert control 141, the compose assistant manager 110 transfers the text from the compose assistant interface 128 to the text field 136 of the web page 134.

[0038] As shown in Figure 1D, the compose assistant interface 128 may include an insert control 141. The insert control 141 inserts text into the text field 136, replacing any text previously written by the user (rather than starting from scratch) if the user has previously written text. If the user has written text, the insert control 141 may display "Replace"; if not, the insert control 141 may display "Insert This". If the user clicks Replace but selects only a portion of the text (for example, by highlighting part of the text), only that selected text (and not all of the text in the field) is replaced. In some examples, the insert control 141 may close the compose assistant interface 128. If the user closes the compose assistant interface 128, for example by selecting the close control 127 before selecting the generate control 131, this may be a local signal used by the individual heuristic, as described above. In other words, at user discretion, the trigger engine 112 may use this type of termination event to determine when to actively display the callout affordance 138. Figure 1E shows the model response 124 inserted into text field 136 of web page 134. The user can edit the response in text field 136 of web page 134.

[0039] The compose assistant interface 128 may also include controls for modifying (editing) the model response 124 using the generated model 152. For example, as shown in Figure 1D, the compose assistant interface 128 may include a tone control 123. The tone control 123 may allow the user to make the response sound more formal, more casual, more humorous, or include emojis. The compose assistant interface 128 may also include a length control 125. The length control 125 may allow the user to shorten or lengthen the model response 124. In response to the user selecting either the tone control 123 or the length control 125, the compose assistant manager 110 may provide a new model response 124. In other words, the selection of the tone control 123 or the length control 125 may cause the generated model 152 to regenerate the model response 124 based on the value of the selected control. In some examples, if the user selects the "Longer" length control 125 multiple times, the Compose Assistant Manager 110 may suggest that the user use a general-purpose large-scale language model (such as Bard, ChatGPT®) to improve the interaction experience.

[0040] In some examples, the compose assistant interface 128 may include a regenerate control 135, as shown in Figure 1D. The regenerate control 135 may generate other text suggestions (e.g., a new model response 124). The compose assistant interface 128 may include a back control 147. The back control 147 may allow the user to return to the prompt 118 of the compose assistant interface in Figure 1D for editing. The compose assistant interface 128 may include a close control 115. The close control 115 may close the compose assistant interface 128 and return the user to the text field 136. In some examples, the compose assistant manager 110 may store information about the use of the compose assistant manager 110 on a web page 134 (e.g., a particular text field 136) to help determine whether to render a callout affordance 138 to that user or other users in the future.

[0041] In some examples, the compose assistant interface 128 includes a feedback mechanism 129. The feedback mechanism 129 may allow the user to rate text suggestions. The rating can be used, at the user's discretion, for additional training (e.g., a thumbs-down or low rating can be used as an example of something not to produce in the prompt). The rating can also be used, at the user's discretion, to trigger the compose assistant manager 110 for this user. Thus, in some embodiments, the user can rate the suggested text output to improve future suggestions. Figure 1E shows a binary (thumbs-up / thumbs-down) feedback mechanism 129, but numerical scales may also be used (e.g., a number of stars, selecting one of several ratings, etc.).

[0042] Referring to Figure 1F, the compose assistant interface 128 may provide the user with one or more suggestion prompts 118a, which the user can select and / or edit. In other words, before the user begins drafting a prompt 118 in the input field 130 of the compose assistant interface 128, the compose assistant manager 110 may provide selectable suggestion prompts 118a, and if a suggestion prompt 118a is selected, the suggestion prompt 118a is entered into the input field 130 of the compose assistant interface 128. These suggestion prompts 118a may be based on context signals 120 obtained from the web page 134. For example, before the submission of a user prompt 118, the compose assistant manager 110 generates a prompt suggestion request with one or more context signals 120 and provides it to the generative model 152, which returns one or more suggestion prompts 118a to be displayed in the compose assistant interface 128. In some examples, the suggestion prompts 118a are selectable elements in the compose assistant interface 128. In some examples, in response to the selection of a suggestion prompt 118a, the compose assistant manager 110 may send the selected (suggestion) prompt 118a and a context signal 120 to the generation model 152.

[0043] In some examples, suggest prompt 118a is a general prompt to show the user that the compose assistant manager 110 can help the user write. In some examples, if the user has not yet started writing and calls the compose assistant manager 110, the compose assistant manager 110 may render a set of rotating suggest prompts 118a based on context signals 120 (for example, there may be about five different sets of prompts depending on whether the user is writing a review, writing a social media caption, or filling out a form). Thus, suggest prompt 118a may use page context or be generic. Page context may include values ​​the user has given to other fields, e.g., the number of stars the user has already given. Page context may include insights from other content on the web page. In some examples, the compose assistant manager 110 may analyze the user's writing history for profile 155 (generated with user permissions) (e.g., local profile) and / or open tabs to provide personalized prompts to ensure relevance and resonance with the user's target audience. Relying on user history can help maintain consistency in the tone of responses generated for the user.

[0044] If a user is reviewing a product, the generative model 152 may return a well-structured review even if the user has not explicitly specified in prompt 118 that it is a review. This context-aware approach can be beneficial even before the user types anything. For example, an embodiment may support a zero-state use case that provides a UI with general text input suggestions. For example, an embodiment may provide a suggestion to “write a constructive review” when a user is browsing a review webpage. In some embodiments, the generative model 152 can be further trained to also provide context-aware input suggestions. In that case, an example of a zero-state suggestion might be “write a 4-star review about this dining table” when the user is on a review page for a wooden dining table, or “write a review about <product> that doesn't work as intended” when the user is browsing a webpage for <product> that is not a review page (e.g., a customer complaints page).

[0045] The disclosed embodiments also reduce interaction between the user and the user device 102 in order to enable the insertion of generated text into a text field 136 of a web page 134. In particular, other large language models are not integrated into the browser application 108 and require the user to copy and paste the generated response. The disclosed embodiments are directly useful wherever the user is writing. The initial text input is taken directly from a text field on the web page the user is working on, and once the generative model 152 provides generated output text (response) that the user deems acceptable, the output text is inserted directly into the same text field 136. The disclosed embodiments can generate relevant ideas for the user to begin writing, adapt to the response to the user's voice, and give the user an initial draft for editing. Whether the user wants to share a witty comment about web content with a friend, file a complaint with a store, or simply craft a heartfelt reply to a wedding invitation, the Compose Assistant Manager 110 can be a reliable writing assistant that is directly integrated into the browser application 108.

[0046] The user device 102 may be any type of computing device, including one or more processors 101, one or more memory devices 103, a display 126, and an operating system 105 configured to run (or assist in running) one or more applications 106, including a browser application 108. In some examples, the browser application 108 is a web browser configured to access information about the Internet. The browser application 108 may launch one or more browser tabs in the context of one or more browser windows on the display 126 of the user device 102. A browser tab may display web documents (e.g., web pages, PDFs, images, videos, or any item that can be generally identified by a resource locator), and / or content associated with web applications, applications such as progressive web applications (PWAs), and / or extensions (e.g., web content). A web application may be an application program stored on a remote server (e.g., a web server) and delivered by the browser application 108 over the network 150. In some examples, a progressive web application is similar to a web application but is (at least partially) stored on the user device 102 and can be used offline. An extension adds features or functionality to a browser application 108. In some examples, an extension may be HTML, CSS, and / or JavaScript-based (in the case of a browser-based extension).

[0047] In some examples, user device 102 is a laptop computer. In some examples, user device 102 is a desktop computer. In some examples, user device 102 is a tablet computer. In some examples, user device 102 is a smartphone. In some examples, user device 102 is a wearable device. In some examples, display 126 is the display of user device 102. In some examples, display 126 may also include one or more external monitors connected to user device 102.

[0048] The processor(s) 101 may be formed on a substrate configured to execute one or more machine-executable instructions, or a portion of software, firmware, or a combination thereof. The processor(s) 101 may be semiconductor-based; that is, the processor may include semiconductor materials capable of performing digital logic. The memory device(s) 103 may include main memory that stores information in a format readable and / or executable by the processor(s) 101. The memory device(s) 103 may store a browser application(s) 108, a compose assistant manager(s) 110 (and, in some examples, a generative model(s) 152) that, when executed by the processor(s) 101, perform certain operations described herein. In some examples, the memory device(s) 103 may include a non-temporary computer-readable medium containing executable instructions that cause at least one processor(s) (e.g., processor(s) 101) to perform an operation. In some examples, the compose assistant manager(s) 110 may be configured to communicate with one or more generative models(s) 152. In some examples, the Compose Assistant Manager 110 allows the user to select one of several generative models 152 to be used to generate input to the text field 136, and the several generative models 152 include different LLMs. For example, the Compose Assistant Interface 128 may provide a first selectable option associated with a first generative model and a second selectable option associated with a second generative model. In response to the selection of the first selectable option, the Compose Assistant Manager may provide a prompt 118 and a context signal 120 to the first generative model. In response to the selection of the second selectable option, the Compose Assistant Manager may provide a prompt 118 and a context signal 120 to the second generative model.

[0049] The server computer(s) 160 can be a computing device in a wide variety of device forms, such as a standard server, a group of such servers, or a rack server system. In some examples, the server computer(s) 160 may be a single system sharing components such as a processor and memory. In some examples, the server computer(s) 160 may be multiple systems that do not share a processor and memory. The network 150 may include the Internet and / or other types of data networks, such as a local area network (LAN), wide area network (WAN), cellular network, satellite network, or other types of data networks. The network 150 may also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) configured to receive and / or transmit data within the network 150. The network 150 may further include any number of wired and / or wireless connections.

[0050] A server computer (or more) 160 may include one or more processors 161 formed on a circuit board, an operating system (not shown), and one or more memory devices 163. The memory devices (or more) 163 may represent any (or more) types of memory (e.g., RAM, flash, cache, disk, tape, etc.). In some examples (not shown), the memory devices may include external storage, such as memory that is physically far from the server computer (or more) 160 but accessible from it. The processors (or more) 161 may be formed on a circuit board and configured to execute one or more machine-executable instructions, or a portion of software, firmware, or a combination thereof. The processors (or more) 161 may be semiconductor-based; that is, the processors may include semiconductor materials capable of performing digital logic. The memory devices (or more) 163 may store information in a format that can be read and / or executed by the processors (or more) 161. In some examples, a memory device 163 may store a generative model 152 that, when executed by a processor 161, performs a specific operation as described herein. In some examples, the memory device 163 includes a non-temporary computer-readable medium containing executable instructions that cause at least one processor (e.g., processor 161) to perform an operation.

[0051] In addition to the above description, the user may be provided with controls that allow the user to make choices regarding whether and when the systems, programs, or functions described herein may enable the collection of user information (e.g., the user's browser usage history, user preferences, information about the user's current location, or other profile information), and whether the features described herein are active. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, user identification information may be processed so that personally identifiable information of the user cannot be determined, or the user's geographical location may be generalized so that the user's specific location cannot be determined if location information is available (e.g., city, zip code, or state level). Thus, the user may have control over what information is collected about them, how that information is used, and what information is provided to them.

[0052] Figure 2 shows a composition assistant interface 228 in one embodiment. The composition assistant interface 228 may be an example of the composition assistant interface 128 in Figures 1A to 1I and may include any of the details described with reference to those figures. As shown in Figure 2, the composition assistant interface 228 includes an input field 230 configured to receive prompts from the user. The composition assistant interface 228 displays a suggestion prompt 218a in the input field 230, which the user can select and / or edit.

[0053] In other words, before the user begins drafting a prompt in the input field 230 on the compose assistant interface 228, the compose assistant manager (e.g., the compose assistant manager 110 in Figures 1A-1I) may provide a suggestion prompt 218a. The suggestion prompt 218a may be generated by a generative model (e.g., the generative model 152 in Figures 1A-1I) based on one or more context signals (e.g., the context signals 120 in Figures 1A-1I).

[0054] Referring to Figure 2, the compose assistant interface 228 may include a generate control 231. In response to user selections in the generate control 231, the compose assistant manager may send prompt and context signals to the generate model. In some examples, the generate control 231 may remain inactive until the user provides text in the input field 230. Thus, after the user enters text in the input field 230, the generate control 231 may become active (and selectable). In some examples, the compose assistant interface 228 may include a close control 217. The close control 217 may close the compose assistant interface 228 and return the user to the text field.

[0055] Figure 3 shows a compose assistant interface 328 in another embodiment. In some examples, a user may invoke a compose assistant manager (e.g., compose assistant manager 110 in Figures 1A-1I) according to any of the techniques described herein, and display the compose assistant interface 328. In some examples, the compose assistant interface 328 may identify a set of categories 362 (e.g., type) of text fields in a web page 334. The user may select one of the categories 362 from the set of categories. In some examples, the compose assistant interface 328 may identify a tone control 364 that allows the user to select the tone of the model response generated by the generative model. The compose assistant interface 328 may include an input field 330 configured to receive prompts from the user. The compose assistant interface 328 may include a generation control 331. When selected by the user, the generation control 331 causes the compose assistant manager to send a prompt, a user selection made via the compose assistant interface 328, and context signals generated by the compose assistant manager.

[0056] Figures 4A to 4C show an example of a compose assistant interface 428 in one embodiment. The compose assistant interface 428 may be rendered on a web page 434 (e.g., a social media web page) to assist a user in writing a passage in a text field 436 on the web page 434. In some examples, the compose assistant interface 428 may be rendered when a compose assistant manager is invoked. The compose assistant manager may be invoked according to any of the techniques described herein.

[0057] As shown in Figure 4, the compose assistant interface 428 includes an input field 430 configured to receive a prompt 418 from the user. The user can provide a prompt 418 to the compose assistant interface 428. The prompt 418 may be a natural language description of the type of content generated by the generative model. For example, the user can type the prompt 418 or provide a voice command to insert the prompt 418 into the input field 430.

[0058] The compose assistant interface 428 may include a generation control 431. In response to a user selection in the generation control 431, the compose assistant manager (e.g., the compose assistant manager 110 in Figures 1A-1I) may send prompts 418 and context signals (e.g., context signals 120 in Figures 1A-1I) to the generation model (e.g., the generation model 152 in Figures 1A-1I). In some examples, the generation control 431 may remain inactive until the user provides text in the input field 430. In response to the prompts 418 and context signals, the generation model may generate a model response 424. As shown in Figure 4C, the compose assistant manager may receive the model response 424 from the generation model and display the model response 424 on the compose assistant interface 128.

[0059] As shown in Figure 4C, the compose assistant interface 428 may include an insert control 441. The insert control 441 inserts text into the text field 436. In some examples, the insert control 441 may close the compose assistant interface 428. The compose assistant interface 428 may also include controls for modifying (editing) the model response 424 using the generated model. For example, the compose assistant interface 428 may include a tone control 423. The tone control 423 may allow the user to make the response sound more formal, more casual, more humorous, or to include emojis. The compose assistant interface 428 may include a length control 425. The length control 425 may allow the user to shorten or lengthen the model response 424. In response to the user selecting either the tone control 423 or the length control 425, the compose assistant manager may provide a new model response 424. In other words, by selecting either the tone control 423 or the length control 425, the generated model can regenerate the model response 424 based on the value of the selected control.

[0060] In some examples, the compose assistant interface 428 may include a regenerate control 435. The regenerate control 435 may generate other text suggestions (e.g., a new model response 424). The compose assistant interface 428 may also include a close control 415. The close control 415 may close the compose assistant interface 428 and return the user to the text field 436.

[0061] Figures 5A to 5F show an example of the Compose Assistant interface 528 of the Compose Assistant Manager in one embodiment. The Compose Assistant interface 528 may be rendered on a display for a text field 136 on a web page. The Compose Assistant interface 528 may be triggered according to any of the techniques described herein. In some examples, the Compose Assistant interface 528 may be a UI dialog.

[0062] As shown in Figure 5B, the compose assistant interface 528 may display a loading status while the initial writing suggestions 524 are being generated. In some examples, the generation of the initial writing suggestions 524 may begin after the user has entered a threshold number of words (the amount of text data that reaches a threshold level) into the text field 538. In some examples, as shown in Figure 5C, the compose assistant interface 528a may appear after the user has selected a threshold number of words in the text field 536, and the compose assistant interface 528a may present the user with a set of actions 550 (e.g., proofreading actions 540 and detailed actions 542). In some examples, the compose assistant interface 528a may include an expander control 544 that presents additional actions when selected.

[0063] In some examples, the initial writing suggestion 524 may be displayed in the compose assistant interface 528, as shown in Figure 5D. In some examples, the compose assistant manager may send text from the text field 536 and context signals (e.g., context signal 120 in Figures 1A-1I) to the generating model. The generating model may generate a model response based on the initial writing suggestion 524. In some examples, as shown in Figure 5E, the user may hover the cursor over the compose assistant interface 528, and a preview of the initial writing suggestion 524 may be displayed in the text field 536. In some examples, the compose assistant manager may detect the cursor position over the suggestion (e.g., the initial writing suggestion 524), and in response to the cursor position within the boundaries of the suggestion, the compose assistant manager may provide a preview of the suggestion in the text field 536. In some examples, as shown in Figure 5F, in response to the user moving the cursor over the expander control 544, the Compose Assistant Manager may render an action menu 562 displaying a set of actions 550.

[0064] Figure 6 shows a composition assistant interface 628 in one embodiment. The composition assistant interface 628 includes a prompt field 618 that displays a prompt and an edit control 660 that, when selected, allows the user to edit the prompt. The composition assistant interface 628 displays a model response 624. The composition assistant interface 628 may display a set of controls, such as a fine-tuning control 670 (which, when expanded, may display controls for shortening, lengthening, tone adjustment, etc.), an undo control 671, and a redo control 635. The composition assistant manager 628 may include an insert control 641 that, when selected, inserts the prompt into a text field on a web page.

[0065] Figure 7 shows a composition assistant interface 728 in one embodiment. The composition assistant interface 728 includes a prompt field 718 that displays a prompt and an edit control 760 that, when selected, allows the user to edit the prompt. The composition assistant interface 728 displays a model response 724. The composition assistant interface 728 may display a set of controls such as a length control 725, a tone control 723, and an undo control 735. The composition assistant manager 728 may include an insert control 741 that, when selected, inserts the prompt into a text field on a web page.

[0066] Figure 8 shows a system 800 comprising a user device 802 and a server computer 860 for implementing the concepts described herein. Generally, the user device 802 can represent any computing device running a browser application 808. As shown in Figure 8, the user device 802 is configured to communicate with the server computer 860 and / or a resource provider (e.g., a web server) via a network 850. The user device 802 includes at least the browser application 808 and other applications (not shown). In some embodiments, the browser application 808 is configured to manage resource content, such as web page content, provided by a resource provider (e.g., a web server). In some embodiments, the browser application 808 is configured to operate as one of several applications running via an operating system (O / S) 802.

[0067] Although not shown in Figure 8, the user device 802 includes several hardware components, including a communication module, one or more cameras, memory, a processing unit 801 such as a central processing unit (CPU) and / or graphics processing unit (GPU), one or more input devices 867 (e.g., a touchscreen, mouse, stylus, microphone, keyboard, etc.), and one or more output devices 868 (e.g., a screen, speaker, vibrator, light emitter, etc.). The hardware components can be used to facilitate the operation of the user device 802, such as a browser application 808. The user device 802 may also include an operating system 805. The browser application 808 includes, for example, a compose component 810 configured to generate a compose assistant user interface, as shown in various figures.

[0068] User device 802 may contain local user profile data 855. Local user profile data 855 may be stored in memory associated with the browser application 808, or in memory accessible to the browser application 808. Local user profile data 855 may be a data source(s) of user-specific information obtained from the user's usage of the browser application 808, collected with user privileges. Local user profile data 855 is on-device storage. In some embodiments, local user profile data 855 may be associated with an account profile, e.g., a user account on server computer 860. In such embodiments, some information may be stored in central user profile data 842. The user controls what information is shared between local user profile data 855 and central user profile data 842 and when. By sharing data from local user profile data 855 (e.g., signals that help the Compose Assistant know when to trigger callout affordances, signals that help define the user's tone) with central user profile data 842, the user can have a consistent experience of the Compose Assistant across user devices.

[0069] The browser application 808 includes a compose-renderer helper 827. The compose-renderer helper 827 runs within the renderer process and performs actions related to the web page 834 and the text field 836. The compose-renderer helper 827 may include a web page interaction component responsible for interactions with the web page 834 necessary for the user experience flow, such as monitoring user interactions with the text field (e.g., text field 836), triggering the display of callout affordances, and extracting and inserting text from the text field. The compose-renderer helper 827 may include a context extraction component instrumented to capture a set of signals that help generate a response (text) to the text field. When the user requests an LLM response, the content extraction component extracts all expected context from the page, which is packed along with the prompt. Context may include the URL, title, and / or page content, as well as / or other signals described herein. For page content, the system may leverage different approaches to determine the most relevant portion of the content. For example, the compose renderer helper 827 may use the DOM (Document Object Model) or accessibility tree to identify which parts of the content are visible or not, which parts of the content enclose an input field (e.g., a text field 836), and other key content parts of the page, such as heading fields. In such embodiments, the content extraction component may extract DOM parts from the DOM representation of the context. In the context of a page related to a conversation, the context includes previous rounds of the conversation.

[0070] The context can include the main entity of a web page. For example, if a web page is a review web page for vacuum cleaners, the main entity could be that vacuum cleaner (or a certain vacuum cleaner). In some embodiments, a generative language model can be trained to recognize the main entity within the content of a web page provided as context. Furthermore, the compose renderer helper 827 can identify text input fields and leverage their metadata to classify their likely purpose in the context of the page paired with user-provided prompts. The browser application 808 extracts raw signals and processes signals useful for creating the correct context to be used by the compose generative language model 852. This includes identifying the type of page, form, and input field the user is typing in. Context can also be obtained from a website (e.g., the domain to which the web page belongs). For example, in an embodiment where a server computer 860 is associated with a search engine and the website is indexed, the content of the domain's search index can be used as context signals.

[0071] The context may also include user history signals with the user's consent. User history signals may include previously generated responses so that, for example, the Compose Generative Language Model 852 can mimic tone. For example, the context may include prompt packing to provide the Compose Generative Language Model 852 with a small number of training shots. Prompt packing is used to bias the Compose Generative Language Model 852 to generate responses that are more similar to how this particular user has formatted responses in the past. In some embodiments, prompt packing may be stored as state in, for example, local user profile data 855. User history signals may include shopping history metadata. For example, if a browser application 808 can access the shopping history and a webpage is a review of a product the user has purchased (for example, if the user clicked an email link requesting them to leave a review, in which case the webpage may be part of a custom tab associated with the email application), the delivery time may be known or computable, and this information can be added as context and made available to the Compose Generative Language Model 852 when generating the review (response). Similarly, flight information may be used when responding to rental car or hotel information. Browser application 808 may include a configuration user interface. The configuration user interface may include a menu that allows the user to enable or disable the compose assistant.

[0072] Figure 8 shows several embodiments of the server computer 860. For example, the server computer 860 may include a compose service 844, a security / policy filter 846, and a compose generation language model 852. The server computer 860 may also include one or more processors (not shown) and one or more memory devices (not shown). The compose service 844 may be server-side business logic responsible for querying all dependent services and data sources to satisfy user requests, including collecting some further user data from central user profile data 842 and requesting inference from the compose generation language model 852. The security / policy filter 846 may ensure that information received from and sent back to the user device 802 complies with all security, policy, and legal requirements, such as avoiding sensitive categories and filtering unsafe content. In some embodiments, the policy filter 846 may be a known classifier that identifies negative / bad prompts and / or types of page context (adult content / violence / offensive). The policy filter 846 can be applied to prompts and their context, and to output from the Compose-Generating Language Model 852. The policy filter 846 can prevent the Compose-Generating Language Model 852 from providing output to the user and instead return an error message indicating that it was unable to process the prompt.

[0073] The Compose Generative Language Model 852 is a generative language model custom-trained for the Compose Assistant to adapt to the use cases targeted by the function. User cases are based on the purpose or type of text field. For example, purpose / type could be reviews (product, place, travel, etc.), comments (e.g., on videos or articles), social media posts, survey responses, forums, conversation replies (e.g., conversations in chatbots or messaging apps), customer complaints, blogs, profile descriptions, etc. Through training, the Compose Generative Language Model 852 will be able to appropriately consider additional context extracted from web pages 834 as well as local user profile data 855 and / or central user profile data 842. Also through training, the Compose Generative Language Model 852 will be able to generate purpose-specific responses to produce responses that are similar in length to, for example, average product reviews, average social media posts, or average forum posts. Thus, the Compose Generative Language Model 852 can leverage input signals (web pages, text input fields, and / or user context) to tailor output based on provided browser signals. Therefore, the Compose Generation Language Model 852 can be fine-tuned to produce the correct writing structure based on the context of the prompt provided by the user. For example, if a user is on a product review page and provides a limited prompt, the system (e.g., a browser application 808) can add sufficient context so that the text generated from the Compose Generation Language Model 852 is a structured review that includes details from the user prompt. Thus, the Compose Assistant leverages the page context to provide solutions that prompt the user based on goals, categories, topics / themes, etc.

[0074] In some embodiments, the Compose Assistant Manager may further personalize the voice for the user by leveraging previous prompts and examples submitted from user interactions (e.g., stored in local user profile data 855). This is called prompt packing. This ensures that the tone and voice are more consistent throughout the individual user experience. User permissions allow these additional user signals to be synchronized with the user profile across the devices the user is signed into (e.g., user device 802), ensuring consistency in tone and voice across user devices.

[0075] In some embodiments, the Compose Generative Language Model 852 may be configured to generate responses with variable placeholders. For example, if a text box is part of a conversation (e.g., a message in an instant messaging conversation or a chat with a chatbot), the user may be responding to a request for specific information (e.g., a fact). The request for the fact may be part of the context provided to the Compose Generative Language Model 852. The Compose Generative Language Model 852 may be configured to generate appropriate variable placeholders for the response generated for the specific information. For example, the response might be "Thank you! Please call [phone number] after 5 p.m." where [phone number] is a variable placeholder that the user can edit.

[0076] In some embodiments, user privileges, along with additional user context available within the browser (represented by local user profile data 855 via profiles and other user data stores), may enable the compose assistant to automatically populate variable placeholders in the generated response. For example, when replying to a post about contact information, the compose generation language model 852 may output a variable placeholder called [user_x_address], which the browser application 808 can then leverage by considering the contact information available in local user profile data 855 and pre-populating the value in the generated response.

[0077] The Compose Generative Language Model 852 can be trained on different types of input fields, i.e., examples of input fields with different types of purposes. Because the Compose Generative Language Model 852 is trained for oriented tasks, it can be smaller than a general-purpose large-scale language model and deliver output faster.

[0078] In some embodiments, the compose generation language model 852 runs on a device, for example, a user device 802, so the server computer 860 is not required. In such embodiments, the functions performed by the compose service 844 and / or policy filter 846 may be performed by one of the compose assistant components of the browser application 808, for example, the compose component 810 and / or the compose renderer helper 827.

[0079] Figure 9 is a flowchart illustrating an exemplary process 900 for providing a compose assistant according to one embodiment. Process 900 may be performed by a browser compose assistant manager, such as browser application 108 in Figures 1A-1I and / or browser application 808 in Figure 9. In step 902, the system receives focus on a text box on a web page. In step 904, the system decides whether to surface a callout affordance configured to initiate the compose assistant for the text box. In step 906, in response to deciding to surface the callout affordance, the system provides the user with a suggestion prompt in the compose assistant interface.

[0080] Figure 10 is a flowchart of an exemplary process 1000 for providing a compose assistant manager according to one embodiment. Process 1000 can be performed by a browser compose assistant manager such as browser application 108 in Figures 1A-1I and / or browser application 808 in Figure 8. In step 1002, the system may receive prompts from the user related to input into text boxes on a web page. In step 1004, the system may generate context signals for the web page. Context signals may include content signals. Context signals may include user signals. In step 1006, the system may provide prompts and context signals to a generative language model trained to provide output for a text box type (purpose). In step 1008, the system may receive a response generated by the generative language model. In step 1010, for example, in response to the selection of an accept control, the system may provide a response as input to a text box. Thus, using process 1000, the user can minimize interaction with the computing device, and the model can generate high-quality output suitable for the purpose of the text box.

[0081] The following are exemplary use cases of the published embodiments. These use cases are not limiting. The embodiments can help users with specific problems, such as writer's block. For example, a user who likes to share content to stay in touch with friends and family may consume interesting content but lack the wit to share it on social platforms. Compose Assistant can help this user draft content to share. Another example is a user who has recently had a bad experience with an airline and wants to file a complaint. Compose Assistant can help express concerns effectively, professionally, and clearly. Another example is a user who may be a blogger and social media influencer but is suffering from writer's block and needs inspiration for their next blog post or social media post. They want to ensure that the content they produce is engaging and relevant to their audience. Another example is a user who has recently moved to an English-speaking country and needs to write emails, job applications, and other documents in English. Compose Assistant can help with these drafts. Another example: the user might be a brand manager who needs to provide daily content inspiration for the company they represent. They must maintain consistency in the brand's voice across various platforms. With consent, the Compose Assistant can ensure consistency across platforms. Another example: the user might be a university student who needs to write a research paper and is struggling to organize their thoughts.

[0082] Figure 11 is a flowchart 1100 illustrating exemplary operation of a system for integrating a language model into a browser application according to one embodiment. Flowchart 1100 may represent the operation of a method performed by a computer. Flowchart 1100 may be applicable to any of the embodiments described herein. Although flowchart 1100 in Figure 11 shows the operation in order, it will be understood that this is merely an example and may include additional or alternative operations. Furthermore, the operation in Figure 11 and related operations may be performed in a different order than illustrated, or in a parallel or overlapping manner.

[0083] Operation 1102 includes receiving prompts from the user related to input into a text field on a web page. In some examples, the prompts are called text data and the web page is called digital content. Operation 1104 includes generating context signals for the web page. In some examples, the context signals are called context data. Operation 1106 includes providing the prompts and context signals to the generating language model. Operation 1108 includes receiving the response generated by the generating language model. Operation 1110 includes providing the response as input to a text field. In some examples, operation 1110 includes providing the response as a suggestion for input to a text field. In some examples, depending on the acceptance of the response, the application may insert the response directly into the text field.

[0084] Clause 1. A method comprising: receiving text data from a user relating to input into a text field of digital content displayed on a user device; generating contextual data relating to the digital content; providing the text data and the contextual data to a generating language model; receiving a response generated by the generating language model; and providing the response as a suggestion for the input to the text field.

[0085] Clause 2. The method according to Clause 1, further comprising detecting interaction with the text field and determining by a model whether to render a callout affordance, wherein the callout affordance, if selected, is configured to render a compose assistant interface for the text field, and the compose assistant interface has an input field configured to receive the text data from the user.

[0086] Clause 3. The method according to Clause 2, further comprising determining whether to render the callout affordance based on a signal, wherein the signal includes one or more signals relating to the text field, one or more signals relating to the digital content, or one or more signals relating to the user and other users of the Compose Assistant.

[0087] Clause 4. The method according to Clause 1, further comprising rendering a callout affordance in response to the amount of text data entered by the user into the text field reaching a threshold level, wherein the callout affordance, when selected, renders a compose assistant interface for the text field, the compose assistant interface having an input field comprising the text data.

[0088] Clause 5. The method according to Clause 1, further comprising receiving a selection of the digital content to a user interface object for the text field, and rendering a compose assistant interface for the text field, wherein the compose assistant interface has an input field configured to receive the text data from the user, and the method further comprises transmitting the text data and the context data to the generate language model in response to a selection of the generate control of the compose assistant interface.

[0089] Clause 6. The method according to Clause 1, further comprising receiving a selection of the text data entered by the user into the text field, and rendering a compose assistant interface having controls that cause the text data and the context data to be sent to the generating language model when selected.

[0090] Clause 7. The method according to Clause 1, further comprising transmitting the text data and the context data in response to the amount of text data entered by the user into the text field reaching a threshold level, and providing the response as a suggestion in the compose assistant interface.

[0091] Clause 8. The method according to Clause 7, further comprising detecting the cursor position on the suggestion and providing a preview of the response in the text field.

[0092] Clause 9. The method according to Clause 1, further comprising inserting the response into the text field.

[0093] Clause 10. The method according to Clause 1, wherein the digital content is a web page, and the method further comprises retrieving a first page content of the web page, retrieving a second page content of the web page embedded in the web page, and generating context data to include the first page content and the second page content.

[0094] Clause 11. A method according to Clause 1, wherein the digital content is a web page, and the method further comprises extracting a document object model (DOM) representation of the web page, extracting DOM portions from the DOM representation, and generating context data to include the DOM portions.

[0095] Clause 12. The method according to Clause 1, wherein the digital content is a web page, and the method further comprises extracting the accessible content structure of the web page and generating the context data to include the accessible content structure.

[0096] Clause 13. An apparatus comprising at least one processor and a non-temporary computer-readable medium storing executable instructions causing the at least one processor to perform an operation, wherein the operation includes receiving text data from a user relating to input into a text field of digital content displayed on a user device; generating contextual data relating to the digital content; providing the text data and the contextual data to a generating language model; receiving a response generated by the generating language model; and providing the response as a suggestion for the input to the text field.

[0097] Clause 14. The apparatus according to Clause 13, wherein the operation further comprises the model determining whether to render a callout affordance based on a signal, the signal including one or more signals relating to the text field, one or more signals relating to the digital content, or one or more signals relating to the user and other users of the compose assistant, the callout affordance, when selected, is configured to render the compose assistant interface for the text field, the compose assistant interface having an input field configured to receive the text data from the user.

[0098] Clause 15. The apparatus according to Clause 13, wherein the operation further comprises rendering a callout affordance in response to the amount of text data entered by the user into the text field reaching a threshold level, the callout affordance being configured to render a compose assistant interface for the text field when selected, the compose assistant interface having an input field comprising the text data.

[0099] Clause 16. The apparatus according to Clause 13, wherein the operation further includes receiving a selection of the digital content to a user interface object for the text field and rendering a compose assistant interface for the text field, the compose assistant interface having an input field configured to receive the text data from the user.

[0100] Clause 17. The apparatus according to Clause 13, wherein the digital content is a web page, and the operation further comprises retrieving a first page content of the web page, retrieving a second page content of the web page embedded in the web page, and generating the context data to include the first page content and the second page content.

[0101] Clause 18. A non-temporary computer-readable medium storing executable instructions causing at least one processor to perform an operation, wherein the operation includes receiving text data from a user relating to input into a text field of digital content displayed on a user device; generating context data for the digital content; providing the text data and the context data to a generating language model; receiving a response generated by the generating language model; and providing the response as a suggestion for the input to the text field.

[0102] Clause 19. The operation further comprises, by the model, determining whether to render a callout affordance, which, if selected, is configured to render a compose assistant interface for the text field, the compose assistant interface having an input field configured to receive the text data from the user, the non-temporary computer-readable medium as described in Clause 18.

[0103] Clause 20. Non-temporary computer-readable media as described in Clause 18, wherein the digital content is a web page, and the operation further comprises retrieving a first page content of the web page, retrieving a second page content of the web page embedded in the web page, and generating the context data to include the first page content and the second page content.

[0104] Various embodiments of the systems and technologies described herein can be realized in digital electronic circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include embodiments in one or more computer programs executable and / or interpretable on a programmable system, comprising at least one programmable processor, the at least one programmable processor may be dedicated or general-purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to them.

[0105] These computer programs (also known as programs, software, software applications, or code) contain machine instructions for programmable processors and can be executed in high-level procedural and / or object-oriented programming languages ​​and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic circuits (PLDs)) used to provide machine instructions and / or data to a programmable processor, including machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signals” refers to any signals used to provide machine instructions and / or data to a programmable processor.

[0106] To provide user interaction, the systems and technologies described herein can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can also be used to provide user interaction; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic, spoken language, or tactile input.

[0107] The systems and technologies described herein may be implemented in a computing system that includes backend components (e.g., as a data server), middleware components (e.g., an application server), or frontend components (e.g., a client computer having a graphical user interface or web browser that allows a user to interact with embodiments of the systems and technologies described herein), or in combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by digital data communications (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), and the Internet.

[0108] Several embodiments are described. Needless to say, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosed embodiments.

[0109] Furthermore, the logic flow shown in the diagram does not require a specific or sequential order to achieve the desired result. Additionally, other steps may be added to the described flow, or steps may be removed from the described flow; and other components may be added to the described system, or other components may be removed from the described system.

[0110] In some embodiments, the techniques described herein include a method for receiving focus on a text box on a web page, determining whether to surface a callout affordance configured to initiate a compose assistant for the text box, and, in response to a decision to surface the callout affordance, providing the user with a suggestion prompt in the compose assistant interface. The suggestion prompt may be based on the context of the web page.

[0111] In some embodiments, the techniques described herein include a method for receiving prompts from a user relating to input into a text box on a web page; generating context signals for the web page; providing the prompts and context signals to a generative language model trained to provide output of a type text box; receiving a response generated by the generative language model; and providing the response as input to a text box.

[0112] In some embodiments, the technology described herein relates to a non-temporary computer-readable medium that stores instructions that, when executed by a processor, perform one of the operations disclosed herein.

[0113] In some embodiments, the technology described herein relates to a computing device comprising at least one processor and a memory for storing instructions causing the computing device to perform any of the operations or methods disclosed herein.

Claims

1. Receiving text data from the user related to input into text fields of digital content displayed on the user device, To generate contextual data relating to the aforementioned digital content, The text data and the context data are provided to the generating language model, Receiving the response generated by the aforementioned generative language model, A method comprising providing the response as a suggestion for the input to the text field.

2. To detect interaction with the aforementioned text field, The method according to claim 1, further comprising determining whether to render a callout affordance by model, wherein the callout affordance, if selected, is configured to render a compose assistant interface for the text field, and the compose assistant interface has an input field configured to receive the text data from the user.

3. The method according to claim 2, further comprising determining whether to render the callout affordance based on a signal, wherein the signal includes one or more signals relating to the text field, one or more signals relating to the digital content, or one or more signals relating to the user and other users of the compose assistant.

4. The method according to any one of claims 1 to 3, further comprising rendering a callout affordance in response to the amount of text data entered by the user into the text field reaching a threshold level, wherein the callout affordance, when selected, is configured to render a compose assistant interface for the text field, the compose assistant interface having an input field comprising the text data.

5. Receiving a selection of a user interface object for the text field of the digital content, The method further includes rendering a compose assistant interface for the text field, wherein the compose assistant interface has an input field configured to receive the text data from the user, and the method further includes The method according to any one of claims 1 to 4, comprising transmitting the text data and the context data to the generating language model in response to the selection of a generation control in the Compose Assistant interface.

6. Receiving the selection of the text data entered into the text field by the user, The method according to any one of claims 1 to 5, further comprising rendering a compose assistant interface having controls, wherein, when selected, the controls cause the text data and the context data to be sent to the generating language model.

7. In response to the amount of text data entered by the user into the text field reaching a threshold level, the text data and the context data are transmitted. The method according to any one of claims 1 to 6, further comprising providing the response as a suggestion in the compose assistant interface.

8. Detecting the cursor position on the aforementioned suggestion, The method according to claim 7, further comprising providing a preview of the response in the text field.

9. The method according to any one of claims 1 to 8, further comprising inserting the response into the text field.

10. The aforementioned digital content is a web page, and the aforementioned method is Extracting the first page content of the aforementioned web page, Extracting the second page content of the web page embedded in the aforementioned web page, The method according to any one of claims 1 to 9, further comprising generating the context data to include the first page content and the second page content.

11. The aforementioned digital content is a web page, and the aforementioned method is Extracting the document object model (DOM) representation of the aforementioned web page, Extracting the DOM portion from the aforementioned DOM expression, The method according to any one of claims 1 to 10, further comprising generating the context data to include the DOM portion.

12. The aforementioned digital content is a web page, and the aforementioned method is Extracting the accessible content structure of the aforementioned web page, The method according to any one of claims 1 to 11, further comprising generating the context data to include the accessible content structure.

13. At least one processor, A device comprising a non-temporary computer-readable medium storing executable instructions that cause at least one processor to perform an operation, wherein the operation is Receiving text data from the user related to input into text fields of digital content displayed on the user device, To generate contextual data relating to the aforementioned digital content, The text data and the context data are provided to the generating language model, Receiving the response generated by the aforementioned generative language model, An apparatus that includes providing the response as a suggestion for the input to the text field.

14. The aforementioned operation is, The apparatus according to claim 13, further comprising determining whether to render a callout affordance based on a signal, wherein the signal includes one or more signals relating to the text field, one or more signals relating to the digital content, or one or more signals relating to the user and other users of the compose assistant, and the callout affordance, when selected, is configured to render a compose assistant interface for the text field, and the compose assistant interface has an input field configured to receive the text data from the user.

15. The aforementioned operation is, The apparatus according to claim 13 or 14, further comprising rendering a callout affordance in response to the amount of text data entered by the user into the text field reaching a threshold level, wherein the callout affordance, when selected, is configured to render a compose assistant interface for the text field, the compose assistant interface having an input field comprising the text data.

16. The aforementioned operation is, Receiving a selection of a user interface object for the text field of the digital content, The apparatus according to any one of claims 13 to 15, further comprising rendering a compose assistant interface for the text field, wherein the compose assistant interface has an input field configured to receive the text data from the user.

17. The aforementioned digital content is a web page, and the aforementioned operation is, Extracting the first page content of the aforementioned web page, Extracting the second page content of the web page embedded in the aforementioned web page, The apparatus according to any one of claims 13 to 16, further comprising generating the context data to include the first page content and the second page content.

18. A non-temporary computer-readable medium that stores executable instructions causing at least one processor to perform an operation, wherein the operation is: Receiving text data from the user related to input into text fields of digital content displayed on the user device, To generate context data for the aforementioned digital content, The text data and the context data are provided to the generating language model, Receiving the response generated by the aforementioned generative language model, A non-temporary computer-readable medium, which includes providing the response as a suggestion for the input to the text field.

19. The aforementioned operation is, The non-temporary computer-readable medium according to claim 18, further comprising determining whether to render a callout affordance by model, the callout affordance being configured, if selected, to render a compose assistant interface for the text field, the compose assistant interface having an input field configured to receive the text data from the user.

20. The aforementioned digital content is a web page, and the aforementioned operation is, Extracting the first page content of the aforementioned web page, Extracting the second page content of the web page embedded in the aforementioned web page, The non-temporary computer-readable medium according to claim 18 or 19, further comprising generating the context data to include the first page content and the second page content.