Writing assistant manager for applications
By integrating generative models into the writing assistant manager, generating context data, and directly inserting responses into webpage text fields, the high computational resource consumption and security issues encountered by users when drafting webpage content are resolved, achieving efficient and secure content generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-26
- Publication Date
- 2026-03-27
AI Technical Summary
When users use generative language models to draft web page content, they need specific terminology and multiple iterations, and there are issues with security and high computational resource consumption.
By integrating a generative model with a writing assistant manager, the generated response is directly inserted into the text fields of the webpage by generating context data and combining it with user input, reducing computational resource consumption and maintaining security.
It provides a user-friendly text input experience, reduces the computational resource consumption of generated content, and improves the security and relevance of generated content.
Smart Images

Figure CN121753030A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 578,816, filed August 25, 2023, the disclosure of which is incorporated herein by reference in its entirety. Background Technology
[0003] Some web pages include text boxes that provide text input from the user. Examples may include web pages that allow users to leave reviews about products, services, places, etc.; web pages that allow users to leave comments or reply to comments; web pages that allow users to post messages (e.g., web pages on social media websites); and / or web pages that include surveys. Users can use generative language models to help draft input for web pages. However, users may need to be relatively specific in their terminology when drafting their prompts, and / or may need to perform multiple iterations using the language model to create the desired comment. Furthermore, obtaining contextual data from the web content used by generative language models may present one or more security-related technical challenges. Summary of the Invention
[0004] This disclosure relates to a compose assistant manager for an application (e.g., a browser application) that integrates a generative model (e.g., a language model) for drafting content as input to text fields of digital content (e.g., a webpage). The compose assistant manager provides one or more technical benefits in maintaining the security of the application content (e.g., a webpage) and / or reducing the amount of computational resources (e.g., memory, CPU) consumed in generating the generative content and inserting it (e.g., directly inserting it) into the text fields of the digital content. The compose assistant manager can provide users with reduced overhead when creating prompts and can customize the generated output for the context of the digital content. The compose assistant manager can generate one or more contextual signals (also referred to as contextual data) about the digital content (e.g., a webpage), and can transmit text data received from the user (e.g., also referred to as prompts or user-provided prompts) and content signals to a generative language model that returns a model response that can be directly inserted into the text fields. In other words, the composition assistant manager helps users enter text into text fields provided by the computer system and does so using technical information—specifically, contextual information about the webpage, which can come from the content of the webpage.
[0005] In some aspects, the techniques described herein relate to a method comprising: receiving from a user text data relating to input in a text field of digital content displayed on a user device; generating context data about the digital content; providing the text data and the context data to a generative language model; receiving a response generated by the generative language model; and providing the response as a suggestion for input in the text field.
[0006] In some aspects, the technology described herein relates to an apparatus comprising: at least one processor; and a non-transitory computer-readable medium storing executable instructions that cause the at least one processor to perform operations including: receiving text data from a user in relation to input of a text field of digital content displayed on a user device; generating context data about the digital content; providing the text data and the context data to a generative language model; receiving a response generated by the generative language model; and providing the response as a suggestion for input to the text field.
[0007] In some aspects, the technology described herein relates to a non-transitory computer-readable medium storing executable instructions that cause at least one processor to perform operations including: receiving text data from a user in relation to input of a text field of digital content displayed on a user device; generating context data of the digital content; providing the text data and the context data to a generative language model; receiving a response generated by the generative language model; and providing the response as a suggestion for input to the text field.
[0008] Details of one or more implementations are set forth in the accompanying drawings and the following description. Other features will be apparent from the description and drawings. Attached Figure Description
[0009] Figure 1A An example callout affordance is shown for invoking the Writer Assistant Manager based on one aspect.
[0010] Figure 1B An example of a call to the Authoring Assistant Manager is shown, based on one aspect.
[0011] Figure 1C The writing assistant interface for receiving prompts is shown, based on one aspect.
[0012] Figure 1D The writing assistant interface for displaying model responses is shown, based on one aspect.
[0013] Figure 1EAn example is shown of a text field input along with the model response, based on one aspect.
[0014] Figure 1F A system with a composition assistant manager for browser applications is shown, which integrates a language model for drafting content as input to text fields of web pages.
[0015] Figure 1G An example of a field signal used to generate a model response based on one aspect is shown.
[0016] Figure 1H An example of a trigger engine based on one aspect is shown.
[0017] Figure 1I An example of a webpage with embedded resources based on one aspect is shown.
[0018] Figure 2 An example of a writing assistant interface based on one aspect is shown.
[0019] Figure 3 An example of a writing assistant interface based on another aspect is shown.
[0020] Figures 4A to 4C The text describes a writing assistant interface rendered on a social media webpage, based on one aspect.
[0021] Figures 5A to 5F The various aspects of the writing assistant interface based on one aspect are shown.
[0022] Figure 6 An example of a writing assistant interface based on another aspect is shown.
[0023] Figure 7 An example of a writing assistant interface based on another aspect is shown.
[0024] Figure 8 This is a diagram illustrating the components of a computing system and server used to implement the concepts described herein, according to one aspect.
[0025] Figure 9 This is a flowchart illustrating an example process for providing a writing assistant manager based on one aspect.
[0026] Figure 10 This is a flowchart illustrating an example process for providing a writing assistant manager, based on another aspect.
[0027] Figure 11 This is a flowchart illustrating an example process for providing a writing assistant manager, based on another aspect. Detailed Implementation
[0028] This disclosure relates to an authoring assistant manager for an application (e.g., a browser application) that integrates a generative model (e.g., a language model) to generate content for text fields and inserts that content (e.g., directly into the text fields). The authoring assistant manager provides one or more technical benefits in maintaining the security of web pages and / or reducing the amount of computational resources (e.g., memory, CPU) consumed in generating generative content and inserting it into one or more text fields of web content. The authoring assistant manager can assist users in leaving comments, reviewing articles, providing survey responses, drafting social media posts, filling out customer complaints, and / or responding to chatbots, etc.
[0029] In some examples, the user can explicitly invoke the writing assistant manager. For example, a user can right-click on a text field on a webpage and select a menu option (e.g., the "Help me write" option), which causes the display of a writing assistant interface for receiving prompts for generative models. In some examples, the text field can be any type of input field configured to receive text from the user (e.g., via keyboard, voice, touchscreen, etc.), with the text received by the user filling the text field. In some examples, the text field is a free-form text field. In some examples, the text field is a structured text field. In some examples, the text field is a multi-line text field. In some examples, the text field is a single-line text field. In some examples, the text field may be filled with data received via a microphone (e.g., via a voice assistant).
[0030] Users can provide prompts in the writing assistant interface (e.g., "write a five-star review about this product"). For example, the writing assistant interface includes input fields that allow users to draft prompts—such as a natural language description of the type of content to be generated by the generative model. In response to a prompt submission, the writing assistant manager can transmit that prompt along with one or more contextual signals (also called contextual data) about the underlying webpage. This contextual data can include information about the topic of the webpage. In response to the prompt and the contextual data, the generative model can generate and return a context-related response that can be directly inserted into a text box on the webpage.
[0031] The Writer Assistant Manager provides a technical solution for generating context data (e.g., one or more context signals) about the underlying webpage, which helps the generative model create context-relevant responses. In some examples, context data includes resource locators, page titles, page content, Document Object Model (DOM) representations, and / or accessibility content structures (e.g., accessibility trees). In some examples, the Writer Assistant Manager can retrieve the first page content of a webpage with text fields and the second page content of one or more embedded webpages, including both in the context signal. In some examples, the webpage includes one or more inline frames (e.g., iframes). An iframe is an HTML element that embeds another Hypertext Markup Language (HTML) document within the current page. Retrieving page content from embedded webpages can present one or more technical challenges to maintaining security.
[0032] However, the Writer Assistant Manager overcomes technical challenges by performing context extraction of context signals, which involves requesting internal text from a specified host and internal text from local same-origin iframes (e.g., all local same-origin iframes). A same-origin iframe can be an iframe that shares the same origin as the main webpage (e.g., an embedded frame within a webpage). The origin of a webpage is determined by its protocol, hostname, and port number. In some examples, the embedded iframe resides on the same server or domain as the main webpage. In some examples, the embedded iframe may have the same protocol, hostname, and / or port number as the main webpage. Internal text can refer to the visible text content within an HTML element and text from one or more child elements of the HTML element. The returned internal text includes the combined internal text of the iframes (e.g., all iframes). The Writer Assistant Manager retrieves the internal text of the webpage, and the webpage's internal text is combined with the internal text of each detected iframe (e.g., the embedded webpage). The Writer Assistant Manager provides context signals and user-provided prompts to the generative model. The generative model generates a model response and returns it to the writing assistant manager, which can then insert the model response directly into the input text box (e.g., with or without user prompts).
[0033] The writing assistant manager can display the model response in a text field. The writing assistant interface can include one or more UI elements that allow the user to adjust the model response (e.g., make it more formal, less formal, expanded, shortened, etc.), causing the generative model to regenerate the model response. In some examples, the user can manually edit the model response. The writing assistant manager can include an insertion control that, when selected, causes the model response to be inserted into a text field on the webpage. For example, in response to the selection of the insertion control, the writing assistant manager transfers the text from the input field of the writing assistant interface to a text field on the webpage.
[0034] In some examples, the writing assistant interface can provide the user with one or more suggested prompts that the user can select and / or edit. In other words, the writing assistant interface can provide selectable suggested prompts before the user begins drafting a prompt in an input field on the writing assistant interface, where selection of a suggested prompt causes the suggested prompt to be populated in the input field of the writing assistant interface. These suggested prompts can be based on contextual signals obtained from a web page. For example, before the user submits a prompt, the writing assistant manager can generate a prompt suggestion request with one or more contextual signals and provide it to a generative model, which returns one or more suggested prompts to be displayed in the writing assistant interface. In some examples, the suggested prompts are selectable elements in the writing assistant interface. In some examples, in response to a selection of a suggested prompt, the writing assistant manager can transmit the selected (suggested) prompt and contextual signals to the generative model.
[0035] In some examples, the Writing Assistant Manager can selectively trigger the display of a pop-up widget, which a user can interact with to invoke the Writing Assistant interface. For instance, instead of the user directly invoking the Writing Assistant Manager (e.g., by selecting a menu item associated with it), the Writing Assistant Manager can selectively display a pop-up widget that informs the user about the Writing Assistant Manager to help draft the content of a text field. The pop-up widget can be a UI object displayed on a webpage near the text field. In response to a user selection of the pop-up widget (or a control on it), the Writing Assistant Manager can display the Writing Assistant interface, allowing the user to submit suggestions to the generative model for creating the text field.
[0036] The writing assistant manager determines whether and / or when to display the call-out feature (or, in some examples, the writing assistant interface). The writing assistant manager may include a grammar and / or machine learning (ML) model that receives one or more signals and determines, based on these signals, whether to render the call-out feature on the webpage. In some examples, these signals may include signals about text fields on the webpage, signals about the page content, and / or signals about previous use of the writing assistant interface for that webpage. In some examples, previous use signals may include one or more signals about whether a user has previously used the writing assistant interface (and / or previously disallowed the writing assistant interface) and / or one or more signals about whether other users have previously used the writing assistant on that particular text field.
[0037] In some examples, the generative model is a machine learning (ML) model. In some examples, the generative model is a pre-trained large language model (LLM). In some examples, the generative model is a specially trained language model. Generative models can generate high-quality responses to text fields. In some examples, generative models can be trained to generate responses to text fields for a specific category (type). Generative models use contextual signals from the webpage to generate the content of the text field. Generative models can use contextual signals from the webpage to determine the category associated with the text field (e.g., which category the text field represents). In some examples, a specially trained generative model for generating responses to text fields for a specific category can be smaller (e.g., in terms of required CPU and memory) and computationally faster (e.g., generating responses in short timeframes such as five or ten seconds) than a general-purpose large language model, and can generate more relevant and higher-quality responses that meet the expectations for the category of the text field. Such relevant and appropriate responses minimize user interaction for generating responses and provide a better human-computer guidance process for generating content.
[0038] Figures 1A to 1I A system 100 is illustrated with a composition assistant manager 110 of a browser application 108 that assists a user in generating content for one or more text fields 136 of a webpage 134. The composition assistant manager 110 can initiate a generative model 152 to generate a model response 124 for the text fields 136 of the webpage 134 and insert that model response 124 into (e.g., by direct input) the text fields 136. For example, a user can interact with the composition assistant manager 110 to help draft the content for the text fields 136 of the webpage 134. In some examples, the webpage 134 may be referred to as digital content. The term digital content can encompass web content, and in some examples, it can encompass non-web content.
[0039] The browser application 108 executable by the user device 102 can render web pages 134 on the display 126, such as... Figure 1A and Figure 1F As shown. Although Figure 1A The example depicts a webpage for writing comments, but webpage 134 can be any type of webpage 134. Furthermore, the techniques discussed herein are not limited to browser application 108, but can be any application that renders web content or, in some examples, non-web content. Webpage 134 includes a text field 136 configured to receive text input from a user. In some examples, text field 136 includes a free-form input field. A free-form input field includes an input field that receives unrestricted input from a user. In some examples, text field 136 includes an input field that receives structured data. In some examples, text field 136 includes a multi-line input field. In some examples, text field 136 includes a single-line input field.
[0040] To access the features of the Authoring Assistant Manager 110, the Authoring Assistant Manager 110 includes a trigger engine 112 configured to render a call-out feature 138 on the display 126 of the user device 102. The call-out feature 138 may be a user interface (UI) element, object, menu item, or control that identifies the Authoring Assistant Manager 110. In some examples, the user may directly access the call-out feature 138 using one or more controls provided by the browser application 108. For example, such as... Figure 1B As shown, trigger engine 112 can render the invoked feature 138 as a menu item 138b from menu 111 (e.g., "help me write"). In some examples, the user can right-click on text field 136 on webpage 134, and browser application 108 can display menu 111 (e.g., right-click menu) near text field 136, such as... Figure 1B As shown. Menu 111 may include menu item 138b, which, when selected, renders the writing assistant interface 128, as... Figure 1A and Figure 1D As shown.
[0041] In some examples, trigger engine 112 can selectively trigger the display of the invoked display element 138. For example, as Figure 1AAs shown, trigger engine 112 can display the invoked display 138 as a selectable UI object 138a. In some examples, trigger engine 112 detects user interaction with text field 136 (e.g., the user focuses on text field 136, such as placing the cursor on text field 136), and in response to the detected interaction, trigger engine 112 can render the selectable UI object 138a. User selection of the selectable UI object 138a causes the writing assistant manager 110 to render the writing assistant interface 128, such as... Figure 1C and Figure 1F As shown.
[0042] In some examples, trigger engine 112 can determine whether and / or when to display call-out display 138 (or, in some examples, writing assistant interface 128). In some examples, trigger engine 112 can detect trigger events for displaying writing assistant interface 128 based on one or more signals 180. In some examples, such as Figure 1H As shown, the trigger engine 112 includes a machine learning (ML) model 114, which is configured to receive signal 180 and calculate whether to display call indicator 138 (e.g., Figure 1B The prediction 188 of the selectable UI object 138a). In some examples, the triggering engine 112 uses one or more heuristics that use signal 180 to actively render the call-out display 138 (e.g., Figure 1B (Selectable UI object 138a). In some examples, the trigger engine 112 uses a combination of semantics and ML prediction to determine whether to display the call-out device 138.
[0043] In some examples, signal 180 includes text field signal 182 (e.g., a signal about text field 136 on webpage 134), content signal 184 (e.g., a signal about page content), and / or prior use signal 186 (e.g., a signal about prior use of the writing assistant manager 110). In some examples, prior use signal 186 may include one or more signals about whether a user has previously used the writing assistant manager 110 (and / or previously disallowed the writing assistant manager 110) and / or one or more signals about whether other users have previously used the writing assistant on that particular text field 136 or webpage 134.
[0044] Heuristics may include the results of existing autofill capabilities. For example, browser application 108 may include an autofill capability for text field 136 that has used various heuristics to identify target text fields important to its purpose. A heuristic for actively triggering a call to feature 138 may be when the autofill capability does not trigger suggestions (e.g., the autofill capability does not determine that text field 136 with focus is suitable for autofill suggestions). Heuristics may include that webpage 134 is a supported language. Heuristics may include that the Writer Assistant Manager 110 is not suppressed for reasons of support (e.g., the feature is disabled by the user, webpage 134 or the website (domain) is considered outside of policy, etc.). Heuristics may include that the use of Writer Assistant Manager 110 does not conflict with another browser feature. Heuristics may include that text field 136 is not relevant to enterprise or work productivity documents (e.g., word processing documents, PowerPoint presentations, etc.). Heuristics may include that text field 136 is not a prompt input box for a large language model (e.g., a text box designed to provide prompts (queries) sent to a large language model). With user permission, heuristics can take into account past user history (e.g., locally stored on the user's device). For example, if a user has used the writing assistant manager 110 on a review website but deactivated the call-out feature 138 on a social media site, the heuristic can enable the triggering engine 112 to render the call-out feature 138 for text fields related to product / service reviews rather than for web pages related to social media.
[0045] Trigger engine 112 may proactively render call-out enabler 138 using one or more heuristics in any combination. In some examples, trigger engine 112 may proactively render call-out enabler 138 in any combination of heuristics in response to trigger engine 112 detecting user interaction with text field 136 (e.g., focus being applied to text field 136). In some examples, trigger engine 112 may proactively render call-out enabler 138 in any combination of heuristics even if no user interaction with text field 136 is detected (e.g., no focus being applied to text field 136). In some examples, trigger engine 112 may render call-out enabler 138 in response to the amount of text data input by the user into text field 136 reaching a threshold level. When selected, the call-out feature 138 is configured to render a writing assistant interface 128 for text field 136, wherein the writing assistant interface 128 has an input field 130 configured to receive prompts 118 from the user.
[0046] In some examples, as indicated above, triggering engine 112 may include ML model 114 (or communicate with such ML model) to generate a prediction 188 regarding whether to render call-out display 138. If prediction 188 includes the probability that the user is likely to use the compose assistant manager 110, then triggering engine 112 may render call-out display 138. In some examples, ML model 114 may be trained using one or more heuristics (or any combination thereof) described herein to determine whether and when to trigger call-out display 138. For example, if the probability is high (a first threshold is met), triggering engine 112 may trigger call-out display 138 (e.g., when text field 136 receives focus). If the probability is neither high nor low (the first threshold is not met but a second threshold is met), then triggering engine 112 may trigger call-out display 138 if the user has typed several characters or words in text field 136 but then stopped.
[0047] refer to Figure 1F The writing assistant manager 110 includes a prompt manager 116. The prompt manager 116 generates a context signal 120 about the webpage 134. The context signal 120 may be referred to as context data. The context data includes information about the topic of the webpage 134. In some examples, the prompt manager 116 generates the context signal 120 in response to the writing assistant manager 110 being invoked (e.g., when call-up widget 138 is selected, and / or when the writing assistant interface 128 is rendered). In some examples, the prompt manager 116 generates the context signal 120 after call-up widget 138 is rendered (e.g., UI object 138a) and before call-up widget 138 is selected. In some examples, the prompt manager 116 generates the context signal 120 in response to the selection of a build control 131 on the writing assistant interface 128.
[0048] like Figure 1G As shown, context signal 120 may include the page title 172 of webpage 134, the page content 170 associated with webpage 134, and / or the resource locator 176 of webpage 134. In some examples, context signal 120 includes DOM representation 178. In some examples, context signal 120 includes an accessible content structure 174. The accessible context structure 174 may be referred to as an accessibility tree.
[0049] The prompt manager 116 provides a technical solution for generating context signals 120 about the underlying webpage 134, wherein the context signals 120 are used to help the generative model 152 create context-relevant responses. The prompt manager 116 performs context extraction, which extracts the page content of the webpage 134 in a manner that maintains the security of the webpage 134.
[0050] In some examples, such as Figure 1I As shown, page content 170 includes page content 170-1 of webpage 134 (e.g., a first webpage) and page content 170a of one or more embedded resources 139 (e.g., embedded in the structure of webpage 134). For example, prompt manager 116 can retrieve page content 170-1 of webpage 134 having a text field 136, and can retrieve page content 170a of one or more embedded resources 139 (e.g., webpages). For example, webpage 134 may embed resource 139-1 (e.g., a second webpage) and resource 139-2 (e.g., a third webpage). Prompt manager 116 can retrieve page content 170-1 of webpage 134, page content 170-2 of resource 139-1, and page content 170-3 of resource 139-2. In some examples, page content 170-1, page content 170-2, or page content 170-3 may be referred to as internal text. Retrieving page content from embedded resources 139 (e.g., webpages) may present one or more technical challenges, such as security risks.
[0051] In other words, webpage 134 includes one or more inline frames (e.g., iframes) (e.g., hypertext markup language (HTML) elements that embed another HTML document (e.g., resource 139-1 or resource 139-2) into the current page (e.g., webpage 134). The prompt manager 116 performs a context extraction on context signal 120, which overcomes technical challenges by requesting the internal text of a specified host (e.g., webpage 134) and the internal text of local same-origin iframes (e.g., all local same-origin iframes). A same-origin iframe can be an iframe that shares the same origin as the main webpage (e.g., webpage 134). The origin of a webpage is determined by its protocol, hostname, and port number. In some examples, the embedded iframe resides on the same server or domain as the main webpage. In some examples, the embedded iframe may have the same protocol, hostname, and / or port number as the main webpage. Internal text can refer to the visible text content within an HTML element and text from one or more child elements of the HTML element. The returned internal text includes the combined internal text of iframes (e.g., all suitable iframes). The prompt manager 116 retrieves the internal text of the webpage 134, and when each iframe is detected, the internal text of the webpage is combined with the internal text of the iframe (e.g., embedded resource 139).
[0052] refer to Figure 1FIn some examples, the prompt manager 116 includes an ML model 122. The ML model 122 may receive context signals 120 as input, such as the page title 172 of webpage 134, page content 170 associated with webpage 134, resource locators 176 of webpage 134, DOM representation 178, and accessible content structure 174. The ML model 122 may use the context signals 120 to generate context data (or use first context data (e.g., a larger set of content data) to generate second context data (e.g., a smaller set of content data)), wherein the context data generated by the ML model 122 includes a smaller subset of information than the context signals 120, and this context data is provided to the generative model 152. In some examples, the ML model 122 selects a subset of the information contained in the context signals 120 and provides that subset to the generative model 152. In some examples, the ML model 122 generates a summary of the context signals 120 and provides that summary to the generative model 152. By using ML model 122 to generate or select a portion of the context signal 120, a smaller set of information can be provided to generative model 152, which can provide one or more technical benefits for reduced computational costs inferred by generative model 152. In other words, the lexical size of the cues provided to generative model 152 can be reduced, which reduces the computational cost of the generative model response 124.
[0053] refer to Figure 1C and Figure 1F The writing assistant interface 128 includes an input field 130 configured to receive prompts 118 from the user. In some examples, the prompt 118 is referred to as text data, such as data entered by the user. The user can provide the prompt 118 in the writing assistant interface 128 (e.g., “write a five-star review about this product”). The prompt 118 can be a natural language description of the type of content to be generated by the generative model 152. For example, the user can type the prompt 118 or provide a voice command to insert the prompt 118 into the input field 130.
[0054] refer to Figure 1CThe writing assistant interface 128 may include a generation control 131. In response to a user selection on the generation control 131, the prompt manager 116 may transmit a prompt 118 and a context signal 120 to the generative model 152. In some examples, the generation control 131 may be inactive until the user provides text in the input field 130. Therefore, the generation control 131 may be active (and selectable) after the user has entered text in the input field 130. In some examples, the writing assistant interface 128 may include options (e.g., in a three-dot menu or similar) to enable or disable the writing assistant manager 110.
[0055] In response to cue 118 and context signal 120, generative model 152 can generate model response 124. Cue manager 116 can receive model response 124 from generative model 152 and display model response 124 in interface 133 of writing assistant interface 128, such as... Figure 1D As shown.
[0056] like Figure 1D As shown, the writing assistant interface 128 may include one or more UI elements that allow the user to adjust the model response 124 (e.g., more formal, less formal, expanded, shortened, etc.), causing the generative model 152 to regenerate the model response 124. In some examples, the user can manually edit the model response 124. (Reference) Figure 1D The writing assistant manager 110 may include an insertion control 141, which, when selected, causes the model response 124 to be inserted into the text field 136 on the webpage 134, such as... Figure 1E and Figure 1F As shown. For example, in response to the selection of the insert control 141, the compose assistant manager 110 transfers text from the compose assistant interface 128 to the text field 136 of the web page 134.
[0057] like Figure 1DAs shown, the writing assistant interface 128 may include an insert control 141. The insert control 141 inserts text into the text field 136, replacing previously written text if the user has previously written text (as opposed to starting from scratch). If the user has written text, the insert control 141 can say "Replace"; otherwise, it can say "Insert this". If the user clicks "Replace" but only selects a portion of the text (e.g., by highlighting a portion of the text), only that text is replaced (as opposed to the entire text in the field). In some examples, the insert control 141 can close the writing assistant interface 128. This could be a local signal used in a personal heuristic approach, as discussed above, if the user closes the writing assistant interface 128, for example, by selecting the close control 127, before selecting the generate control 131. In other words, with the user's permission, the triggering engine 112 can use this type of close event to determine when to actively display the call-out feature 138. Figure 1E The model response 124 inserted into text field 136 of webpage 134 is shown. Users can edit the response in text field 136 of webpage 134.
[0058] The writing assistant interface 128 may also include controls for modifying (editing) the model response 124 using the generative model 152. For example, as Figure 1D As shown, the writing assistant interface 128 may include a tone control 123. The tone control 123 allows the user to make the response sound more formal, casual, fun, or even include emojis. The writing assistant interface 128 may include a length control 125. The length control 125 allows the user to shorten or lengthen the model response 124. In response to the user selecting either the tone control 123 or the length control 125, the writing assistant manager 110 can provide a new model response 124. In other words, selecting either the tone control 123 or the length control 125 can cause the generative model 152 to regenerate the model response 124 based on the value of the selected control. In some examples, if the "lengthen" length control 125 is selected more than once, the writing assistant manager 110 may suggest that the user use a general large language model (such as Bard, chat GPT) for a better back-and-forth interactive experience.
[0059] In some examples, such as Figure 1D As shown, the writing assistant interface 128 may include a regeneration control 135. The regeneration control 135 can generate another text suggestion (e.g., a new model response 124). The writing assistant interface 128 may include a back control 147. The back control 147 allows the user to go back and... Figure 1DThe prompts 118 are edited in the writing assistant interface. The writing assistant interface 128 may include a close control 115. The close control 115 can close the writing assistant interface 128 and return the user to the text field 136. In some examples, the writing assistant manager 110 may store information related to the writing assistant manager 110's use of the webpage 134 (e.g., a specific text field 136) to help determine whether to render the call-out presentation 138 for the user or other users in the future.
[0060] In some examples, the writing assistant interface 128 includes a feedback mechanism 129. Feedback mechanism 129 allows users to rate text suggestions. With user permission, the rating can be used for additional training (e.g., a dislike or low rating can be used as an example of a suggestion not to be generated). With user permission, the rating can also be used to trigger the writing assistant manager 110 for that user. Therefore, some implementations allow users to rate suggested text output to help improve future suggestions. Although in Figure 1E The diagram illustrates a binary (like / dislike) feedback mechanism 129, but numerical scales (e.g., the number of stars, the selection of one of multiple ratings, etc.) can also be used.
[0061] refer to Figure 1F The writing assistant interface 128 can provide the user with one or more suggested prompts 118a that the user can select and / or edit. In other words, before the user begins drafting a prompt 118 in the input field 130 on the writing assistant interface 128, the writing assistant manager 110 can provide selectable suggested prompts 118a, where selection of a suggested prompt 118a causes the suggested prompt 118a to be populated in the input field 130 of the writing assistant interface 128. These suggested prompts 118a may be based on context signals 120 obtained from the webpage 134. For example, before the user submits a prompt 118, the writing assistant manager 110 may generate a prompt suggestion request with one or more context signals 120 and provide it to a generative model 152, which returns one or more suggested prompts 118a to be displayed in the writing assistant interface 128. In some examples, the suggested prompts 118a are selectable elements in the writing assistant interface 128. In some examples, in response to the selection of suggested cue 118a, the writing assistant manager 110 can transmit the selected (suggested) cue 118a and context signal 120 to the generative model 152.
[0062] In some examples, the suggested prompt 118a is a generic prompt instructing the writing assistant manager 110 to help them write. In some examples, if the user has not yet started writing and invokes the writing assistant manager 110, the writing assistant manager 110 may render a rotating set of suggested prompts 118a based on context signal 120 (e.g., the set of approximately 5 prompts may vary if the user is writing a comment, social media commentary, or filling out a form). Therefore, the suggested prompt 118a may use a page context or may be generic. The page context may include values that the user has already provided for other fields, such as the number of stars the user has provided. The page context may include insights from other content on the webpage. In some examples, the writing assistant manager 110 may analyze the user's writing history in profile 155 (e.g., a local profile) (generated with the user's permission) and / or open tabs to provide personalized prompts, thereby ensuring relevance and resonance with the user's intended audience. Reliance on user history helps maintain a consistent tone in the responses generated for the user.
[0063] If a user is reviewing a product, generative model 152 can return a well-structured review, even if the user hasn't explicitly specified it as a review in prompt 118. This context-aware approach is beneficial even before the user types anything. For example, implementations can support zero-state use cases that provide UI suggestions that include general text input. For instance, when a user is viewing a review page, the implementation could provide the suggestion "write a constructive review." In some implementations, generative model 152 can be further trained to provide input suggestions that are also context-aware. In this case, a zero-state example could be "write a 4-star review about this dining table" if the user is on a review page for a wooden dining table, or "write a review about..." if the user is viewing a webpage for a <product> that is not a review page (e.g., a customer complaint page). <product>that does not work as intended (write a comment about the product not working as expected).
[0064] The disclosed implementation also reduces user interaction with user device 102 to insert generated text into text field 136 of webpage 134. Specifically, other large language models are not integrated into browser application 108, and generated responses must be copied and pasted. The disclosed implementation assists the user directly where they are writing. Initial text input comes directly from the text field of the webpage the user is working on, and once the generative model 152 provides generated output text (response) deemed acceptable to the user, that output text is directly inserted into the same text field 136. The disclosed implementation can generate relevant ideas to get users started writing, adapt responses to user speech, and provide users with a draft for editing. Whether the user is someone who likes to share a humorous commentary on a piece of web content with their friends, someone trying to file a complaint with a store, or someone who simply wants to create a more personal RSVP note for a wedding invitation, the writing assistant manager 110 can be a reliable writing assistant directly built into browser application 108.
[0065] User device 102 can be any type of computing device including one or more processors 101, one or more memory devices 103, a display 126, and an operating system 105 configured to execute (or assist in executing) one or more applications 106 (including browser application 108). In some examples, browser application 108 is a web browser configured to access information on the Internet. Browser application 108 can launch one or more browser tabs in the context of one or more browser windows on the display 126 of user device 102. Browser tabs can display content (e.g., web content) associated with web documents (e.g., web pages, PDFs, images, videos, or any items typically identifiable by resource locators) and / or applications such as web applications, progressive web applications (PWAs), and / or extensions. Web applications can be applications stored on a remote server (e.g., a web server) and delivered over network 150 via browser application 108. In some examples, progressive web applications are similar to web applications but can also be (at least partially) stored on user device 102 and used offline. Extensions add features or functions to browser application 108. In some examples, extensions can be based on HTML, CSS, and / or JavaScript (for browser-based extensions).
[0066] In some examples, user device 102 is a laptop computer. In some examples, user device 102 is a desktop computer. In some examples, user device 102 is a tablet computer. In some examples, user device 102 is a smartphone. In some examples, user device 102 is a wearable device. In some examples, display 126 is the display of user device 102. In some examples, display 126 may also include one or more external monitors connected to user device 102.
[0067] Processor 101 may be formed in a substrate configured to execute one or more machine-executable instructions or one or more pieces of software, firmware, or a combination thereof. Processor 101 may be semiconductor-based—that is, the processor may include semiconductor material capable of performing digital logic. Memory device 103 may include main memory storing information in a format that can be read and / or executed by processor 101. Memory device 103 may store browser application 108, a composition assistant manager 110 (and in some examples, generative model 152), which, when executed by processor 101, performs certain operations discussed herein. In some examples, memory device 103 includes a non-transitory computer-readable medium comprising executable instructions that cause at least one processor (e.g., processor 101) to perform operations. In some examples, composition assistant manager 110 may be configured to communicate with one or more generative models 152. In some examples, composition assistant manager 110 may enable a user to select one of a plurality of generative models 152 for generating input to text field 136, wherein the plurality of generative models 152 include different LLMs. For example, the writing assistant interface 128 may provide a first selectable option associated with a first generative model and a second selectable option associated with a second generative model. In response to the selection of the first selectable option, the writing assistant manager may provide a prompt 118 and a context signal 120 to the first generative model. In response to the selection of the second selectable option, the writing assistant manager may provide a prompt 118 and a context signal 120 to the second generative model.
[0068] Server computer 160 may be a computing device in various forms (e.g., a standard server, a group of such servers, or a rack server system). In some examples, server computer 160 may be a single system sharing components such as processors and memory. In some examples, server computer 160 may be multiple systems that do not share processors and memory. Network 150 may include the Internet and / or other types of data networks, such as local area networks (LANs), wide area networks (WANs), cellular networks, satellite networks, or other types of data networks. Network 150 may also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) configured to receive and / or transmit data within network 150. Network 150 may further include any number of hardwired and / or wireless connections.
[0069] Server computer 160 may include one or more processors 161, an operating system (not shown), and one or more memory devices 163 formed in a substrate. Memory devices 163 may represent any type (or multiple types) of memory (e.g., RAM, flash memory, cache, disk, tape, etc.). In some examples (not shown), the memory device may include external storage, such as memory physically located away from server computer 160 but accessible by the server computer. Processor 161 may be formed in a substrate configured to execute one or more machine-executable instructions or one or more pieces of software, firmware, or a combination thereof. Processor 161 may be semiconductor-based—that is, the processor may include semiconductor material capable of performing digital logic. Memory device 163 may store information in a format readable and / or executable by processor 161. In some examples, memory device 163 may store a generative model 152 that, when executed by processor 161, performs certain operations discussed herein. In some examples, memory device 163 includes a non-transitory computer-readable medium comprising executable instructions that cause at least one processor (e.g., processor 161) to perform operations.
[0070] In addition to the description above, users can be provided with controls that allow them to choose whether and when the system, program, or feature described herein can enable the collection of user information (e.g., information about the user's browser history, user preferences, current location, or other profile information) and whether the functions described herein are active. Furthermore, before storing or using certain data, it can be processed in one or more ways to remove personally identifiable information. For example, user identity can be processed to make it impossible to determine the user's personally identifiable information, or, if location information is available, the user's geographic location can be generalized (e.g., to the city, zip code, or state level) to make it impossible to determine the user's specific location. Therefore, users can control what information is collected, how that information is used, and what information is provided to the user.
[0071] Figure 2 A writing assistant interface 228 is shown based on one aspect. The writing assistant interface 228 can be... Figures 1A to 1I This is an example of the writing assistant interface 128, and any details discussed with reference to these figures can be included. Figure 2 As shown, the writing assistant interface 228 includes an input field 230 configured to receive prompts from the user. The writing assistant interface 228 displays suggested prompts 218a in the input field 230 that the user can select and / or edit.
[0072] In other words, before the user begins drafting suggestions in input field 230 on the writing assistant interface 228, the writing assistant manager (e.g., Figures 1A to 1I The writing assistant manager 110 can provide suggested hints 218a. The suggested hints 218a can be generated by a generative model (e.g., Figures 1A to 1I The generative model 152) is based on one or more field signals (e.g., Figures 1A to 1I The field signal 120 is used to generate it.
[0073] refer to Figure 2 The writing assistant interface 228 may include a generation control 231. In response to a user selection on the generation control 231, the writing assistant manager may transmit prompts and context signals to the generative model. In some examples, the generation control 231 may be inactive until the user provides text in the input field 230. Therefore, the generation control 231 may be active (and selectable) after the user has entered text in the input field 230. In some examples, the writing assistant interface 228 includes a close control 217. The close control 217 can close the writing assistant interface 228 and return the user to the text field.
[0074] Figure 3 An authoring assistant interface 328 according to another aspect is shown. In some examples, the user can invoke an authoring assistant manager according to any of the technologies discussed herein (e.g., Figures 1A to 1I The writing assistant manager 110 can display a writing assistant interface 328. In some examples, the writing assistant interface 328 can identify a set of categories 362 (e.g., types) of text fields of web page 334. The user can select one of the categories 362 from the set of categories. In some examples, the writing assistant interface 328 can identify a tone control 364 that allows the user to select the tone of the model response to be generated by the generative model. The writing assistant interface 328 may include an input field 330 configured to receive prompts from the user. The writing assistant interface 328 may include a generation control 331. When the user selects the generation control 331, the generation control causes the writing assistant manager to transmit prompts, user selections made via the writing assistant interface 328, and context signals generated by the writing assistant manager.
[0075] Figures 4A to 4C An example of a writing assistant interface 428 according to one aspect is shown. The writing assistant interface 428 can be rendered on a webpage 434 (e.g., a social media webpage) to help a user write paragraphs for text field 436 on webpage 434. In some examples, the writing assistant interface 428 can be rendered when the writing assistant manager is invoked. The writing assistant manager can be invoked according to any of the techniques discussed herein.
[0076] like Figure 4A As shown, the writing assistant interface 428 includes an input field 430 configured to receive prompts 418 from a user. The user can provide prompts 418 within the writing assistant interface 428. Prompts 418 can be a natural language description of the type of content to be generated by the generative model. For example, the user can type prompts 418 or provide a voice command to insert prompts 418 into the input field 430.
[0077] The writing assistant interface 428 may include a generation control 431. In response to a user selection on the generation control 431, the writing assistant manager (e.g., Figures 1A to 1I The writing assistant manager 110 can provide prompts 418 and contextual signals (e.g., Figures 1A to 1I The field signal 120) is transmitted to the generative model (e.g., Figures 1A to 1I (Generative model 152). In some examples, the generation control 431 may be inactive until the user provides text in the input field 430. In response to prompts 418 and context signals, the generative model can generate a model response 424. The writing assistant manager can receive the model response 424 from the generative model and display the model response 424 in the writing assistant interface 128, such as... Figure 4C As shown.
[0078] like Figure 4C As shown, the writing assistant interface 428 may include an insertion control 441. The insertion control 441 inserts text into the text field 436. In some examples, the insertion control 441 can close the writing assistant interface 428. The writing assistant interface 428 may also include controls for modifying (editing) the model response 424 using a generative model. For example, the writing assistant interface 428 may include a tone control 423. The tone control 423 allows the user to make the response sound more formal, casual, fun, or include emojis. The writing assistant interface 428 may include a length control 425. The length control 425 allows the user to shorten or lengthen the modal response 424. In response to the user selecting either the tone control 423 or the length control 425, the writing assistant manager can provide a new model response 424. In other words, the selection of the tone control 423 or the length control 425 can cause the generative model to regenerate the model response 424 based on the value of the selected control.
[0079] In some examples, the writing assistant interface 428 may include a regeneration control 435. The regeneration control 435 may generate another text suggestion (e.g., a new model response 424). The writing assistant interface 428 may include a close control 415. The close control 415 may close the writing assistant interface 428 and return the user to the text field 436.
[0080] Figures 5A to 5F An example of a writing assistant interface 528 according to one aspect of a writing assistant manager is shown. The writing assistant interface 528 can be rendered on a display relative to a text field 136 of a webpage. The writing assistant interface 528 can be triggered according to any of the techniques discussed herein. In some examples, the writing assistant interface 528 can be a UI dialog box.
[0081] like Figure 5B As shown, the writing assistant interface 528 can display a loading status while generating initial writing suggestions 524. In some examples, the initial writing suggestions 524 can begin to be generated after the user has already written a threshold level of words (the amount of text data that reaches the threshold level) in the text field 538. In some examples, such as Figure 5C As shown, after the user selects a threshold number of words in text field 536, a writing assistant interface 528a can be displayed, whereby the writing assistant interface 528a can provide the user with a set of actions 550, such as proofreading action 540 and polishing action 542. In some examples, the writing assistant interface 528a may include an expander control 544 that provides additional actions when selected.
[0082] In some examples, such as Figure 5D As shown, initial writing suggestions 524 can be displayed in the writing assistant interface 528. In some examples, the writing assistant manager can combine text in text field 536 with contextual signals (e.g., Figures 1A to 1I The field signal 120 is transmitted to the generative model. The generative model can generate a model response with initial writing suggestions 524. In some examples, such as Figure 5E As shown, the user can hover the cursor over the writing assistant interface 528, which provides a preview of the initial writing suggestion 524 in the text field 536. In some examples, the writing assistant manager can detect cursor positioning on the suggestion (e.g., the initial writing suggestion 524), and in response to the cursor being positioned within the boundaries of the suggestion, the writing assistant manager can provide a preview of the suggestion in the text field 536. In some examples, such as Figure 5F As shown, in response to the user moving the cursor on the expander control 544, the writing assistant manager can render and display an action menu 562 that shows a set of actions 550.
[0083] Figure 6 A writing assistant interface 628 is shown according to one aspect. The writing assistant interface 628 includes a prompt field 618 that displays a suggestion, and an editing control 660 that allows the user to edit the prompt when selected. The writing assistant interface 628 displays a model response 624. The writing assistant interface 628 may display a series of controls, such as a refinement control 670 (which, when expanded, can display controls related to shortening, length, tone adjustment, etc.), an undo control 671, and a redo control 635. The writing assistant interface 628 may include an insert control 641 that, when selected, inserts the prompt into a text field of the webpage.
[0084] Figure 7 A writing assistant interface 728 is shown according to one aspect. The writing assistant interface 728 includes a prompt field 718 that displays a suggestion, and an editing control 760 that allows the user to edit the prompt when selected. The writing assistant interface 728 displays a model response 724. The writing assistant interface 728 may display a series of controls, such as a length control 725, a tone control 723, and a redo control 735. The writing assistant interface 728 may include an insert control 741 that inserts the prompt into a text field of the webpage when selected.
[0085] Figure 8 This is a diagram illustrating a system 800 having a user device 802 and a server computer 860 for implementing the concepts described herein. Typically, the user device 802 can represent any computing device executing a browser application 808. Figure 8 As shown, user device 802 is configured to communicate with server computer 860 and / or resource provider (e.g., web server) via network 850. User device 802 includes at least browser application 808 and other applications (not shown). In some implementations, browser application 808 is configured to manage resource content, such as web page content, provided by resource provider (e.g., web server). In some implementations, browser application 808 is configured to operate as one of several applications executed via operating system (O / S) 802.
[0086] although Figure 8 Not shown, but user device 802 includes several hardware components, including a communication module, one or more cameras, memory, a processing unit 801 such as a central processing unit (CPU) and / or a graphics processing unit (GPU), one or more input devices 867 (e.g., touchscreen, mouse, stylus, microphone, keyboard, etc.), and one or more output devices 868 (screen, speaker, vibrator, light emitter, etc.). The hardware components can be used to facilitate the operation of a browser application 808 and / or other applications on user device 802. User device 802 may also include an operating system 805. Browser application 808 includes a writing component 810 configured to generate, for example, a writing assistant user interface as shown in the various figures.
[0087] User device 802 may include local user profile data 855. Local user profile data 855 may be stored in memory associated with browser application 808, or in memory accessible to browser application 808. Local user profile data 855 may be a data source (or multiple sources) of user-specific information collected with user permission from the user's use of browser application 808. Local user profile data 855 is stored on the device. In some implementations, local user profile data 855 may be associated with an account profile, such as a user account on server computer 860. In such implementations, some information may be stored in central user profile data 842. The user can control what information is shared between local user profile data 855 and central user profile data 842, and when information is shared. Sharing data from local user profile data 855 (e.g., signals that help the writing assistant know when to trigger a call, signals that help define tone for the user) with central user profile data 842 allows the user to have a consistent experience with the writing assistant across user devices.
[0088] Browser application 808 includes a composer renderer assistant 827. The composer renderer assistant 827 runs in the renderer process and performs operations related to webpage 834 and text field 836. The composer renderer assistant 827 may include a webpage interaction component responsible for the user experience flow's interaction with webpage 834, such as monitoring user interaction with text fields (e.g., text field 836), triggering the rendering of call-up widgets, extracting and inserting text from text fields, etc. The composer renderer assistant 827 may include a context extraction component equipped to capture a set of signals to help generate a response (text) to the text field. Once the user requests an LLM response, the context extraction component extracts all expected context from the page to package it along with the prompt. Context may include URLs, titles, and / or page content and / or other signals described herein. For page content, the system may utilize different methods to determine the most relevant parts of the content. For example, the composer renderer assistant 827 can utilize the DOM (Document Object Model) or accessibility tree to identify which parts of the content are visible or invisible, which parts of the content surround input fields (e.g., text fields 836), and other key content parts of the page (such as title fields). In such an implementation, the content extraction component can extract DOM portions from the DOM representation of the context. In the context of a page related to a dialogue, this context includes previous turns in the dialogue.
[0089] The context may include the main entity of the webpage. For example, if the webpage is a review page for a vacuum cleaner, the main entity could be the vacuum cleaner or a particular vacuum cleaner. In some implementations, the generative language model can be trained to recognize the main entity in the content of the webpage provided as a context. Additionally, the composer renderer assistant 827 can identify text input fields and utilize their metadata, which can be used to categorize the possible purposes of the text input fields within the context of the page paired with the user-provided prompts. The browser application 808 will extract the raw signals and process signals useful for creating the correct context to be used by the composer generative language model 852. This includes recognizing the page, form, and type of input field the user is typing. The context can also be obtained from a website (e.g., a domain in which the webpage is part). For example, in an implementation where the server computer 860 is associated with a search engine and the website is indexed, the content of the search index from that domain can be used as a context signal.
[0090] With the user's consent, the context may also include user history signals. User history signals may include previously generated responses, for example, so that the composition generative language model 852 can mimic tone. For example, the context may include cue packs to provide few-shot training for the composition generative language model 852. Cue packs are used to bias the composition generative language model 852 towards generating responses that are more similar to how that particular user has formatted responses in the past. In some implementations, cue packs may be stored as state, for example, in local user profile data 855. User history signals may include metadata from shopping history. For example, if a browser application 808 is enabled to access shopping history, and the webpage is a review of a product purchased by the user (e.g., the user clicked a link in an email requesting the user to leave a review; in this case, the webpage may be part of a custom tab associated with the email application), the shipping time may be known or computable, and this information may be added as part of the context and utilized by the composition generative language model 852 when generating the review (response). Similarly, flight information may be used to respond to car rental or hotel instructions. Browser application 808 may include a settings user interface. The user interface settings can include menus where users can enable and disable the writing assistant.
[0091] Figure 8 Some aspects of server computer 860 are illustrated. For example, server computer 860 may include an authoring service 844, a security / policy filter 846, and an authoring generative language model 852. Server computer 860 also includes one or more processors (not shown) and one or more memory devices (not shown). Authoring service 844 may be server-side business logic responsible for querying all dependent services and data sources to serve a user's request, including collecting any further user data from central user profile data 842 and requesting inference from authoring generative language model 852. Security / policy filter 846 may ensure that information received from and sent back to user device 802 complies with all security, policy, and legal requirements, such as avoiding sensitive categories and filtering unsafe content. In some implementations, policy filter 846 may be a known classifier that identifies the type of context (adult content / violence / aggression) of negative / inappropriate prompts and / or pages. Policy filter 846 may be targeted at the prompt and its context and operate on the output from authoring generative language model 852. The policy filter 846 can prevent the writing generative language model 852 from providing output to the user, and instead return an error message indicating that the prompt cannot be processed.
[0092] The writing generative language model 852 is a generative language model custom-trained for a writing assistant to adapt to the use cases targeted by the features. This user case is based on the purpose or type of the text fields. For example, purpose / type could be reviews (products, locations, travel, etc.), comments (e.g., on videos or articles), social media posts, survey responses, forums, replies in conversations (e.g., conversations with chatbots or messaging apps), customer complaints, blogs, profile descriptions, etc. Training enables the writing generative language model 852 to appropriately consider additional context extracted from web pages 834 and from local user profile data 855 and / or central user profile data 842. Training also enables the writing generative language model 852 to generate responses tailored to this purpose, for example, to generate responses similar in length to average product reviews, average social media posts, average forum contributions, etc. Therefore, the writing generative language model 852 can utilize input signals (context about web pages, text input fields, and / or the user) to customize the output based on provided browser signals. The generative writing language model 852 can therefore be fine-tuned to produce the correct writing structure based on the context and the prompts provided to the user. For example, if the user is on a product review page and is given limited prompts, the system (e.g., browser application 808) can add enough context that the text generated from the generative writing language model 852 will be a structured review containing details from the user's prompts. The writing assistant thus provides a solution for leveraging the page context to prompt the user based on goals, categories, topics / themes, etc.
[0093] In some implementations, the composing assistant manager can leverage previous prompts and submission examples from user interactions (e.g., stored in local user profile data 855) to further personalize the voice for the user. This is known as prompt packaging. This ensures that tone and voice are more consistent across individual user experiences. With the user's permission, these additional user signals can be synchronized with the user profile across devices (e.g., user device 802) on which the user is logged in to ensure consistency of tone and voice across user devices.
[0094] In some implementations, the compose generative language model 852 can be configured to generate responses with variable placeholders. For example, if the textbox is part of a conversation (e.g., a message in an instant messaging conversation or a chatbot chat), the user may be responding to a request for a specific piece of information (e.g., a fact). The request for a fact can be part of a context provided to the compose generative language model 852. The compose generative language model 852 can be configured to generate appropriate variable placeholders in the generated response for the specific piece of information. Thus, for example, the response could be "Thanks! You can reach me at [phone number] after 5pm", where [phone number] These are variable placeholders that can be edited by the user.
[0095] In some implementations, with user permission, additional user context available within the browser (via profiles and other user data storage—represented by local user profile data 855) can allow the writing assistant to automatically fill in variable placeholders in the generated response. For example, when replying to a post about contact information, the writing generative language model 852 can output... [user_x_address] The browser application 808 then utilizes the variable placeholder by viewing the available contact information in the local user profile data 855 to pre-populate the value into the generated response.
[0096] The Generative Language Model 852 can be trained using examples of different types of input fields; that is, it can be trained using input fields with different purposes. Because the Generative Language Model 852 is trained for a specific task, it can be small and provide output faster than general-purpose large language models.
[0097] In some implementations, a server computer 860 is not required because the generative language model 852 runs on the device—for example, on the user device 802. In such an implementation, the functions performed by the writing service 844 and / or the policy filter 846 can be performed by one of the writing assistant components of the browser application 808 (e.g., writing component 810 and / or writing renderer assistant 827).
[0098] Figure 9 This is a flowchart illustrating an example process 900 for providing a writing assistant, depending on the implementation method. Process 900 can be executed by a browser's writing assistant manager, such as... Figures 1A to 1I Browser application 108 and / or Figure 9 The browser application 808. At step 902, the system receives focus on the text box of the webpage. At step 904, the system determines whether to display a call-out feature configured to initiate a writing assistant for the text box. At step 906, in response to determining to display the call-out feature, the system provides the user with suggested prompts in the writing assistant interface.
[0099] Figure 10 This is a flowchart illustrating a sample process 1000 for providing a composition assistant manager, based on its implementation. Process 1000 can be executed by the browser's composition assistant manager, such as... Figures 1A to 1I Browser application 108 and / or Figure 8 The browser application 808. At step 1002, the system can receive prompts from the user related to input into a text box on a webpage. At step 1004, the system can generate context signals for the webpage. Context signals may include content signals. Context signals may include user signals. At step 1006, the system can provide the prompts and context signals to a generative language model trained to provide output for the type (purpose) of the text box. At step 1008, the system can receive a response generated by the generative language model. At step 1010, for example, in response to a selection of an accept control, the system can provide a response as input to the text box. Therefore, using process 1000, the user minimizes interaction with the computing device, and the model is able to generate high-quality output suitable for the purpose of the text box.
[0100] Below are example use cases for the disclosed implementations. These use cases are non-limiting examples. These implementations can assist users in solving specific problems, such as writer's block. For example, a user who likes to share content to stay connected with friends and family sees something interesting but doesn't know how to share it humorously on a social media platform. A writing assistant can help this user draft the content to share. As another example, a user may have recently had a negative experience with an airline and wants to file a complaint. A writing assistant can help in expressing their concerns effectively and professionally. As another example, a user may be a blogger and social media influencer but is experiencing a writing block and needs inspiration for their next blog post or social media post. They want to ensure that the content they produce is engaging and relevant to their audience. As another example, a user may have recently moved to an English-speaking country and needs to write emails, job applications, and other documents in English. A writing assistant can help them draft these. As another example, a user may be a brand manager who needs to generate daily content inspiration for the company they represent. They must maintain a consistent brand voice across various platforms. A writing assistant can ensure cross-platform consistency with their consent. As another example, the user might be a college student who needs to write a research paper and is struggling to organize their thoughts.
[0101] Figure 11 This is a flowchart 1100 illustrating example operations of a system for integrating a language model in a browser application, based on one aspect. Flowchart 1100 can depict the operations of a computer-implemented method. Flowchart 1100 can be applied to any implementation discussed herein. Although Figure 11 Flowchart 1100 illustrates the operations in sequence; however, it should be understood that this is merely an example and may include additional or alternative operations. Furthermore, Figure 11 The operations and related operations can be performed in a different order than those shown, or in a parallel or overlapping manner.
[0102] Operation 1102 includes receiving a prompt from the user in relation to input to a text field for a webpage. In some examples, the prompt is referred to as text data, and the webpage is referred to as digital content. Operation 1104 includes generating a context signal for the webpage. In some examples, the context signal is referred to as context data. Operation 1106 includes providing the prompt and context signal to a generative language model. Operation 1108 includes receiving a response generated by the generative language model. Operation 1110 includes using the response as input to a text field. In some examples, operation 1110 includes providing a suggestion of the response as input to a text field. In some examples, in response to acceptance of the response, the application can directly insert the response into the text field.
[0103] Clause 1. A method comprising: receiving from a user text data relating to input to a text field of digital content displayed on a user device; generating context data about the digital content; providing the text data and the context data to a generative language model; receiving a response generated by the generative language model; and providing the response as a suggestion for input to the text field.
[0104] Clause 2. The method as described in Clause 1 further includes: detecting interaction with the text field; and determining by the model whether to render a call-out feature, which, when selected, is configured to render a writing assistant interface for the text field, the writing assistant interface having an input field configured to receive text data from the user.
[0105] Clause 3. The method as described in Clause 2 further includes: determining whether to render the call-out device based on signals, including one or more signals regarding text fields, one or more signals regarding digital content, or one or more signals regarding the writing assistant's user and other users.
[0106] Clause 4. The method of Clause 1 further includes: in response to the amount of text data entered by the user into the text field reaching a threshold level, rendering a call-out display, which, when selected, is configured to render a writing assistant interface for the text field, the writing assistant interface having an input field with the text data.
[0107] Clause 5. The method of Clause 1 further includes: receiving a selection of a user interface object for a text field concerning digital content; rendering a writing assistant interface for the text field, the writing assistant interface having an input field configured to receive text data from a user; and, in response to a selection of a generation control of the writing assistant interface, transferring the text data and the context data to a generative language model.
[0108] Clause 6. The method of Clause 1 further includes: receiving a selection of text data entered by a user into a text field; and rendering a writing assistant interface with a control that, when selected, causes the text data and the context data to be transferred to a generative language model.
[0109] Clause 7. The method as described in Clause 1 further includes: transmitting the text data and the context data in response to the amount of text data entered by the user into the text field reaching a threshold level; and providing a response as a suggestion in the writing assistant interface.
[0110] Clause 8. The method described in Clause 7 further includes: detecting cursor positioning on the suggestion; and providing a preview of the response in the text field.
[0111] Clause 9. The method described in Clause 1 further includes: inserting the response into a text field.
[0112] Clause 10. The method of Clause 1, wherein the digital content is a webpage, the method further comprising: retrieving first page content of the webpage; retrieving second page content of a webpage embedded in the webpage; and generating context data to include the first page content and the second page content.
[0113] Clause 11. The method of Clause 1, wherein the digital content is a webpage, the method further comprising: retrieving a document object model (DOM) representation of the webpage; extracting DOM portions from the DOM representation; and generating context data to include the DOM portions.
[0114] Clause 12. The method of Clause 1, wherein the digital content is a webpage, the method further comprising: retrieving an accessible content structure of the webpage; and generating context data to include the accessible content structure.
[0115] Clause 13. An apparatus comprising: at least one processor; and a non-transitory computer-readable medium storing executable instructions that cause the at least one processor to perform operations including: receiving from a user text data relating to input to a text field of digital content displayed on a user device; generating context data about the digital content; providing the text data and the context data to a generative language model; receiving a response generated by the generative language model; and providing the response as a suggestion for the input to the text field.
[0116] Clause 14. The device as described in Clause 13, wherein the operation further comprises: the model determining, based on signals, whether to render a call-out display, including one or more signals relating to a text field, one or more signals relating to digital content, or one or more signals relating to a writing assistant user and other users, wherein the call-out display, when selected, is configured to render a writing assistant interface for the text field, the writing assistant interface having an input field configured to receive text data from a user.
[0117] Clause 15. The device as described in Clause 13, wherein the operation further comprises: in response to the amount of text data input by a user into a text field reaching a threshold level, rendering a call-out display, which, when selected, is configured to render a writing assistant interface for the text field, the writing assistant interface having an input field with the text data.
[0118] Clause 16. The device as described in Clause 13, wherein the operation further includes: receiving a selection of a user interface object for a text field concerning digital content; and rendering a writing assistant interface for the text field, the writing assistant interface having an input field configured to receive text data from a user.
[0119] Clause 17. The device as described in Clause 13, wherein the digital content is a webpage, wherein the operation further includes: retrieving first page content of the webpage; retrieving second page content of a webpage embedded in the webpage; and generating context data to include the first page content and the second page content.
[0120] Clause 18. A non-transitory computer-readable medium storing executable instructions that cause at least one processor to perform operations including: receiving text data from a user in relation to input of a text field of digital content displayed on a user device; generating context data of the digital content; providing the text data and the context data to a generative language model; receiving a response generated by the generative language model; and providing the response as a suggestion for input to the text field.
[0121] Clause 19. A non-transitory computer-readable medium as described in Clause 18, wherein the operation further comprises: determining by the model whether to render a call-out display, which, when selected, is configured to render a writing assistant interface for text fields having input fields configured to receive text data from a user.
[0122] Clause 20. A non-transitory computer-readable medium as described in Clause 18, wherein the digital content is a web page, wherein the operation further includes: retrieving first page content of the web page; retrieving second page content of a web page embedded in the web page; and generating context data to include the first page content and the second page content.
[0123] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be dedicated or general-purpose, coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to transmit data and instructions to the storage system, at least one input device, and at least one output device.
[0124] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0125] To provide interaction with the user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor for displaying information to the user) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including sound, speech, or tactile input.
[0126] The systems and technologies described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with the implementation of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), and the Internet.
[0127] Several implementations have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the disclosed implementations.
[0128] Furthermore, the logical flow depicted in the diagram does not require the desired result to be achieved in the specific order or sequence shown. Additionally, other steps may be provided, or steps may be removed from the described flow, and other components may be added to or removed from the described system.
[0129] In some aspects, the technology described herein relates to a method comprising: receiving focus on a text box of a webpage; determining whether to display a call-out feature configured to initiate a writing assistant for the text box; and, in response to determining to display the call-out feature, providing a suggested suggestion to the user in a writing assistant interface. The suggested suggestion may be based on the context of the webpage.
[0130] In some respects, the techniques described herein relate to a method comprising: receiving from a user a cue related to input to a text box on a webpage; generating a contextual signal for the webpage; providing the cue and the contextual signal to a generative language model trained to provide output for the type of the text box; receiving a response generated by the generative language model; and providing the response as input to the text box.
[0131] In some respects, the techniques described herein relate to a non-transitory computer-readable medium that stores instructions that, when executed by a processor, perform any of the operations or methods disclosed herein.
[0132] In some aspects, the techniques described herein relate to a computing device including at least one processor and a memory storing instructions that cause the computing device to perform any of the operations or methods disclosed herein.< / product>
Claims
1. A method comprising: Receive text data from the user in relation to input in text fields of digital content displayed on the user's device; Generate contextual data about the digital content; The text data and the context data are provided to the generative language model; Receive the response generated by the generative language model; as well as The response is provided as a suggestion for the input in the text field.
2. The method of claim 1, further comprising: Detect interactions with the text field; as well as The model determines whether to render a call-out display, which, when selected, is configured to render a writing assistant interface for the text field, the writing assistant interface having an input field configured to receive the text data from the user.
3. The method of claim 2, further comprising: The decision to render the call-out display is based on signals, including one or more signals relating to the text field, one or more signals relating to the digital content, or one or more signals relating to the writing assistant, the user, and other users.
4. The method according to any one of claims 1 to 3, further comprising: In response to the amount of text data input by the user into the text field reaching a threshold level, a call-out display is rendered. When selected, the call-out display is configured to render a writing assistant interface for the text field, the writing assistant interface having an input field with the text data.
5. The method according to any one of claims 1 to 4, further comprising: Receive a selection of a user interface object for the text field concerning the digital content; Render a writing assistant interface for the text field, the writing assistant interface having an input field configured to receive the text data from the user; as well as In response to the selection of the generation control in the writing assistant interface, the text data and the context data are transmitted to the generative language model.
6. The method according to any one of claims 1 to 5, further comprising: Receive a selection of the text data input by the user into the text field; as well as Render a writing assistant interface with controls, which, when selected, cause the transfer of text data and context data to the generative language model.
7. The method of any one of claims 1 to 6, further comprising: In response to the amount of text data input by the user into the text field reaching a threshold level, the text data and the context data are transmitted. as well as Provide the aforementioned response as a suggestion in the writing assistant interface.
8. The method of claim 7, further comprising: Detect the cursor positioning on the suggested location; as well as A preview of the response is provided in the text field.
9. The method of any one of claims 1 to 8, further comprising: The response is inserted into the text field.
10. The method of any one of claims 1 to 9, wherein the digital content is a webpage, the method further comprising: Retrieve the content of the first page of the webpage; Retrieve the content of the second page of the webpage embedded in the webpage; as well as The scene data is generated to include the content of the first page and the content of the second page.
11. The method of any one of claims 1 to 10, wherein the digital content is a webpage, the method further comprising: Retrieve the Document Object Model (DOM) representation of the webpage; Extract the DOM portion from the DOM representation; as well as Generate the context data to include the DOM portion.
12. The method of any one of claims 1 to 11, wherein the digital content is a webpage, the method further comprising: Retrieve the accessible content structure of the webpage; as well as The context data is generated to include the accessible content structure.
13. An apparatus comprising: At least one processor; as well as A non-transitory computer-readable medium storing executable instructions that cause at least one processor to perform operations, the operations including: Receive text data from the user in relation to input in text fields of digital content displayed on the user's device; Generate contextual data about the digital content; The text data and the context data are provided to the generative language model; Receive the response generated by the generative language model; and The response is provided as a suggestion for the input in the text field.
14. The device as claimed in claim 13, wherein, The operation further includes: The model determines whether to render a call-out display based on signals, including one or more signals related to the text field, one or more signals related to the digital content, or one or more signals related to the writing assistant user and other users. When selected, the call-out display is configured to render a writing assistant interface for the text field, the writing assistant interface having an input field configured to receive the text data from the user.
15. The device as claimed in claim 13 or 14, wherein, The operation further includes: In response to the amount of text data input by the user into the text field reaching a threshold level, a call-out display is rendered. When selected, the call-out display is configured to render a writing assistant interface for the text field, the writing assistant interface having an input field with the text data.
16. The device as claimed in any one of claims 13 to 15, wherein, The operation further includes: Receive selection of a user interface object for the text field regarding the digital content; and Render a writing assistant interface for the text field, the writing assistant interface having an input field configured to receive the text data from the user.
17. The device as claimed in any one of claims 13 to 16, wherein, The digital content is a webpage, and the operation further includes: Retrieve the content of the first page of the webpage; Retrieve the content of the second page of the webpage embedded in the webpage; and The scene data is generated to include the content of the first page and the content of the second page.
18. A non-transitory computer-readable medium storing executable instructions that cause at least one processor to perform operations, the operations including: Receive text data from the user in relation to input in text fields of digital content displayed on the user's device; Contextual data for generating the digital content; The text data and the context data are provided to the generative language model; Receive the response generated by the generative language model; as well as The response is provided as a suggestion for the input in the text field.
19. The non-transitory computer-readable medium of claim 18, wherein the operation further comprises: The model determines whether to render a call-out display, which, when selected, is configured to render a writing assistant interface for the text field, the writing assistant interface having an input field configured to receive the text data from the user.
20. The non-transitory computer-readable medium as claimed in claim 18 or 19, wherein, The digital content is a webpage, and the operation further includes: Retrieve the content of the first page of the webpage; Retrieve the content of the second page of the webpage embedded in the webpage; and The scene data is generated to include the content of the first page and the content of the second page.