Input mechanism for generative models
Patent Information
- Application Number
- CN202580014459.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-14
- Filing Date
- 2025-02-12
- Publication Date
- 2026-09-11
AI Technical Summary
因此,鉴于包括在LLM中的大量参数,使用LLM来执行NLP任务可能消耗相对大量的资源(例如,就完成NLP任务时所使用的计算资源、完成NLP任务的执行所花费的时间、完成NLP任务的执行所消耗的能量等而言)
[0005]For example, following the example above, the techniques described in this paper can be used to enrich the initial query to read: "Help meto configure my smart home ecosystem. An unregistered smart light device is located in the master bedroom. Please detect and register this device with the smart home ecosystem, and add it to the device group named 'masterbedroom'. Please name the new device in accordance with the naming scheme of the 'master bedroom' device group and apply all routines associated with the other smart light devices in the 'master bedroom' device group. Show me any other configuration settings available for the new device." Therefore, the response generated based on processing this enriched input can be expected to be highly relevant to the user's intended outcome from the interaction (e.g., and thus minimize subsequent computationally expensive follow-up interactions).
Smart Images

Figure CN122743501A_ABST
Abstract
Description
Background Technology
[0001] Various generative models have been proposed that can be used to process natural language (NL) content and / or other inputs to generate outputs that reflect generative content in response to the input. For example, large language models (LLMs) are a specific type of machine learning model capable of performing various natural language processing (NLP) tasks such as language generation, machine translation, and question answering. These LLMs are typically trained on massive amounts of diverse data, including, but not limited to, data from web pages, ebooks, software code, e-news articles, and machine-translated data. Therefore, these LLMs leverage the underlying data on which they are trained to perform these various NLP tasks. For example, when performing a language generation task, these LLMs can process natural language (NL)-based input received from a client device and generate NL-based outputs that are in response to the NL-based inputs and are to be rendered at the client device.
[0002] In some cases, an LLM can include millions, hundreds of millions, billions, or even a hundred billion or more parameters. Therefore, given the large number of parameters included in an LLM, using an LLM to perform NLP tasks can consume relatively large amounts of resources (e.g., in terms of the computational resources used to complete the NLP task, the time spent executing the NLP task, the energy consumed in executing the NLP task, etc.). Therefore, in terms of computational resource usage, it is advantageous that an LLM generates responses to NL-based inputs that do not require additional follow-up. Summary of the Invention
[0003] The implementation disclosed herein involves enriching (or otherwise termed refining) the input provided by the user to the LLM. Because the input has been enriched, the response to that input (e.g., generated using the LLM) can be more relevant to the user's intended outcome from the interaction. For example, an initial query entered by the user may be flawed in one or more ways. For instance, the initial query may contain multiple ambiguities, or may lack details necessary to complete the user's intended task, or may otherwise lack other crucial information. Therefore, a response generated using such an initial query may be unsatisfactory (e.g., leading to subsequent computationally expensive follow-up interactions to complete the user's intended task). The implementation described herein enables (e.g., leveraging the LLM) enrichment of the initial query to correct such flaws. Thus, a response generated based on the enriched input can be expected to be highly relevant to the user's intended outcome from the interaction (e.g., and therefore minimize subsequent computationally expensive follow-up interactions).
[0004] As an example, an initial query might be provided by a user: "Add a bedroom light." This initial query contains multiple ambiguities and lacks key information and any details, and doesn't even clearly specify the expected task to be completed. For example, it's unclear whether the user wants to configure a smart device (e.g., a smart light fixture) or add a light to a shopping list or cart, or which bedroom the user refers to. Therefore, it's unlikely to achieve the user's intended outcome based solely on this initial input. However, according to the implementation described herein, the initial query can be enriched to address these shortcomings. As discussed herein, in some implementations, this enrichment can be based on contextual information, such as data associated with the user's profile (e.g., the existence of a smart home ecosystem associated with the user, including a group of devices called "Master Bedroom"), historical conversation data (e.g., if the user mentioned in previous input that they had already added a new smart light fixture to their Master Bedroom), etc.
[0005] For example, following the example above, the techniques described in this paper can be used to enrich the initial query to read: "Help meto configure my smart home ecosystem. An unregistered smart light device is located in the master bedroom. Please detect and register this device with the smart home ecosystem, and add it to the device group named 'masterbedroom'. Please name the new device in accordance with the naming scheme of the 'master bedroom' device group and apply all routines associated with the other smart light devices in the 'master bedroom' device group. Show me any other configuration settings available for the new device." Therefore, the response generated based on processing this enriched input can be expected to be highly relevant to the user's intended outcome from the interaction (e.g., and thus minimize subsequent computationally expensive follow-up interactions).
[0006] In these and other ways, LLM can be used to quickly and efficiently achieve a user's intended outcome, where the user only needs to provide initial input, which may be flawed in one or more ways. This allows users to provide relatively simple input in terms of the number of characters to achieve their intended outcome. Therefore, for a given intended outcome, the amount of input required by the user (e.g., keystrokes on a keyboard) and the associated computational resources required to facilitate that input can be reduced.
[0007] In other words, the technology described herein can be said to provide a mechanism that enables user input. For example, the technology described herein can be said to help users enter text by providing a predictive input mechanism (e.g., enriched or refined input can be predicted based on the initial input).
[0008] Furthermore, determining the content of input to an LLM to achieve a specific result may require trial and error, and / or may require a high level of skill, training, and / or familiarity with LLM (or LLM in general). The techniques described in this paper mitigate these obstacles by enriching (or refining) the initial input from the user (who may not be an expert). Moreover, by presenting the user with enriched input and allowing for further editing, the user retains control and understanding of the information provided to the LLM in generating the response content, while being guided throughout the interaction and also educated on how to interact effectively with an LLM. In other words, the techniques described in this paper can be said to assist users in performing technical tasks through a guided human-computer interaction process.
[0009] The implementation described in this paper can also be used to reduce the number of follow-up inputs that can be received by an LLM. Since the inputs processed by the LLM to generate response content have been refined (or enriched), it can be assumed that the response content is highly relevant to the user's intended outcome (at least relative to the initial input provided by the user). Therefore, the generated response content is unlikely to prompt subsequent follow-up inputs (e.g., to correct or clarify aspects of the response content), at least relative to response content generated based on the initial input provided by the user. While any given user can decide to provide follow-up NL-based inputs, any "average" reduction in the number of follow-up NL-based inputs can be extremely beneficial in terms of computational resource usage.
[0010] Furthermore, in some cases, the techniques described herein can be used as part of a NL-based response system for engaging in dialogue with human users (e.g., involving multiple inputs and responses). For example, NL-based response systems can be provided as part of automated assistants, chatbots, etc. In some cases, users can provide one or more commands to be performed as part of the dialogue (e.g., to control smart devices, generate code, generate commands to control robots, assist with navigation in vehicles, etc.). Therefore, the use of the techniques described herein can also assist users in performing these technical tasks through a continuous and guided human-computer interaction process. Moreover, since responses generated using LLM can reliably have higher quality, the human-computer interaction process can end quickly and efficiently.
[0011] Some implementations described in this paper involve generating training sets and training or fine-tuning the LLM to refine (or enrich) the input, as described in this paper.
[0012] The above is presented as an overview of only some of the implementations disclosed in this article. This article discloses these and other implementations in more detail. Attached Figure Description
[0013] Figure 1 A block diagram depicts an example environment that illustrates various aspects of this disclosure and in which some of the implementations disclosed herein can be implemented.
[0014] Figure 2 An overview of various example methods, including using LLM to refine natural language-based input, is provided.
[0015] Figure 3A , Figure 3B , Figure 3C , Figure 4A , Figure 4B , Figure 5A , Figure 5B and Figure 6 A sample client device is depicted that renders a sample graphical user interface.
[0016] Figure 7A A flowchart illustrating an example method for determining the content of a response to an NL-based input is shown.
[0017] Figure 7B A flowchart illustrating an example method for fine-tuning an LLM is shown.
[0018] Figure 8 Example architectures of computing devices based on various implementations are described. Detailed Implementation
[0019] Go to Figure 1 This diagram depicts a block diagram of an example environment 100 illustrating various aspects of this disclosure and in which implementations disclosed herein may be implemented. Example environment 100 includes a client device 110 and a natural language (NL)-based response system 120. In some implementations, all or aspects of the NL-based response system 120 may be implemented locally at the client device 110. In additional or alternative implementations, all or aspects of the NL-based response system 120 may be implemented from, for example... Figure 1The depicted client device 110 is implemented remotely (e.g., at a remote server). In those implementations, the client device 110 and the NL-based response system 120 may be communicatively coupled to each other via one or more networks 199, such as one or more wired or wireless local area networks (“LANs”, including Wi-Fi LANs, mesh networks, Bluetooth, near field communication, etc.) or wide area networks (“WANs”, including the Internet)).
[0020] Client device 110 may be one or more of the following, for example: desktop computer, laptop computer, tablet computer, mobile phone, vehicle computing device (e.g., in-vehicle communication system, in-vehicle entertainment system, in-vehicle navigation system), independent interactive speaker (optionally with a display), smart home appliance (such as a smart TV), and / or user's wearable device including a computing device (e.g., user's watch with computing device, user's glasses with computing device, virtual or augmented reality computing device). Additional and / or alternative client devices may be provided.
[0021] Client device 110 may execute one or more software applications via application engine 115, through which NL-based input can be submitted, and / or NL-based output and / or other output can be rendered in response to the NL-based input (e.g., audibly and / or visually). Application engine 115 may execute one or more software applications separate from the operating system of client device 110 (e.g., software applications installed "on top" of the operating system), or alternatively, may be implemented directly by the operating system of client device 110. For example, application engine 115 may execute a web browser or automation assistant installed on top of the operating system of client device 110. As another example, application engine 115 may execute a web browser software application or automation assistant software application integrated as part of the operating system of client device 110. Application engine 115 (and one or more software applications executed by application engine 115) may interact with NL-based response system 120.
[0022] In various implementations, client device 110 may include a user input engine 111 configured to detect user input provided by a user of client device 110 using one or more user interface input devices. For example, client device 110 may be equipped with one or more microphones that capture audio data, such as audio data corresponding to the user's spoken words or other sounds in the environment of client device 110. Alternatively, client device 110 may be equipped with one or more visual components configured to capture visual data corresponding to images and / or movements (e.g., gestures) detected in the field of view of one or more visual components. Alternatively, client device 110 may be equipped with one or more touch-sensitive components (e.g., keyboard and mouse, stylus, touchscreen, touch panel, one or more hardware buttons, etc.) configured to capture signals corresponding to touch input directed at client device 110. Some examples of queries described herein may be queries formulated based on user input provided by a user of client device 110 and detected via user input engine 111. For example, a query can be a typed query entered via a physical or virtual keyboard, a suggested query selected via a touchscreen or mouse, a spoken voice query detected by the microphone of the client device, or an image query based on an image captured by the visual component of the client device.
[0023] In various implementations, client device 110 may include rendering engine 112 configured to provide content (e.g., NL-based summaries) for audible and / or visual presentation to a user of client device 110 using one or more user interface output devices. For example, client device 110 may be equipped with one or more speakers enabling the provision of content for audible presentation to a user via client device 110. Alternatively, client device 110 may be equipped with a display or projector enabling the provision of content for visual presentation to a user via client device 110.
[0024] In various implementations, client device 110 may include a context engine 113 configured to determine the context (e.g., current or recent context) of client device 110 and / or its user. In some of these implementations, the context engine 113 may determine the context using profile data of the client device 110's current or recent interactions, the client device 110's location, the user of the client device 110 (e.g., an active user when multiple profiles are associated with the client device 110), and / or other data accessible to the context engine 113. For example, the context engine 113 may determine the current context based on the current state of a query session (e.g., considering one or more recent queries in the query session), profile data, and / or the current location of the client device 110. For example, the context engine 113 may determine the current context "looking for a healthy lunchrestaurant in Louisville, Kentucky" based on a recently issued query, profile data, and the client device 110's location. As another example, the context engine 113 may determine the current context based on which application is active in the foreground of the client device 110, the current or recent state of the active application, and / or the content currently or recently rendered by the active application. The context determined by the context engine 113 may be utilized, for example, in supplementing or rewriting queries formulated based on user input, in generating implicit queries (e.g., queries formulated independently of user input), and / or in determining the submission of implicit queries and / or rendering results for implicit queries (e.g., NL-based summaries).
[0025] In various implementations, client device 110 may include implicit input engine 114, configured to: generate an implicit query independently of any user input intended to formulate the implicit query; submit the implicit query, optionally independently of any user input requesting submission of the implicit query; and / or cause the rendering of the result of the implicit query, optionally independently of any user input requesting rendering of the result. For example, implicit input engine 114 may use the current context from context engine 113 to generate the implicit query, determine to submit the implicit query, and / or determine to cause the rendering of the result of the implicit query. For example, implicit input engine 114 may automatically generate and automatically submit the implicit query based on the current context. Therefore, implicit input engine 114 may automatically push the result of the implicit query so that the result is automatically rendered, or may automatically push notifications of the result, such as optional notifications, which cause the result to be rendered when selected. As another example, the implicit input engine 114 can generate implicit queries (e.g., implicit queries related to the user's interests) based on profile data, submit the queries at regular or irregular intervals, and cause corresponding results (or notifications) to be automatically provided for the submissions. For example, the implicit query could be "patent news" based on profile data indicating interest in patents, which is submitted periodically, and the corresponding NL-based summary results are automatically rendered. Note that, given the existence of, for example, new / fresh search result documents over time, the provided NL-based summary results may change over time.
[0026] Furthermore, client device 110 and / or NL-based response system 120 may include one or more memories for storing data and / or software applications, one or more processors for accessing data and executing software applications, and / or other components facilitating communication via one or more networks in network 199. In some implementations, one or more software applications may be locally installed at client device 110, while in other implementations, one or more software applications may be remotely hosted (e.g., by one or more servers) and may be accessible from client device 110 via one or more networks in network 199.
[0027] although Figure 1The aspects described herein are illustrative or representative of a single client device with a single user, but it should be understood that this is for illustrative purposes and is not intended to be limiting. For example, one or more additional client devices of the user and / or additional users may also implement the techniques described herein. For example, client device 110, one or more additional client devices, and / or any other computing devices of the user may form a device ecosystem that can employ the techniques described herein. These additional client devices and / or computing devices may communicate with client device 110 (e.g., via network 199). As another example, a given client device may be utilized by multiple users in a shared setup (e.g., user group, family).
[0028] The NL-based response system 120 is exemplified as including an LLM selection engine 124, an LLM input engine 126, an LLM response generation engine 128, a cue refinement engine 130, and a training engine 132. Some of these engines may be omitted in various implementations. In some implementations, the engines of the NL-based response system are distributed across one or more computing systems.
[0029] The LLM selection engine 124 can determine, in response to a received query, which of a plurality of generative models (LLM 150 and / or other generative models, if any) will be used to generate a response for rendering in response to the query. For example, the LLM selection engine 124 can select zero, one or more generative models to be used to generate a response for rendering in response to the query. The LLM selection engine 124 may optionally utilize one or more classifiers and / or rules (not illustrated).
[0030] LLM input engine 126 can generate LLM input in response to a received query. This LLM input will be processed using LLM to generate an NL-based response to the query. As described herein, such content may include query-based content and / or additional content, such as contextual information derived from user data and / or historical conversation data.
[0031] LLM response generation engine 128 can use LLM to process LLM inputs generated by LLM input engine 126 to generate NL-based responses. In various implementations, LLM response generation engine 128 can perform... Figure 7A Method 700A includes all or all aspects of boxes 720 and 740. The LLM response generation engine 128 can utilize one or more LLM 150s.
[0032] The prompt refinement engine 130 can generate refined input based on NL-based input prompts. For example, the prompt refinement engine 130 can use an LLM response generation engine 128 (e.g., utilizing one or more LLMs 150) to process the LLM input generated by the LLM input engine based on NL-based input to generate refined input prompts. In various implementations, the prompt refinement engine 130 can perform... Figure 7A Method 700A of all or all aspects of frame 720.
[0033] Training engine 132 can train one or more LLM 150s. For example, training engine 132 can use training data from the training dataset (e.g., training data 152) to retrain / fine-tune the parameters of one or more LLM 150s. In various implementations, training engine 132 can perform... Figure 7B Method 700B, including all or all aspects of boxes 760 and 770.
[0034] Now go to Figure 2 It provides an overview of various example methods, including using LLM to refine natural language-based inputs.
[0035] like Figure 2 As shown, NL-based input 210 can be received (e.g., from a user of client device 110). In some examples, NL-based input 210 is received in the form of an input text query. For example, NL-based input 210 can be manually initiated as text input by a user of a user application running at client device 110. Alternatively or additionally, NL-based input 210 can originate from verbal input to a user application running at client device 110, such as verbal query input following invocation of the user application. Verbal input is converted into an input query by a speech-to-text engine running on client device 110 (as part of the user application, or accessible by the user application). In some examples, NL-based input 210 is part of an ongoing human-computer dialogue, such as a series of input queries and their corresponding responses from NL-based response system 120. Although NL-based input 210 is generally described as originating from input text (or verbal input), it should be understood that NL-based input can originate from any suitable format (e.g., an image).
[0036] The NL-based input 210 can be processed to generate a corresponding refined input hint 220 (e.g., using the hint refinement engine 130). For example, the hint refinement engine 130 can use the LLM input engine 126 to generate LLM input based on the NL-based input 210. Then, the hint refinement engine 130 can use the LLM response generation engine 128 to generate the refined input hint 220 based on the LLM input (e.g., which may include at least the NL-based input 210). For example, the LLM response generation engine 128 can use (e.g., in LLM 150) an LLM to process the LLM input to generate an LLM output corresponding to the refined input hint 220.
[0037] LLM input can be generated to include requests to refine a given NL-based input. In some implementations, the LLM input can be customized based on input prompts and / or other contextual information such as the task to be completed (e.g., whether contextual data 214 is considered, whether the user is instructed to select terms, etc.). In some implementations, given a corresponding NL-based input, the LLM input can include one or more examples of corresponding refined input prompts. For example, these examples can be selected or compiled by a human expert. In some implementations, the LLM input can be formatted according to a template. The template may not be provided by the user providing the initial NL-based input 210, and can even be provided to the LLM without user knowledge. The template may include, for example, spaces or input fields following the request, to prompt the LLM to fill the spaces or input fields with output in response to the request.
[0038] The refined input prompts 220 can be processed (e.g., by the NL-based response system 120) to generate response content 230. For example, the LLM input engine 126 can generate LLM input based on the refined input prompts 220. The LLM response generation engine 128 can then be used to generate response content 230 based on processing the LLM input (which may include at least the refined input prompts). For example, the LLM response generation engine 128 can use an LLM (e.g., the same LLM used to generate the refined input prompts 220) and / or another LLM (e.g., from LLM 150) to process the LLM input to generate LLM output corresponding to the response content. The response content 230 may include a natural language response to the refined input prompts 220 (and correspondingly, the NL-based input 210). The response content 230 can then be rendered at the client device 110 (e.g., via a user application), for example, as text in a text-based conversation / chat application, converted to speech using a text-to-speech engine, etc.
[0039] about Figures 3A to 3C Described Figure 2 The example shown is an example of the example method. For example, and go to... Figure 3A The document depicts an example client device 310 (which may be similar to client device 110) rendering an example graphical user interface 350. The graphical user interface 350 may include an input field 320 for receiving NL-based input from a user. For example, the input field 320 may receive text data from a virtual or physical keyboard. In some implementations, the graphical user interface may include a graphical element 322 (represented in this case by a microphone icon) that, upon selection, can begin recording (e.g., via one or more microphones of client device 310) spoken words from the user. The spoken words can then be transcribed and entered into the input field 320. In some implementations, the graphical user interface may include a graphical element 324 (in this case, an image icon) that, upon selection, allows the user to submit an image for processing (e.g., an image retrieved from storage on client device 310, an image retrieved from the Internet, etc.). In some implementations, a model may be used to generate NL-based input based on the submitted image to be entered into the input field (e.g., based on text detected in the image, people and / or objects detected in the image, etc.). The graphical user interface 350 may also include a graphical element 326 that, when selected, can submit NL-based input currently displayed in the input field 320 (or in other words, selection of the graphical element 326 can indicate acceptance of NL-based input currently displayed in the input field 320).
[0040] In some implementations, the graphical user interface 350 may include a graphical element 330. The graphical element 330 may be configured to refine (e.g., as described herein) the NL-based input (e.g., text currently displayed in input field 320) upon selection. In some implementations, the graphical element 330 is provided (e.g., rendered) in the graphical user interface 350 solely based on a determination that it should be provided (e.g., rendered). For example, the determination that the graphical element 330 should be provided may be based on a quality metric of the NL-based input. A quality metric may include, for example, the length of the NL-based input (e.g., currently displayed in input field 320). For example, if the NL-based input includes a number of characters below a minimum threshold and / or above a threshold number of characters, it may be determined that the graphical element 330 is not provided. As another example, a quality metric may include, for example, an indication that the NL-based input cannot be meaningfully refined. For example, it can be determined that the NL-based input cannot be meaningfully refined based on the processing of the input by an LLM (or another model) before the selection of graphical element 330, or in response to a previous selection of graphical element 330. For example, previous iterations of the generation of refined input suggestions (which may have already resulted in the NL-based input currently included in input field 320) can provide metadata indicating that the NL-based input cannot be meaningfully refined further, or that the maximum threshold number of refinement iterations may have been reached. In some implementations, it can be determined not to provide graphical element 330 based on the determination that the text included in input field 320 contains harmful content. For example, when text is entered into input field 320, the text included in input field 320 can be processed by a model (e.g., trained to identify harmful content) to generate an indication of whether the text contains harmful content. In some implementations, harmful content can be based on the detection of the presence of keywords. In some implementations, when refining NL-based input, harmful content detection can be performed as an intermediate step (e.g., in response to a user selecting a graphical element that causes the NL-based input to be refined, but before generating a refined input hint). If harmful content is detected, subsequent processing can be blocked and / or a warning message indicating that harmful content has been identified can be rendered to the user.
[0041] like Figure 3A As depicted, a user can enter the NL-based input "Find me flights to Miami" into input field 320. Assuming the user then selects graphical element 330, the NL-based input can be refined. Therefore, the graphical user interface 350 can be updated to replace the NL-based input rendered in input field 320 with refined input prompts (e.g., ...). Figure 3B (As depicted). While it is generally described that NL-based input can be refined based on the selection of graphical elements (e.g., graphical element 330 in this case), it should be understood that this is not necessary. For example, NL-based input can be automatically refined in response to interaction with a physical input device (e.g., selection of a key on a physical keyboard), in response to spoken words, etc. (e.g., when it is determined that the user has stopped entering text for a predetermined period of time).
[0042] like Figure 3B The described field 320 now includes the text “Find me flights from NYC to Miami for this summer. Give me an itinerary that includes breakfast, lunch, and dinner plans. Recommend places that are well-rated and no more than 30 minutes away from each other. Give me the suggestions in bullet points.” This can be further refined (e.g., Figure 3A The text is generated from NL-based input previously included in input field 320, as described herein. In some implementations, the graphical user interface 350 may also be updated to include a graphical element 328, which, upon selection, causes the text currently included in input field 320 to revert to the text included in the input field before refinement (e.g., an "Undo" button). Furthermore, the text included in the input field may be editable (e.g., via user selection and input via a virtual or physical keyboard, via one or more spoken words, etc.). Note that the user can further refine the text now included in input field 320 (e.g., by selecting graphical element 330). However, assuming the user indicates acceptance of the text included in input field 320 (e.g., based on the selection of graphical element 326), the text can be submitted for processing by the NL-based response system 120 (or its LLM).
[0043] For example, such as Figure 3C The graphical user interface 350 described has (for example, based on) Figure 3BThe selection of graphic element 326 in the text is updated to indicate the text 360 submitted to the NL-based response system 120 (which was previously included in the text). Figure 3B (In the input field 320). Alternatively, the graphical user interface 350 may be updated to include content 362 in response to submitted text 360 (and correspondingly initial NL-based input). For example, response content 362 may include the text “Okay. Here are the available flights:...”, as well as various available flight options based on the user’s detailed instructions, itinerary plans based on the user’s detailed instructions, etc.
[0044] Return now Figure 2 In some implementations, the refined input hints 220 can also be generated based on context data 214. For example, context data 214 can be obtained (e.g., by hint refinement engine 130) and used (e.g., by LLM input engine 126) to generate LLM input. LLM can be used (e.g., by LLM response generation engine 128) to process (at least including NL-based input 120 and context data 214) LLM input to generate refined input hints 220.
[0045] Context data can be profile data based on, for example, current or recent interactions via client device 110, the location of client device 110, profile data of the user of client device 110 (e.g., an active user when multiple profiles are associated with client device 110), and / or other data accessible to the NL-based response system 120 (e.g., via context engine 113). For example, assuming the user has authorized the NL-based response system 120 to have access to such data, context data could include the user's home and / or work address, calendar entries associated with the user, names of family members and their relationships to the user, smart home configuration data associated with the user, user preference information, etc. In some implementations, context data could include historical data associated with one or more previous rounds of conversation with the user. For example, Figure 4A and 4B An illustrative example of input refinement based on contextual data obtained from multi-turn dialogue sessions is described.
[0046] Go to Figure 4A The illustration depicts an example client device 410 (which may be similar to client device 110 and client device 310) rendering an example graphical user interface 450. The graphical user interface 450 can be largely compared to... Figure 3A The graphical user interface is similar to that of the 350, and includes largely the same... Figure 3AThe input field 320 is similar to the input field 420, and is largely the same as... Figure 3A The graphic elements 322, 324, 326, and 330 are similar to the graphic elements 422, 424, 426, and 430.
[0047] like Figure 4A As depicted, the user and the NL-based response system 120 have engaged in a multi-turn dialogue. This multi-turn dialogue is represented by: user submission 460 including the text "What is a good destination for a summer vacation?", response 462 including the text "London is a good spot for a summer vacation", user submission 464 including the text "I will be flying from NYC and would like a short flight", and response 466 including the text "Based on that, I recommend Miami". Also, Figure 4A What is described, and Figure 3A Similarly, the user has entered "Find me flights to Miami" as an NL-based input.
[0048] Go to Figure 4B Suppose the user has requested refinement of the NL-based input (e.g., with...). Figure 4B Similarly, the NL-based input "Find me flights to Miami" could be replaced with the text "Find me flights from NYC to Miami for this summer. Give me an itinerary that includes breakfast, lunch, and dinner plans. Recommend places that are well-rated and no more than 30 minsaway from each other. Give me the suggestions in bullet points." The graphical user interface 450 can be updated accordingly.
[0049] In this scenario, it can be assumed that the NL-based input is refined at least in part based on contextual data, which includes historical data associated with one or more previous rounds of the dialogue between the user and the NL-based input system 120 (e.g., specific information specified by the user and / or provided by the NL-based response system in previous rounds of the dialogue). For example, the contextual data may include an indication that the user is departing from New York (e.g., based on input 464), and the refined input may accordingly include a request for a flight departing from New York. Alternatively, the contextual data may include an indication that the user is looking for summer flights (e.g., based on input 460), and the refined input may accordingly include a request for a summer flight. Alternatively, the contextual data may include an indication that the user is on vacation (e.g., based on input 460), and the refined input may accordingly include a request for itinerary planning for various dining arrangements. Alternatively, contextual data may also include indications that users typically do not travel more than 30 minutes for a meal and / or that they prefer to be presented as key points (e.g., based on user-associated preference data, recorded user behavior history, etc.). Refined input may accordingly include these constraints. In this way, users do not need to repeatedly submit information to the NL-based response system 120.
[0050] Return now Figure 2 In some implementations, one or more terms (e.g., characters, words, phrases, sentences, etc.) from the refined input prompt 220 can be selected as user-selectable terms. Indications for the user-selectable terms can then be rendered. During user interaction with a specific user-selectable term (e.g., selection, cursor hovering, etc.), one or more alternative terms for that specific user-selectable term can be rendered. These alternative terms can be selectable, allowing the selected alternative term to replace the corresponding user-selectable term upon selection.
[0051] The selection of user-selectable terms can be based on metadata indicating user-selectable terms generated, for example (e.g., by an LLM), using refined input prompts 220. For example, the LLM can be fine-tuned to identify user-selectable terms in a given refined input prompt, and / or the LLM input processed by the LLM to generate the refined input may include a request to identify user-selectable terms. Alternatively, user-selectable terms can be selected after the generation of refined input prompts 220 (e.g., separately from the generation of refined input prompts 220 by another model). In some implementations, user-selectable terms can be selected based on a confidence metric associated with the term. For example, if the confidence metric associated with a given term is below a confidence threshold, it can be determined that the given term is selected as a user-selectable term (e.g., because the user is more likely to want to change the term). As another example, user-selectable terms can be selected based on the identification that the term in question has at least one possible alternative. For example, based on the existence of alternative terms that can be presented upon user request, and the possibility that one of the alternative terms is correct (e.g., because there is a finite and discrete number of alternatives, because the user has previously selected a particular alternative, etc.), the term "Monday" can be selected as a user-selectable term. As another example, user-selectable terms can be selected based on contextual data. For instance, based on (e.g., based on user data) determining that the user frequently travels with people associated with the title "daughter" and occasionally with people associated with the title "son," the term "daughter" can be selected as a user-selectable term. In this case, the alternative terms provided to the user could include, for example, "daughter and son," "son," "alone," etc.
[0052] The determination of alternative terms can be performed in any suitable manner. For example, similar to the selection of user-selectable terms, alternative terms can be indicated in the metadata generated with the refined input prompt 220. Alternatively, alternative terms can be determined after the generation of the refined input prompt 220 (e.g., before or during interaction with the corresponding user-selectable term). For example, alternative terms can be determined using another model (e.g., based on processing the corresponding user-selectable term). In some implementations, context data can be used to determine alternative terms. For example, following the example above of selecting “daughter” as a user-selectable term, alternative terms provided to the user could include “daughter and son,” “son,” “alone,” etc., based on user data indicating that the user frequently travels with people associated with the title “daughter” and occasionally with people associated with the title “son.” In some implementations, user selections of specific alternatives can be stored for future refinement of NL-based input. For example, following the example above again, suppose the user selects the alternative term "daughter and son". The indication of this selection can be stored and provided when a refined input suggestion is subsequently generated, such that the refined input suggestion includes "daughter and son" instead of "daughter" (which may have already been included without this indication).
[0053] For example, Figure 5A and Figure 5B Examples of user-selectable terms and corresponding alternative terms rendered to the user are shown. Go to Figure 5A The illustration depicts an example client device 510 (which may be similar to client device 110 and client device 310) rendering an example graphical user interface 550. The graphical user interface 550 can be largely compared to... Figure 3B The graphical user interface is similar to that of the 350, and includes largely the same... Figure 3B The entry field 320 is similar to the entry field 520, and is largely the same as... Figure 3B The graphic elements 322, 324, 326 and 330 are similar to the graphic elements 522, 524, 526 and 530.
[0054] like Figure 5AThe described input field 520 includes refined input prompts such as "Find me flights from NYC to Miami for this summer. Give me an itinerary that includes breakfast, lunch, and dinner plans. Recommend places that are well-rated and no more than 30 minutes away from each other. Give me the suggestions in bullet points." It also includes... Figure 5A The depicted graphical user interface 550 also includes various user-selectable instructions 564, 566, 568, and 570. Figure 5A In this context, the system indicates to the user selectable terms based on underlined terms (i.e., “NYC”, “Miami”, “summer”, “30mins”). However, it should be understood that any suitable method can be used to indicate selectable terms to the user.
[0055] Now go to Figure 5B The graphical user interface 550 has been updated to include various alternative terms 568A, 568B, and 568C for the corresponding user-selectable term 568. Specifically, alternative terms "Winter" 568A, "Spring" 568B, and "Fall" 568C can be presented to the user as alternatives to the user-selectable term "summer" 568. The graphical user interface 550 can be updated in this way based on detected user interaction with the user-selectable term "summer" 568 (e.g., user selection of that user-selectable term). When one of the alternative terms 568A, 568B, and 568C is selected, the corresponding user-selectable term 568 can be replaced with the selected alternative term 568A, 568B, and 568C. The graphical user interface 550 can be updated accordingly to include one of the alternative terms 568A, 568B, and 568C in the input field 520 to replace the user-selectable term 568.
[0056] Return now Figure 2In some implementations, the NL-based input 210 may be stored together with the refined input cues 220 as training examples (e.g., in the training data database 152). Furthermore, in some implementations, additional signaling information (e.g., from one or more user inputs 212 as described herein) may be stored in the corresponding training examples.
[0057] For example, in some implementations, multiple options for response content 230 can be generated and rendered to the user. For instance, since LLMs can operate probabilistically, multiple refined input cues 220 can be generated (e.g., using a cue refinement engine 130 as described herein), and these multiple refined input cues can be expected to differ from each other, even when generated using the same input (e.g., based on the same NL-based input 210). Therefore, each of the multiple options for response content 230 can be generated based on a corresponding one of the multiple refined input cues. Alternatively, the multiple options for response content 230 can include response content generated based on (e.g., by LLM) directly processing the NL-based input 210. For example, a first option for response content 230 (as described herein) can be generated based on refined input cues 220, and a second option for response content can be generated based on directly processing the NL-based input 210. The multiple options for response content 230 can then be rendered at the client device 110. Each of the multiple options in response content 230 can be selectable (e.g., for selecting which option best achieves the user's intended outcome, for determining which option to continue with in a multi-turn dialogue, etc.). Indications of the selected option can be stored in corresponding training examples (e.g., as user input 212 in training data database 152). For example, the indication of the selected option can be stored as a training example along with NL-based input 210 and a corresponding refined input cue 220 (e.g., a refined input cue for generating response content associated with the selected option). This indication can be viewed as an additional signal that the corresponding refined input cue 220 is an appropriate refined input cue given NL-based input 210 (and optionally available context data 214) (e.g., because the user explicitly indicates that the obtained response content achieves their intended outcome at least relative to the other options offered).
[0058] Although only one example of user input 212 is described here, it should be understood that any suitable user input 212 can be used for these purposes, as described in more detail herein. For example, multiple options of refined input prompts can be provided to the user, and user input can be provided to select one (or more) of the refined input prompts. Other examples may include, for example: user input for editing refined input prompts, user input for reverting from refined input prompts to NL-based input, user input for providing subsequent NL-based input (e.g., to clarify or correct aspects of the response content), etc.
[0059] exist Figure 6 The text describes an example of rendering multiple options to the user. Go to... Figure 6 The illustration depicts an example client device 610 (which may be similar to client device 110 and client device 310) rendering an example graphical user interface 650. The graphical user interface 650 can be largely compared to... Figure 3C The graphical user interface is similar to that of the 350, and includes largely the same... Figure 3C The input field 620 is similar to the input field 620, and is largely the same as... Figure 3C The graphic elements 322, 324, 326 and 330 are similar to the graphic elements 622, 624, 626 and 630.
[0060] like Figure 6 The user may have submitted the text 660, “Find me flights from NYC to Miami for this summer. Give me an itinerary that includes breakfast, lunch, and dinner plans. Recommend places that are well-rated and no more than 30 minsaway from each other. Give me the suggestions in bullet points.” to an NL-based response system 120 (e.g., as...). Figures 3A to 3C (As described). Based on the user-submitted text 660, a first option 662A and a second option 662B of the response content can be provided in the graphical user interface 650. The first option 662A of the response content can be generated based on processing the submitted text 660 (e.g., it can be refined based on the initial NL-based input, e.g., as about). Figures 3A to 3C(As described). In some implementations, the second option 662B of the response content may be generated based on processing the NL-based input used to generate the submitted text 660 (e.g., by refining the NL-based input). In some implementations, the second option 662B of the response content may be generated based on processing an alternative refined input prompt that is generated based on processing the initial NL-based input. For example, since LLM is probabilistic, it can be expected that when multiple refined input prompts are generated based on the same NL-based input, they will be different from each other.
[0061] Various options 662A and 662B of the response content can be presented as user-selectable elements. These user-selectable elements can be configured such that selecting an option instructs the user to choose the option that better achieves their intended outcome. For example, a selected option can be selected to continue in a multi-turn dialogue (e.g., the graphical user interface 650 can be updated to retain the selected option and remove options that have not yet been selected). As another example, the response content may include one or more actions that can be performed by the NL-based response system 120, and the selection of an option can cause the NL-based response system 120 to execute the corresponding action. In response to the selection of one of the options 662A and 662B, the indication of the selection can be stored as training data, as described herein.
[0062] Return now Figure 2 In some implementations, training data obtained from a training data database (e.g., training data database 152) can be used (e.g., using training engine 132) to train one or more LLMs and / or one or more additional LLMs. For example, an LLM can be fine-tuned to generate refined input cues based on a given NL-based input. In some examples, the training data can be used to fine-tune an existing LLM, such as one of LLMs 150. Alternatively or additionally, the training data can be used to train a new LLM from scratch. The training data can be used by training engine 132 to determine a parameter update set for one or more of the parameters of LLMs 150 or additional LLMs.
[0063] Training data may include one or more training examples. Each training example may include a NL-based input and a corresponding refined input cue. Training examples can be used during LLM fine-tuning. Typically, it can be assumed that the training examples include refined input cues that accurately reflect the intended input based on the corresponding NL-based input. Therefore, the LLM can be fine-tuned towards generating similar refined input cues given the same NL-based input (e.g., by providing a reward during fine-tuning when the LLM generates semantically similar refined input cues given the same NL-based input). However, in some implementations, some or all of the training examples may include cues that do not accurately reflect the intended input based on the corresponding NL-based input. In these cases, the LLM can be fine-tuned away from generating similar refined input cues given the same NL-based input of these training examples (e.g., by providing a negative reward or penalty during fine-tuning when the LLM generates semantically similar refined input cues given the same NL-based input).
[0064] In some implementations, training examples may include additional signals that can be used during LLM fine-tuning (e.g., based on user input 212 received at client device 110). As an example, the additional signals may include an indication of whether the user manually edited the refined input prompt. For example, if the user manually edited the refined input prompt, this could be an indication that the refined input prompt accurately reflects the user's intended input based on the initial NL-based input. As another example, the additional signals may include an indication of whether the user (e.g., by selecting an "undo" button) reverted the generated refined input prompt to its initial NL-based input. For example, it can be assumed that if the user reverted the refined input prompt to its initial NL-based input, this would be an indication that the refined input prompt did not accurately reflect the user's intended input. As yet another example, the additional signals may include an indication of whether the user selected response content generated based on processing the refined input prompt (e.g., when response content generated based on processing NL-based input and / or another refined input prompt is also provided). For example, it can be assumed that if a user selects a response based on the processed refined input prompt, this is an indication that the refined input prompt accurately reflects the user's intended input. As another example, additional signals may include indications of whether one or more follow-up inputs have been received from the user. For instance, if a user has to clarify or correct one or more aspects of a follow-up NL-based input, this could be an indication that the refined input prompt did not represent the user's intended input based on the initial NL-based input. Although several additional signals have been discussed, it will be understood that any suitable additional signal may be used.
[0065] It should be understood that the LLM can be fine-tuned in any suitable manner. For example, NL-based inputs and corresponding refined input cues can be obtained from specific training examples (e.g., retrieved from training data database 152). Refined input cues can be generated based on processing the NL-based inputs using the LLM. The generated refined input cues can then be compared with the obtained refined input cues to generate a training loss. Comparing the generated refined input cues with the obtained refined input cues can include, for example, lexicalization, Natural Language Understanding (NLU), Natural Language Processing (NLP), etc. For example, instead of comparing the refined input cues themselves, the embeddings generated based on the responses (in any suitable manner) can be compared to generate the training loss. Furthermore, the LLM can be updated based on training. In some implementations, additional signals (e.g., as described herein) can also be obtained from the training examples and used to generate the training loss.
[0066] Once the LLM has been fine-tuned, it can be deployed to generate refined input cues based on NL-based inputs. For example, the fine-tuned LLM can be used to update the NL-based response system 120. In some implementations, the fine-tuned LLM can be used to further generate training data and fine-tune (e.g., as described herein).
[0067] Alternatively or additionally, training examples can be used to evaluate the performance of one or more LLM 150 or other LLMs. For example, the performance of an LLM can be evaluated based on training examples, such as through human evaluation of refined input cues, or by comparing refined input cues from the LLM output with ground truth output provided by a human annotator.
[0068] Now go to Figure 7A This document depicts flowcharts illustrating example methods 700A for determining the content of a response to an NL-based input, according to various implementations. For convenience, the operation of method 700A is described with reference to the system performing the operation. This system of method 700A includes one or more processors, memories, and / or other components of a computing device (e.g., client device 110, NL-based response system 120, computing device 810, one or more servers, and / or other computing devices). Furthermore, although the operations of method 700A are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.
[0069] At box 710, the system receives NL-based input associated with the client device (e.g., regarding...). Figure 1 and Figure 2 (The same or similar manner as described and / or in a similar manner).
[0070] In some implementations, the system may cause graphical user interface elements to be rendered at the client device. The graphical user interface may be configured to cause refined input prompts to be generated upon selection (e.g., as described below with respect to box 720). In some implementations, the system may make a determination regarding whether to cause the rendering of graphical user interface elements. This determination may be based on, for example, a quality metric of NL-based input and / or on determining whether the NL-based input contains harmful content.
[0071] At box 720, the system generates a refined input prompt corresponding to the NL-based input based on a first LLM output, which is generated based on at least NL-based input processed using LLM.
[0072] In some implementations, the system can generate refined input prompts based on processing at least (i) NL-based input and (ii) one or more exemplary refined input prompts using an LLM. Alternatively, the LLM can be fine-tuned to generate corresponding refined input prompts for a given NL-based input.
[0073] In some implementations, the system may generate refined input prompts based on at least (i) NL-based input and (ii) context data processed using LLM. Context data may include user data associated with the user of the client device. Alternatively, context data may include historical data associated with one or more previous turns of the conversation between the user and the client device.
[0074] At box 730, the system causes the refined input prompts to be rendered on the client device.
[0075] In some implementations, the system may select one or more terms from the refined input suggestions as user-selectable terms. The system may then cause the indication of the user-selectable terms from the refined input suggestions to be rendered at the client device. In some implementations, the selection of one or more terms from the refined input suggestions as user-selectable terms may be based on metadata included in the first LLM output.
[0076] In some implementations, the system can detect user interaction with a specific user-selectable term. In response, the system can cause one or more alternative terms corresponding to the specific user-selectable term to be rendered at the client device. The system can then detect a user selection of the specific alternative term at the client device. In response, the system can replace the specific user-selectable term with the specific alternative term in a refined input prompt to generate an updated refined input prompt. Response content can then be generated based on processing the updated refined input prompt (e.g., in response to user input indicating acceptance of the updated refined input prompt received at the client device). In some implementations, the system can store the indication of selection of the alternative term for use in generating subsequent refined input prompts using LLM.
[0077] At box 740, in response to user input indicating acceptance of refined input prompts received at the client device, the system generates response content for NL-based input based on a second LLM output generated based on refined input prompts processed using LLM.
[0078] In some implementations, in response to a user input indicating acceptance of the refined input prompt received on the client device, the system may store the refined input prompt and the NL-based input together as training examples for fine-tuning the LLM. In some implementations, in response to a user input requesting restoration to the NL-based input received on the client device (e.g., selection of an "undo" button), the system may bypass storing the refined input prompt and the NL-based input as training examples (or store training examples as examples of poorly refined input prompts).
[0079] At box 750, the system causes the response content to NL-based input to be rendered at the client device.
[0080] In some implementations, the system can generate a second response content to the NL-based input based on a third LLM output, which is generated by processing the NL-based input using an LLM. The second response content to the NL-based input can be rendered at the client device. The system can then detect user input at the client device indicating a selection of one of the first and second response contents. In response to the first response content being selected, the system can store the refined input hint along with the NL-based input as training examples to fine-tune the LLM to generate corresponding refined input hints based on a given NL-based input. In response to the second response content being selected, the system can bypass storing the refined input hint along with the NL-based input as training examples, or the refined input hint along with the NL-based input can be stored together as an example of a poorly quality refined input hint.
[0081] In some implementations, the user can modify the refined input hints at the client device to generate updated refined input hints. In response to user input received at the client device indicating acceptance of the updated refined input hints, the system can generate response content for NL-based input based on processing the updated refined input hints using LLM. The response content for the NL-based input can then be rendered at the client device; and the system can store the updated refined input hints and the NL-based input together as training examples for fine-tuning the LLM.
[0082] In some implementations, the system can fine-tune the LLM based on one or more training examples to generate corresponding refined input cues based on a given NL-based input, where each training example includes the NL-based input and the corresponding refined input cues (e.g., to relate to...). Figure 7B (The same or similar manner as described herein, and / or any other manner as described herein).
[0083] Figure 7B A flowchart illustrating the example method is provided.
[0084] Now go to Figure 7B A flowchart illustrating an example method 700B for fine-tuning an LLM, based on various implementations, is depicted. For convenience, the operation of method 700B is described with reference to the system performing the operation. This system of method 700B includes one or more processors, memories, and / or other components of a computing device (e.g., client device 110, NL-based response system 120, computing device 810, one or more servers, and / or other computing devices). Furthermore, although the operations of method 700B are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, and / or added.
[0085] At box 760, the system obtains one or more training examples (e.g., with respect to...). Figure 1 and Figure 2 (In the same or similar manner as described herein, and / or in other manner as described herein). Each training example may include NL-based input and a corresponding refined input cue.
[0086] At box 770, the system fine-tunes the LLM based on one or more training examples to generate a corresponding refined input cue based on a given NL-based input associated with a client device. This refined input cue can be used in response to user input received at the client device indicating acceptance of the corresponding refined input cue. Response content for a given NL-based input is generated based on the LLM output, which is generated by processing the corresponding refined input cue using the LLM. This response content is then rendered at the client device.
[0087] Figure 8 Example architectures of computing devices based on various implementations are described.
[0088] Now go to Figure 8 This diagram depicts a block diagram of an example computing device 810 that can be optionally used to perform one or more aspects of the techniques described herein. In some implementations, one or more of a client device, a cloud-based automation assistant component, and / or other components may include one or more components of the example computing device 810.
[0089] Computing device 810 typically includes at least one processor 814 that communicates with a plurality of peripheral devices via a bus subsystem 812. These peripheral devices may include a storage subsystem 824 (including, for example, a memory subsystem 825 and a file storage subsystem 826), a user interface output device 820, a user interface input device 822, and a network interface subsystem 816. The input and output devices allow interaction with a user of computing device 810. The network interface subsystem 816 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0090] User interface input device 822 may include a keyboard, pointing devices such as a mouse, trackball, touchpad or graphics tablet, scanner, touchscreen integrated into a display, audio input devices such as a voice recognition system, microphone, and / or other types of input devices. Generally, the term "input device" is used to encompass all possible types of means and methods for inputting information into computing device 810 or into a communication network.
[0091] User interface output device 820 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual displays, such as via an audio output device. Generally, the term "output device" is used to encompass all possible types of means and methods for outputting information from computing device 810 to a user or to another machine or computing device.
[0092] Storage subsystem 824 stores the functional programming and data structures of some or all of the modules described herein. For example, storage subsystem 824 may include selected aspects for implementing the methods disclosed herein and implementations. Figure 1 The logic of the various components described in the text.
[0093] These software modules are typically executed by processor 814 alone or in combination with other processors. The memory 825 used in storage subsystem 824 may include multiple memories, including main random access memory (RAM) 830 for storing instructions and data during program execution and read-only memory (ROM) 832 for storing fixed instructions. File storage subsystem 826 can provide persistent storage for program and data files and may include hard disk drives, floppy disk drives, and associated removable media, CD-ROM drives, optical disk drives, or removable media cartridges. Modules implementing certain functionalities may be stored by file storage subsystem 826 within storage subsystem 824 or in other machines accessible to processor 814.
[0094] The bus subsystem 812 provides mechanisms for enabling the various components and subsystems of the computing device 810 to communicate with each other as intended. Although the bus subsystem 812 is schematically shown as a single bus, alternative implementations of the bus subsystem 812 may use multiple buses.
[0095] The computing device 810 can be of different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the constantly evolving nature of computers and networks, [the following applies]. Figure 8 The description of the computing device 810 depicted herein is intended only as a specific example to illustrate some implementation methods. Many other configurations of the computing device 810 may have... Figure 8 The computing device depicted in the text has more or fewer components.
[0096] Where the systems described herein collect or otherwise monitor personal information about users, or may utilize personal information and / or monitored information, they may provide users with the opportunity to control whether programs or functions collect user information (e.g., information about the user's social networks, social behaviors or activities, occupation, user preferences, or the user's current geographic location) or to control whether and / or how content that may be more relevant to the user is received from content servers. Additionally, certain data may be processed in one or more ways before it is stored or used, such that personally identifiable information is removed. For example, a user's identity may be processed to the point that the user's personally identifiable information cannot be determined, or, where geographic location information is available, the user's geographic location may be generalized (e.g., to the city, zip code, or state level) to the point that the user's specific geographic location cannot be determined. Therefore, users can control how information about themselves is collected and / or used.
[0097] In some implementations, a method implemented by one or more processors is provided, the method comprising: receiving NL-based input associated with a client device; generating a refined input prompt corresponding to the NL-based input based on a first LLM output, the first LLM output being generated based on processing at least the NL-based input using an LLM; causing the refined input prompt to be rendered at the client device; in response to user input received at the client device indicating acceptance of the refined input prompt, generating response content for the NL-based input based on a second LLM output, the second LLM output being generated based on processing the refined input prompt using an LLM; and causing the response content for the NL-based input to be rendered at the client device.
[0098] These and other implementations of the techniques disclosed herein may optionally include one or more of the following features.
[0099] In some implementations, generating refined input hints may be based on using an LLM to process at least (i) NL-based input and (ii) one or more exemplary refined input hints. In some additional or alternative implementations, the LLM may be fine-tuned to generate corresponding refined input hints for a given NL-based input.
[0100] In some additional or alternative implementations, generating refined input cues may be based on processing at least (i) NL-based input and (ii) context data using LLM. In some versions of those implementations, the context data may include user data associated with the user of the client device. In some additional or alternative implementations, NL-based input may be received as part of a multi-turn dialogue with the user of the client device, and the context data may include historical data associated with one or more previous turns of the multi-turn dialogue.
[0101] In some additional or alternative implementations, the method further includes: selecting one or more terms from the refined input prompt as user-selectable terms; and causing an indication of the user-selectable terms from the refined input prompt to be rendered at the client device. In some versions of those implementations, the selection of one or more terms from the refined input prompt as user-selectable terms may be based on metadata included in the first LLM output.
[0102] In some additional or alternative versions of those implementations, the method further includes: in response to detecting user interaction at the client device with a specific user-selectable term among one or more user-selectable terms, causing one or more alternative terms corresponding to the specific user-selectable term to be rendered at the client device; and in response to detecting user selection of the specific alternative term at the client device, replacing the specific user-selectable term with the specific alternative term in a refined input prompt to generate an updated refined input prompt, wherein the response content may be generated using LLM processing of the updated refined input prompt based on user input received at the client device indicating acceptance of the updated refined input prompt. In some further versions of those implementations, the method further includes: storing an indication of selection of an alternative term for use in generating subsequent refined input prompts using LLM.
[0103] In some additional or alternative implementations, the method further includes causing a graphical user interface element to be rendered at the client device, wherein refined input prompts are generated in response to a user selection of the graphical user interface element at the client device. In some versions of those implementations, the method further includes determining whether to cause the rendering of the graphical user interface element based on a quality metric of the NL-based input and / or based on determining that the NL-based input contains harmful content.
[0104] In some implementations, the response content may be a first response content, and the method further includes: generating a second response content for an NL-based input based on a third LLM output, the third LLM output being generated based on processing the NL-based input using an LLM; causing the second response content for the NL-based input to be rendered at a client device; detecting user input at the client device indicating a selection of one of the first and second response contents; and in response to the selection of the first response content: storing the refined input prompt and the NL-based input together as training examples for fine-tuning the LLM to generate a corresponding refined input prompt based on a given NL-based input.
[0105] In some implementations, the method further includes: modifying a refined input prompt based on user input received at the client device to generate an updated refined input prompt; generating response content for NL-based input based on processing the updated refined input prompt using LLM in response to user input received at the client device indicating acceptance of the updated refined input prompt; causing the response content for the NL-based input to be rendered at the client device; and storing the updated refined input prompt and the NL-based input together as training examples for fine-tuning the LLM to generate corresponding refined input prompts based on given NL-based inputs.
[0106] In some implementations, the method further includes: responding to user input indicating acceptance of the refined input prompt received at the client device; storing the refined input prompt and the NL-based input together as training examples for fine-tuning the LLM to generate corresponding refined input prompts based on given NL-based input. In some versions of those implementations, the method further includes: bypassing storing the refined input prompt and the NL-based input as training examples in response to user input requesting restoration to the NL-based input received at the client device.
[0107] In some implementations, the method further includes: fine-tuning the LLM based on one or more training examples to generate corresponding refined input cues based on a given NL-based input, wherein each training example includes the NL-based input and the corresponding refined input cues.
[0108] In some implementations, a method implemented by one or more processors includes: obtaining one or more training examples, each training example including an NL-based input and a corresponding refined input cue; fine-tuning an LLM based on the one or more training examples to generate a corresponding refined input cue based on a given NL-based input associated with a client device, the corresponding refined input cue being available in response to user input received at the client device indicating acceptance of the corresponding refined input cue; generating response content for the given NL-based input based on the LLM output, the LLM output being generated based on processing the corresponding refined input cue using the LLM; and causing the response content to be rendered at the client device.
[0109] Additionally, some implementations include one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in associated memory, and wherein the instructions are configured to cause any of the methods described above to be performed. Some implementations also include one or more computer-readable storage media (e.g., transient and / or non-transient) storing computer instructions executable by one or more processors to perform any of the methods described above. Some implementations also include a computer program product comprising instructions executable by one or more processors to perform any of the methods described above.
Claims
1. A method implemented by one or more processors, the method comprising: Receive natural language (NL) based input associated with the client device; Refined input prompts corresponding to the NL-based input are generated based on the output of a first large-scale language model (LLM), wherein the first LLM output is generated based on processing at least the NL-based input using an LLM. This causes the refined input prompts to be rendered on the client device; In response to a user input received at the client device indicating acceptance of the refined input prompt, response content for the NL-based input is generated based on a second LLM output, which is generated based on processing the refined input prompt using an LLM. as well as This causes the response content to the NL-based input to be rendered at the client device.
2. The method of claim 1, wherein generating the refined input prompt is based on processing at least (i) the NL-based input and (ii) one or more exemplary refined input prompts using the LLM.
3. The method of claim 1 or 2, wherein the LLM has been fine-tuned to generate corresponding refined input prompts for a given NL-based input.
4. The method as described in any of the preceding claims, wherein the generation of the refined input prompt is based on processing at least (i) the NL-based input and (ii) the context data using the LLM.
5. The method of claim 4, wherein the context data includes user data associated with a user of the client device.
6. The method of claim 4 or 5, wherein the NL-based input is received as part of a multi-turn dialogue with a user of the client device, and wherein the context data includes historical data associated with one or more previous turns of the multi-turn dialogue.
7. The method as described in any of the preceding claims, further comprising: Select one or more of the refined input suggestions as user-selectable terms; as well as This causes the indication of the user-selectable terms from the refined input prompts to be rendered on the client device.
8. The method of claim 7, wherein selecting one or more terms from the refined input prompt as user-selectable terms is based on metadata included in the first LLM output.
9. The method of claim 7 or 8, further comprising: In response to detecting user interaction with a specific user-selectable term among the one or more user-selectable terms at the client device, one or more alternative terms corresponding to the specific user-selectable term are rendered at the client device; as well as In response to detecting a user selection of a specific alternative term at the client device, the specific user-selectable term is replaced with the specific alternative term in the refined input prompt to generate an updated refined input prompt. The response content is generated by processing the updated refined input prompt using the LLM in response to user input received at the client device indicating acceptance of the updated refined input prompt.
10. The method of claim 9, further comprising: The selection of the alternative terms is stored for use in generating subsequent refined input prompts using the LLM.
11. The method as described in any of the preceding claims, further comprising: This causes graphical user interface elements to be rendered on the client device, wherein the refined input prompts are generated in response to user selection of the graphical user interface elements on the client device.
12. The method of claim 11, further comprising: Whether rendering of the graphical user interface elements is caused by quality metrics based on the NL-based input and / or by determining that the NL-based input contains harmful content.
13. The method of any of the preceding claims, wherein the response content is a first response content, the method further comprising: A second response to the NL-based input is generated based on a third LLM output, which is generated based on processing the NL-based input using an LLM. This causes the second response content to the NL-based input to be rendered at the client device; The client device detects user input indicating a selection of one of the first response content and the second response content; as well as In response to the first response content being selected: The refined input prompts and the NL-based inputs are stored together as training examples to fine-tune the LLM to generate corresponding refined input prompts based on the given NL-based inputs.
14. The method as described in any of the preceding claims, further comprising: The refined input prompt is modified based on user input received at the client device to generate an updated refined input prompt; In response to a user input received at the client device indicating acceptance of the updated refined input prompt, response content for the NL-based input is generated based on processing the updated refined input prompt using the LLM; This causes the response content to the NL-based input to be rendered at the client device; as well as The updated refined input hints and the NL-based input are stored together as training examples to fine-tune the LLM to generate corresponding refined input hints based on a given NL-based input.
15. The method as described in any one of the preceding claims, further comprising: The accepted user input in response to an instruction received at the client device for the refined input prompt; The refined input prompts and the NL-based inputs are stored together as training examples to fine-tune the LLM to generate corresponding refined input prompts based on the given NL-based inputs.
16. The method of claim 15, further comprising: In response to a user input request received at the client device to restore the NL-based input, the refined input prompts and the NL-based input are bypassed and stored as training examples.
17. The method as described in any of the preceding claims, further comprising: The LLM is fine-tuned based on one or more training examples to generate corresponding refined input prompts based on a given NL-based input, wherein each training example includes the NL-based input and the corresponding refined input prompt.
18. A method implemented by one or more processors, the method comprising: Obtain one or more training examples, each of which includes NL-based input and a corresponding refined input cue; The LLM is fine-tuned based on one or more training examples to generate corresponding refined input prompts based on a given NL-based input associated with a client device. The corresponding refined input prompts are usable in response to user input received at the client device indicating acceptance of the corresponding refined input prompts. Response content for the given NL-based input is generated based on the LLM output, which is generated based on processing the corresponding refined input prompts using the LLM. The response content is then rendered at the client device.
19. A system comprising: One or more hardware processors; as well as A memory that stores instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations according to any one of claims 1 to 17.
20. A temporary or non-temporary computer-readable storage medium for storing instructions, which, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations according to any one of claims 1 to 17.