Streaming chat in SERPs
By integrating a GLM with a search engine to generate streaming content on the SERP using user input and web page information, the system addresses the limitations of conventional search engines and GLMs, offering accurate and comprehensive information directly on the SERP.
Patent Information
- Application Number
- JP2025538603
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-06-15
- Filing Date
- 2024-01-04
- Publication Date
- 2026-02-10
AI Technical Summary
Conventional search engines and generative language models (GLMs) are deficient in providing accurate and comprehensive information in response to certain types of user queries, often requiring users to navigate multiple web pages to gather information and generating incorrect outputs.
Integrating a GLM with a search engine to generate streaming content on a search engine results page (SERP) using user input, web page information, and previous dialogue turns, summarizing and formatting the content with citations and credibility, and leveraging the search engine for accurate information retrieval.
Enhances user experience by providing concise summaries and accurate answers directly on the SERP, reducing the likelihood of false outputs and improving information accessibility.
Smart Images

Figure 2026504814000001_ABST
Abstract
Description
[Background technology]
[0001] background A conventional computer-implemented search engine is configured to receive a search query and infer the information retrieval intent of the user issuing the query (e.g., ascertain whether the user wants to navigate to a particular page, whether the user intends to purchase an item or service, whether the user is searching for a fact, whether the user is searching for images or videos, etc.). The search engine identifies results based on the inferred information retrieval intent and returns a search engine results page (SERP) to the user's computing device. The SERP may include links to web pages, snippets of text extracted from web pages, images, videos, knowledge cards (graphical items containing information about entities such as people, places, companies, etc.), instant answers (graphical items that depict answers to questions posed in the query), widgets (e.g., a graphical calculator with which the user can interact), supplemental content (e.g., advertisements related to the query), etc.
[0002] While search engines are frequently updated with features aimed at improving the user experience (and providing users with more relevant results), they are not fully equipped to provide certain types of information. For example, search engines are not configured to provide output that requires inference on the content of a web page or that is based on several different sources of information. For example, when a user queries, "How many home runs did Babe Ruth hit by the time she turned 30?", a traditional search engine will return, among other information, a knowledge card about Babe Ruth (which may feature a picture of Babe Ruth, her birth date, etc.), suggested alternative queries (e.g., "How many hits did Babe Ruth hit in her career?"), and links to web pages containing statistics. To get the answer, the user must visit the web page containing the statistics and calculate the answer themselves.
[0003] In another example, upon receiving the query "Give me a list of famous people born in Seattle and Chicago," a traditional search engine may return knowledge cards about the cities of Chicago and Seattle, a link to a first web page containing a list of people from Chicago, and a link to a second web page containing a list of people from Seattle. However, the search engine cannot infer the content of the two web pages to result in a list that includes identities of people from both Chicago and Seattle.
[0004] Generative language models (GLMs), also known as large-scale language models (LLMs), have been developed relatively recently. One example of a GLM is the Generative Pre-trained Transformer 3 (GPT-3). Another example of a GLM is the BigScience Language Open-science Open-access Multilingual (BLOOM) model, which is also a transformer-based model. Briefly, a GLM is configured to generate output (human language text, source code, music, video, etc.) in near real time (e.g., within a few seconds of receiving the prompt) based on prompts set by a user. The GLM generates content based on the training data on which it is trained. Thus, in response to receiving the prompt, "How many home runs did Babe Ruth hit before she turned 30?", the GLM can output, "Babe Ruth hit 94 home runs before she turned 30." In another example, in response to receiving the prompt "Please provide a list of famous people born in Seattle and Chicago," a GLM can output two separate lists of people (one in Seattle and one in Chicago), where the list of people born in Chicago includes Barrack Obama. However, in both of these examples, the GLM outputs incorrect information, such as that Babe Ruth hit over 94 home runs before turning 30 and that Barrack Obama was born in Hawaii (not Chicago). Thus, both traditional search engines and GLMs are deficient in identifying and / or generating appropriate information in response to certain types of user input. Summary of the Invention
[0005] summary The following is a brief summary of what is described more fully herein. This summary is not intended to limit the scope of the claims.
[0006] This specification describes various techniques related to providing streaming content by a GLM on a SERP and / or a web page provided within the SERP. Information provided as input to the GLM that the GLM uses to generate output is called a prompt. According to the techniques described herein, the prompts that the GLM uses to generate output can include 1) user input, such as a query, and 2) information from the web page the user is viewing or information obtained by a search engine. As described in more detail herein, the prompts can also include previous dialogue turns.
[0007] In one example, a browser on a client computing device loads a web page of a search engine and receives a query submitted by a user of the client computing device. The browser transmits the query to a computing system running the search engine, and the search engine identifies search results and generates a search engine results page (SERP) based on the query. The search results may include web pages related to the query, knowledge cards, instant answers, entity descriptions, supplemental content, etc. The search engine returns the SERP to the browser, and when the client computing device is powered on, the SERP is displayed on the display of the client computing device.
[0008] In another example, the system organizes and summarizes information from classical retrieval-based search engines into a semantically meaningful format, making the information easier for search engine users to understand and navigate. The system does this by first creating a summary that provides an overview of the information, for example, from the top N (e.g., 10 or other number) search results, and then creating disambiguating subsections about various aspects of the original search query based on the intent of the original search query. These subsections use citation links to attribute the summarized information to its source and provide credibility. The goal of the system is to help users quickly find and understand the information they are looking for by providing a curated and structured view of search engine results pages.
[0009] The system retrieves relevant information from a search engine based on a user's search query. It then leverages GLM to summarize the content according to the intent detected from the query. In some cases, the system can generate a direct answer to the query and provide related literature supporting the information. Additionally, the system uses information from reference documents to provide a concise summary of key facts or aspects related to the user's query. The model accesses data such as the date and location of the query, the top N web results, and surrounding information for each result. After a user enters a search query, the system then uses a search engine to retrieve relevant web pages. It then uses a large-scale language model to detect the user's intent, summarizes content from the retrieved documents, generates a direct answer, formats the generated content (bolding text and grouping content under different headers), cites reference documents, and provides a concise summary of key facts, events, or aspects of the user's query based on information from the reference documents. The model is provided with date and location information and the top N web results along with relevant portions of those results.
[0010] The systems described herein go beyond the capabilities of traditional search engines by summarizing and generating answers to user input and providing concise summaries of key facts, aspects, or other disambiguating information relevant to a query. Traditional search engines typically only retrieve and rank relevant content based on a user's query, without providing additional information or analysis. The described systems and methods achieve new capabilities by leveraging large-scale language models.
[0011] As an example, a search engine may receive the query, "How many home runs did Babe Ruth hit before she turned 30?" and the search results identified by the search engine include Babe Ruth's birth date and her season-by-season statistics. The GLM obtains such information as part of a prompt along with the query. Because the prompt includes Babe Ruth's season-by-season home run totals, the GLM can infer about such data and provide output based on the information identified by the search engine as relevant to the query. Thus, the GLM may output, "Babe Ruth hit 284 home runs before she turned 30." This information may be streamed in a chat window presented on or alongside the SERP the user is viewing. In another example, the chat window is presented alongside a web page the user navigates to from the SERP.
[0012] The techniques described herein offer various advantages over conventional search engine and / or GLM techniques. Specifically, integration with a GLM allows a search engine to provide information to end users that conventional search engines cannot provide. Additionally, the GLM described herein is provided with information obtained by the search engine for use in generating output, thereby reducing the likelihood that the GLM will issue false or irrelevant output.
[0013] The foregoing summary presents a simplified summary to provide a basic understanding of some aspects of the systems and / or methods discussed herein. This summary is not an extensive overview of the systems and / or methods discussed herein. It is not intended to identify key / critical elements or delineate the scope of such systems and / or methods. Its sole purpose is to present some concepts in a simplified form as a prelude to the more detailed description that is presented later. [Brief explanation of the drawings]
[0014] BRIEF DESCRIPTION OF THE DRAWINGS [Figure 1] FIG. 1 is a functional block diagram of a computing system in accordance with various aspects described herein. [Figure 2] 1 illustrates a computing system with additional elements for providing streaming chat in a SERP or web page. [Figure 3] 1 illustrates a GUI of an operating system installed on a client computing device in accordance with various aspects described herein. [Figure 4] 1 illustrates a GUI displaying a streaming semantic SERP according to various aspects described herein. [Figure 5] 1 illustrates a GUI on a communication device, such as a tablet, mobile phone, or smartphone, according to one or more aspects described herein. [Figure 6] 1 illustrates a GUI displaying a web page with a streaming SERP overlaid, according to one or more aspects described herein. [Figure 7] 7 depicts a flow diagram illustrating a method 700 for providing a streaming experience in a SERP in a computing system, according to one or more aspects described herein. [Figure 8] 1 depicts a high-level diagram of an exemplary computing device that can be used in accordance with the systems and methodologies disclosed herein. DETAILED DESCRIPTION OF THE INVENTION
[0015] Detailed Description Various techniques relating to streaming information from a GLM into a SERP on a computing device will now be described with reference to the drawings, wherein like reference numerals are used to refer to like elements throughout. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of one or more aspects. It may be apparent, however, that such aspects can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form to facilitate describing one or more aspects. Furthermore, it should be understood that functionality described as being performed by a certain system component may be performed by multiple components. Similarly, a component may be configured to perform functions described, for example, as being performed by multiple components.
[0016] Furthermore, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, the phrase "X uses A or B" is intended to mean any of the natural inclusive permutations. That is, the phrase "X uses A or B" is satisfied by either the instance where X uses A, where X uses B, or where X uses both A and B. Additionally, as used in this application and the appended claims, the articles "a" and "an" should be interpreted broadly to mean "one or more" unless otherwise specified or clear from the context that the reference is to the singular.
[0017] Furthermore, as used herein, the terms "component" and "system" are intended to encompass a computer-readable data storage comprised of computer-executable instructions that, when executed by a processor, cause certain functions to be performed. The computer-executable instructions may include routines, functions, etc. It should also be understood that a component or system may be localized on a single device or distributed across several devices. Furthermore, as used herein, the term "exemplary" is intended to mean serving as an example or instance of something and is not intended to indicate preference.
[0018] Described herein are various techniques related to providing streaming chat within search engines and SERPs and / or related web pages using generative language models (GLMs), also known as large-scale language models (LLMs). The described systems and methods enable search engine results pages (SERPs) to enable streaming content, particularly streaming result summaries, streaming layout and organization, and streaming chat / interactions from the generative language models (GLMs), and more broadly, streaming from any model or content source, such as images, video generators, audio, and other chained or pipelined models. In another example, real-time conditional (adaptive) text generation is provided, which, along with other model-level streaming optimizations, allows responses to adapt to dynamic events.
[0019] A GLM can take several minutes to generate a response. Typical GPU load times for prompts are on the order of 1,000 tokens per second, with output generation times on the order of 5–10 tokens per second. A typical token averages roughly three Latin characters in English. A prompt can be 2,000 tokens long, and a response can be 1,000 tokens long. These times are too slow to provide responses on search engines that traditionally return responses in a few seconds at most. Furthermore, many prompts are chained in that they gather and organize information by invoking sub-prompts, making API calls, launching search engines, etc. Various aspects of this specification define several user interface, user experience, and system-level innovations that alleviate much of this latency.
[0020] 1, a functional block diagram of a computing system 100 is shown in accordance with various aspects described herein. While shown as a single system, it should be understood that computing system 100 may include several different server computing devices, may be distributed across a data center, etc. Computing system 100 is configured to obtain information based on a query written by a user, and is further configured to provide the obtained information as part of a prompt to the GLM.
[0021] A client computing device 102 operated by a user (not shown) communicates with computing system 100 over network 104. Client computing device 102 may be any suitable type of client computing device, such as a desktop computer, a laptop computer, a tablet (slate) computing device, a video game system, a virtual reality or augmented reality computing system, a mobile phone, a smart speaker, or other suitable computing device.
[0022] Computing system 100 includes processor 106 and memory 108, where memory 108 contains instructions executed by processor 106. More specifically, memory 108 includes search engine 110 and GLM 112, the operation of which is described in more detail below. Computing system 106 also includes data stores 114-122, which store data accessed by search engine 110 and / or GLM 112. More specifically, data stores 114-122 include web index data store 114, instant answers data store 116, knowledge graph data store 118, supplemental content data store 120, and interaction history data store 122. Web index data store 114 includes a web index that indexes web pages by keywords contained in or associated with the web pages. The instant answer data store 116 includes an index of instant answers indexed by queries, query terms, and / or terms semantically similar or equivalent to the queries and / or query terms. For example, the instant answer "2.16 meters" may be indexed by the query "How tall is Shaquille O'Neal" (and semantically similar or equivalent queries such as "How tall is Shaquille O'Neal").
[0023] The knowledge graph data store 118 includes a knowledge graph, which includes data structures about entities (people, places, things, etc.) and their relationships to each other, thereby representing the relationships between the entities. The search engine 110 can use the knowledge graph in connection with presenting entity cards on a search engine results page (SERP). The supplemental content data store 120 includes supplemental content that the search engine 110 can return based on a query.
[0024] The interaction history data store 122 includes an interaction history, which includes interaction information regarding the user and the GLM 112. For example, the interaction history may include, with respect to a user, an identification of a conversation that took place between the user and the GLM 112, input provided by the user to the GLM 112 for multiple interaction turns during the conversation, responses during the conversation generated by the GLM 112 in response to input from the user, queries generated by the GLM during the conversation that are used by the GLM 112 to generate the responses, etc. Additionally, the interaction history may include context obtained by the search engine 110 during the conversation; for example, with respect to a conversation, the interaction history 122 may include content from SERPs generated based on queries posted by the user and / or GLM 112 during the conversation, content from web pages identified by the search engine 110 based on queries posted by the user and / or GLM 112 during the conversation, etc. It should be understood that data stores 114-122 are presented to illustrate a representative sample of the types of data accessible to search engine 110 and / or GLM 112, and that there are many other data sources accessible to search engine 110 and / or GLM 112, such as data stores containing real-time financial information, data stores containing real-time weather information, data stores containing real-time sports information, data stores containing images, data stores containing video, data stores containing maps, etc. Such information sources may be available to search engine 110 and / or GLM 112.
[0025] The search engine 110 includes a web search module 124, an instant answer search module 126, a knowledge module 128, a supplemental content search module 130, and a SERP constructor module 132. The web search module 124 is configured to search the web index data store 114 based on queries received by a user, queries generated by the search engine 110 based on queries received by a user, and / or queries generated by the GLM 112 based on a user's interaction with the GLM 112. Similarly, the instant answer search module 126 is configured to search the instant answer data store 116 based on queries received by a user, queries generated by the search engine 110 based on queries received by a user, and / or queries generated by the GLM 112 based on a user's interaction with the GLM 112. The knowledge module 128 is configured to search the knowledge graph data store 118 based on queries received by a user, queries generated by the search engine 110 based on queries received by a user, and / or queries generated by the GLM 112 based on a user's interaction with the GLM 112. Similarly, the supplemental content search module 130 is configured to search the supplemental content data store 120 based on queries received by a user, queries generated by the search engine 110 based on queries received by a user, and / or queries generated by the GLM 112 based on a user's interaction with the GLM 112.
[0026] The SERP constructor module 132 is configured to construct a SERP based on information identified by the searches performed by modules 124-130. For example, the SERP may include links to web pages identified by the web search module 124, instant answers identified by the instant answer search module 126, entity cards (containing information about entities) identified by the knowledge module 128, and supplemental content identified by the supplemental content search module 130. The SERP may also include widgets, cards showing the current weather, etc. The SERP constructor module 132 may also generate structured, semi-structured, and / or unstructured data representing the content of the SERP or portions of the content of the SERP. For example, the SERP constructor module 132 generates a JSON document containing information obtained by the search engine 110 based on one or more searches performed on the data stores 114-120 (or other data stores). In one example, the SERP constructor module 132 generates data in a structure / format used by the GLM 112 as part of a prompt.
[0027] As discussed above, the operation of search engine 110 is improved based on GLM 112, and the operation of GLM 112 is improved based on search engine 110. For example, search engine 110 may provide output that search engine 110 was not previously able to provide (e.g., based on output generated by GLM 112), and GLM 112 is improved by using information obtained by search engine 110 to generate the output (e.g., information identified by search engine 110 may be included as part of the prompts used by GLM 112 to generate the output). Specifically, GLM 112 generates results based on information obtained by search engine 110, and the results are likely to be more accurate when compared to results generated by GLM 112 not based on such information, because search engine 110 is designed over time to curate information sources to ensure its accuracy.
[0028] With continued reference to Figure 1, Figure 2 illustrates a computing system 100 that includes a processor 106 and a memory 108 as described with respect to Figure 1, where the memory 108 includes instructions that are executed by the processor 106. More specifically, the memory 108 includes a search engine 110 and a GLM 112, the operation of which is described in more detail above with respect to Figure 1. The computing system 106 also includes data stores 114-122, which store data that is accessed by the search engine 110 and / or the GLM 112 as described above.
[0029] The search engine 110 includes a web search module 124 , an instant answer search module 126 , a knowledge module 128 , a supplemental content search module 130 , and a SERP constructor module 132 .
[0030] In addition to the elements described with respect to FIG. 1 , the computing system of FIG. 2 includes one or more chat buffers 202 that store response data (e.g., text, images, video, graphics, etc.) received from the GLM 112 for transmission to the client computing device 102 in response to a user query. The response data is used by the SERP constructor module 132 to generate a SERP that includes a streaming conversation with the user, such as summary information, instant answers, entity descriptions, search results, supplemental content, etc. Before transmitting the response data to the client device 102, the response data is parsed or chunked into data subsets for streaming transmission. A chunk analyzer module 204 analyzes each data chunk to identify objectionable content (e.g., language, images, etc.). If objectionable content is identified, the objectionable content can be removed before transmitting the chunk to the client device 102. In another embodiment, the entire chunk can be removed, or the entire set of response data to which the chunk belongs can be removed. In such cases, the search system can resume searching to respond to the user's query.
[0031] The search system also includes one or more caches 206 that cache user inputs and responses generated by the GLM 112 so that responses to relatively frequently listed user inputs are readily available. This functionality improves response time. Additionally, the search system includes a synchronization module 208 that synchronizes multiple data streams being rendered using metadata, etc.
[0032] 1 and 2 store cached information, such as cached web pages or other data sources, for responding to chat queries and / or providing search results or other information in the SERPs provided to client devices. Cached web pages are periodically updated and / or invalidated to maintain current source data.
[0033] 3, a schematic diagram illustrating a GUI 300 of a SERP displayed on a client computing device is shown, according to various aspects described herein. The computing device may be any suitable computing device, including, but not limited to, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone, a wearable computing device (including a watch or headgear), etc.
[0034] A pointer 310 can be used in connection with selecting graphical elements displayed within the GUI 300, and the position of the pointer 310 is based on the corresponding position of a mouse, the position of the user's gaze relative to the display, the position of numbers on a touch-sensitive display, etc.
[0035] The SERP includes an input field 314 into which a user can enter user input. The user input is submitted to a search engine (not shown), which returns instant answers 316 (if applicable), entity descriptions 318 (if applicable), one or more search results 320, and supplemental content 322. The instant answers 316 are answers provided by the search engine in response to a query, for example, without the user having to navigate away from the search results 320. The instant answers may be cached answers that have been previously seen in response to the same user input.
[0036] It will be understood that the illustrated orientation of the input field 314, instant answer 316, entity description 318, search results 320, and supplemental content 322 relative to one another is presented by way of example only and is not intended to limit the particular placement of these elements within a SERP. For example, the input field 314 may be presented below the results 320 and entity description 318. In another example, the position of the instant answer 316 may be swapped with the position of the entity description 318. In another example, the query field 314, instant answer 316, entity description 318, results 320, and supplemental content 322 may be presented in any order as a vertical stack of fields.
[0037] In one embodiment, a user can hover a pointer 310 over a particular result in the results field 320, and the system retrieves and presents additional content related to the hovered result in the supplemental content field 322. In another embodiment, when a user hovers a pointer over a particular result, a pop-up window appears showing the source of the information and / or supplemental information, such as advertisements (images or videos) presented on the source page. The hover content can be generated based on the context around the link along with the paragraph content and query. This content can be generated by parallel calls that stream in asynchronously from the main content.
[0038] In the example of FIG. 3, the user places the pointer over search result C, and the system retrieves supplemental content C' and C'' related to search result C.
[0039] The SERP 312 also includes a conversation field 336 that allows the user to interact with the GLM 112 (FIGS. 1 and 2). The conversation field 336 shows user input and the natural language responses that are provided to the user. The conversation field 336 also includes a slider bar 338 that allows the user to scroll up and down within the conversation field. Additionally, one or more chat images 340 can be included within the conversation field 336, which relate to various natural language responses that are provided by the system.
[0040] For example, user input is submitted to the GLM 112 (FIG. 1), which returns natural language responses, images, etc. in response to the user query. The user is allowed to respond to the responses provided by the GLM 112 as if they were conversing with another human. The GLM 112 then provides a second natural language response to the user, and the conversation continues. Meanwhile, the user can place a pointer 310 (see FIG. 1) over one or more of the natural language responses provided by the GLM 112, and the system forms and submits a query that returns supplemental content 322 for presentation on the GUI 300. The supplemental content 322 typically includes one or more selectable links to articles or web pages related to the natural language response over which the pointer 310 is placed, and may also include one or more supplemental images 334. In another embodiment, the supplemental images 334 are displayed as chat images 340 in the conversation field 336. A slider bar 328 is provided in the supplemental content field 322 to allow the user to scroll through the supplemental content and / or images 334 in the supplemental content field.
[0041] In another embodiment, if the client computing device includes a touchscreen, the user can simply swipe up or down using a finger, stylus, or other device. In this embodiment, slider bars 324, 326, 328, and 330 are optional. Alternatively, the slider bars may be held or may appear when the user's finger or stylus is touching the screen in one of the results field 320, entity description field 318, supplemental content field 322, or main display area 304, respectively. When the user releases the touchscreen, the individual slider bars disappear.
[0042] In another embodiment, the client computing device 302 includes a microphone and one or more speakers (not shown) for communicating with the search engine and / or the GLM.
[0043] 4, another GUI 400 of a web page is shown, which may be presented, for example, when a search result corresponding to the web page is selected. Information 404 is displayed on the web page, including, for example, text, images, selectable graphical icons, etc. Additionally, a semantic SERP 406 constructed by the GLM 112 is overlaid on the web page 402.
[0044] The semantic SERP 406 includes several areas, including a summary field 408 and a summary area 410 that (optionally) includes one or more images 412, a text description 414, and additional facts 416. The semantic SERP 406 also includes suggestion chips 418, 420 and a text field 422 where a user can enter text manually or by voice. In addition, a graphical icon for a microphone 424 is provided. When a user selects the microphone icon 424, the user can speak a query into the text field. The semantic SERP 406 is described in more detail below by way of example.
[0045] According to one example, a user may enter an input about Rome, and the search engine 110 may return search results corresponding to a web page that the user can select. The summary field 418 may be generated as follows: "You seem interested in Rome. Here's a history of how the Eternal City came to be." In one example, the text presented to the user is streamed onto the user's display so that it appears as if it were being typed out in real time. The summary area 410 may present one or more images 412 of Rome along with some description 414 about the city's history. For example, the description 414 may be generated as follows: "A brief overview of the history of Rome: a fascinating and complex topic spanning thousands of years and encompassing the rise and fall of one of the most influential civilizations...more." The phrase "more" at the end of the description is clickable by the user to reveal additional information.
[0046] Additional facts or insights can be displayed in the additional facts field 416. For example, text such as "Rome was founded in 753 BC by the twins Romulus and Remus. Rome became the capital of Italy in 1870 and has remained so ever since" can be generated and displayed by the GLM 112. These additional facts can be listed, numbered, itemized, etc. Additionally, source information can be provided in the additional effect field 416, for example, "Source: rome.net, wikipedia.org." Additionally or alternatively, the source information can be revealed in a pop-up or text bubble that appears when the user hovers the pointer 310 over a fact.
[0047] Also provided within the semantic SERP 406 are suggestion chips 418, 420 that the user can click to go to additional information. For example, suggestion chip 418 may have the text "Historical Events in Rome," while suggestion chip 420 displays the text "Historical Sites in Rome." The GLM 112 generates suggestion chips based on user input, web page content, etc.
[0048] The summary 406 is a conversational response generated by the GLM 112 in response to the user's input. The response is short, informative, and demonstrates the system's reasoning capabilities. When rendered, a "poll" section 426 on the left side of the screen displays a row of "suggestion pills" 428, which are conversation starters generated by the GLM 110. Clicking on one of the suggestion pills enters conversation mode with the selected user input as the user's turn, and the chatbot's response is displayed.
[0049] Suggestion pills represent suggestions that are creative and helpful in continuing the conversation. Such suggestions typically relate to ideas and concepts that immediately follow the current concept. Therefore, they are highly personalized using suggestions the user has previously consumed. Unlike PAA or RS, conversation enablers are somewhat informal and retain context over longer time steps. They can also include anaphoric instructions for the current input. For example, if the input is "When is the best time to go to the Grand Canyon?", the pill could be "How do I get to the Grand Canyon information center?" Current technology demonstrates the following types of suggestions: query auto-completion, questions about what others also asked, and related searches. The primary goal is to help users quickly find what they are looking for. The goal of a conversation enabler is to engage users in longer sessions by motivating them to explore further by suggesting highly interesting recommendations that engage them in a conversation with the GLM, thereby learning more in the process. In one embodiment, conversation enablers can be longer suggestions compared to the typically short current suggestions. In a sense, the idea is to generate idea paths for various concepts. Based on the user's current state on the idea path, suggest the next idea on that path. For example, a user may be searching and reading web pages about ideas such as "artificial neuron" and "multilayer perceptron". Related to this deep learning concept, it may be desirable to show the user ideas such as "convolutional neural network" or "recurrent neural network". The next pill is generated by retaining previously suggested pills, pills consumed, and previous user queries. Thus, the pill generator relies on richer context. Conversation enabler pills should not be bland answers such as "sounds good" or "seems okay".
[0050] The SERP conversation is based on information in the poll 426 (e.g., story, creative content, semantic SERP, etc.). To demonstrate the inferential power of the GLM 112, a "peek" of the semantic SERP can be provided for certain inputs. For example, a user can be shown a peek (small set) of the semantic SERP 406. This is a highly visual, story-driven experience in which classic SERP data is inferred, reorganized, and streamed into a visual, story-driven experience, along with summaries 410 that entice the user to explore analytics and links in each category. The semantic SERP "streams" word by word, drawing the user's attention to anything active happening on the page. In one embodiment, the semantic SERP can remain a constant size over the page load time without user input to reduce page clutter.
[0051] The system generates content that streams in visually character by character, token by token, word by word, or segment by segment. The backend of the system can segment chunks to speak or display while streaming. This feature reduces TTS latency. Results can be streamed in at the rate of probability as they arrive. In another embodiment, controlled generation is performed using a buffer of content. The system can return text word by word or speak the response word by word.
[0052] Streaming output from one model can be fed into another model, for example, text produced by model my can be streamed into a Responsible AI (RAI) component, then into a speech generation model, etc.
[0053] When a bot is reading a story or generating content, there can be dynamic panes that update along with the stream. For example, if a bot is reading a story, there can be a side pane that displays images generated by a generative model related to the current paragraph being streamed. Similarly, an advertising pane can update ads based on what is currently being streamed in another stream (e.g. a summary of the main topic). Multiple streams can be synchronized in time when they are rendered.
[0054] SERPs do not have to be one-shot GET requests, but contain dynamic content that is continually fetched and updated by the GLM. The connection to the server can be maintained until the user navigates away completely. Images / ads can be rotated as they are generated. Streaming supports left-to-right and right-to-left languages.
[0055] Speculative pre-fetching of the next input given a user conversation is supported, e.g., "I thought you'd ask that." Pre-computed / pre-aggregated results are also provided in a semantically consistent cache, allowing for retroactive fact-checking, for example, to inform the user if the bot's opinion has changed due to new information.
[0056] Other features include: streaming results for smart find on the current page content, streaming results for long documents displayed in a web browser with a side pane for sending chunks of the current page to the model, a Map Reduce paradigm using LLM, iteratively refining a summary answer to the user as results arrive, different parts of the page that match the query criteria being highlighted as the answer streams in, summaries for different parts of the document being displayed can be streamed in, and the system can send parallel requests to summarize different parts of a document and stream them all back in parallel.
[0057] Further features include popup text and hidden content that can be generated by the GLM on demand in subsequent passes as needed. Interaction with certain parts of the page can trigger additional streaming requests, and dedicated tokens can be injected into the streaming response for a single prompt to instruct the model to stop processing the last section / request and start streaming tokens for a new section of the larger response. For example, a response may have several search results and their summaries, so the first pass may only generate 100 tokens for each of the 10 summaries. When the user hovers (or clicks) on Summary 2, the generation can update with ACTION-HOVER-2 and begin generating more tokens for Summary 2. This resembles a UI metadata interaction with the page (the streaming model's results) behind the scenes for the user. This metadata can include tokens or soft tokens (embedding, e.g., changes in the user's context) that are injected as the streaming progresses. Other actions can include the user clicking on certain other parts of the page, which helps refine the user's intent as the bot generates text / voice.
[0058] Another embodiment provides dynamic, real-time response modification by the GLM 112 based on user feedback. As results are streamed back via text or voice, the user can interject questions like "Tell me more about X" or "More about Y" to direct subsequent responses in real time. The user can communicate their latest intent to the system using text, voice, or other methods (eye tracking, user nods or shakes of the head, visual tracking of the user's emotions (e.g., smiles, frowns, etc.), EEG, or user actions within Windows / Edge / Page (e.g., user asks for help with something, bot starts explaining how to do it, user understands, bot says "Ah, looks like you understand," etc.). This feature saves further computation of previous responses.
[0059] An example of the dialogue is: Bott: "There are many breeds of cattle, raised for a variety of purposes, including dairy, beef, and draft cattle. Some of the main breeds of cattle include: 1. Holstein: This breed is known for its high milk production and is the most common breed of dairy cattle. 2Jersey: This breed is known for the high milk fat content, 3 and is a popular breed for dairy farming. Angus: This breed is known for... User: "Tell me more about Holsteins."
[0060] Voice and text models have an adaptive way of gracefully shortening speech or text so that new requests can be processed, as opposed to abruptly stopping. While the model searches and composes a new response, smaller, faster models / prompts can be used to generate quick filler responses such as "Holstein? Yes, what a great breed..." ... "According to Wikipedia, Holsteins are typically black and white..." Additionally, bot responses can be split into multiple smaller bubbles to reduce TTS latency.
[0061] Other features include the ability to paraphrase for responsible artificial intelligence (RAI). If RAI is triggered, the same adaptive response can be used while the system regenerates a more desirable response. The system may also stream backwards, i.e., use the backspace key, if it detects the need to restate something it has already displayed. This can be difficult with voice, and in such cases the system can say, "Oops, I meant to say .... Here's a better answer: ..." or some other phrase to correct the error. The client-side audio buffer can be large enough so that these corrections can occur before being heard / acknowledged by the user. The bot can decide not to answer instead of restating, or answer with a more generic / safe response that can be predetermined.
[0062] With adaptive aggregation, if the GLM 112 determines it has enough evidence to support a fact, it can stop making background search requests to obtain more information. With database aggregation, or streaming statistics in general, once the model has enough evidence sampled from an IID distribution, it can state an answer with a certain confidence limit. Once a certain bound is reached, the model can stop processing. Similarly, an active information agent can stop once it has convincing evidence about the answer to a particular subproblem.
[0063] Asynchronous chatbot requests are provided for dynamically updating the ranked list of results, with answers coming back at a later point in time, driven by, for example, a list of conversations or tasks and their progress. In another example, the system can stop processing requests that the user has scrolled out of view. The system can speculatively begin streaming results outside the current viewport if the user is about to scroll.
[0064] Another feature uses eye tracking to determine where to focus the production budget. For example, the system determines answers to questions such as: does this paragraph need more text tokens? What production rate is best? How fast can the user read / absorb? Can appropriate speeds, pauses, etc. increase information retention?
[0065] In another embodiment, input is streamed to the GLM 112 as the user types. The model already knows what the user wants and can provide an answer more quickly. The model can also autocomplete the user's sentences (autocomplete the question and show a tentative answer if the user accepts the completion). Example: "How tall is Barack Obama? Answer: Barack Obama is 6'1". He served as the 44th President of the United States..."
[0066] The bot can actively stream summaries of k real-time streams (meetings, video, audio, chat, other event streams (sporting events, TV shows, etc.)) in less than real time. Thus, the user can participate in the activities of k simultaneous meetings. A dashboard of live semantic summaries can also be provided. The summaries can include text and various visual elements. The summaries can include condensed clips of the most interesting recent moments from a basketball game. This functionality can enable a user to watch 10 games at once and not miss anything interesting, or to serve 10 customers simultaneously with text-based chat support.
[0067] A model can review its output as it is being generated and initiate additional background processing, such as RAI (Responsible AI Classifier), fact-checking processors, etc. This can be done token-by-token, at sentence boundaries, etc. This functionality applies to multimodal generation (text, images, video, audio, etc.). A buffer is also provided to call downstream models with the output and get feedback from the downstream models before the user sees it. If the RAI detects an issue, it can dynamically send that feedback to the streaming output / input of the first model so that the response can change to be more desirable. There can be many simultaneous generation processes incorporating these near-real-time feedback signals.
[0068] There can also be complex workflows of streaming data entering models. Ads generated by one expensive model can be streamed into a second prompt running on an expensive model that is also streaming. A page can visually display multiple streams from different models and model invocations. As above, some of these streams can be internal and only recognized by other models.
[0069] Pipelining multiple prompts is like chaining n prompts together into t * It allows rendering in t time instead of n time, where t is the full generation time. At the hardware and system level, it allows multiple continuation queries to communicate with each other. Given two prompts a, b, where b depends on the output of a, both prompts can start loading in parallel. As output from a arrives, it can be appended / loaded into b, and once enough has been received, b can start generating.
[0070] There can be hot spots where certain experts in a Mixture of Experts (MoE) model become overloaded. To improve throughput, parallel copies of experts can be created. Furthermore, distilled models can be used that are specialized for their function while requiring less computational resources (and possibly time) compared to a general-purpose GLM. For example, a distilled GLM can be configured to perform text summarization (but not generate queries). Another distilled GLM can be configured to generate queries (but not perform text summarization).
[0071] For iteratively generated content, such as diffusion models, intermediate steps of the computation can be streamed: for example, while streaming images, a single image may start out as white noise but sharpen into a picture of a cat as the model iterates.
[0072] Outpainting techniques can also be streamed in. As the user pans the image, the model can begin to outpaint the missing areas, generating new details that match the original image. Zooming an image can also be streamed in, with the zoomed image being streamed in.
[0073] Generated text for writing assistance can also be streamed in. Streaming can occur within a sentence or between existing paragraphs, etc. This is an example of text "inpainting," where the user has areas they want to fill in with a model. The user can initiate multiple simultaneous requests, such as "add a paragraph about bulldog health issues after the paragraph about dog breed costs," or "add an introductory paragraph saying this is a buying guide for dog breeds." These can be streamed in simultaneously.
[0074] There can be many UI elements that indicate that the bot or page is working on something: "..." or "I think..." or "Searching for x" or "Checking prices" or "Finding the best deal." This description of what the system is doing is generated by a model and can be conditional on the query / task.
[0075] There can be more than two participants in a conversation, and possibly multiple bots. This feature can handle situations where there are long prompts and long responses; instead of timing out and discarding all partially generated tokens, the system can use the partial output and get the job done.
[0076] In another embodiment, a "home page" SERP is provided, providing a TV newscast-like experience where users can interact with the news anchor. For example, the news anchor can be represented as an avatar dynamically generated in near real time. When a user asks about the weather, the transition between queries can be animated, such as "Switching to weather forecast." The system can pose questions to the user to help guide the interactive broadcast. In this regard, the system functions as a virtual operating system, where the user can ask for various information to be displayed, summarized, etc. Some jobs may take a long time to process, in which case the system can switch topics, such as "Now, back to your question about the cow..." Streams can contain metadata for synchronization with other streams, such as tone of voice, facial expression, and style, which can be input to other widgets, models, avatars, or renderers that process or combine one or more streams. For example, there may be a cue to flip to a certain advertisement when a particular sentence is being rendered. A page can also have multiple open connections for pulling different streams. Some streaming may occur on the client, and some stream processing may occur on the server.
[0077] Referring now to FIG. 5 , a GUI 500 is shown on a communication device 502, such as a tablet, mobile phone, or smartphone, according to one or more aspects described herein. The GUI 500 includes a SERP interface 504, which includes a query field 506, an instant answer field 508 (if applicable), an entity description field 510 (if applicable), and a results field 512. Displayed within the results field 512 are one or more results (labeled A-D) returned in response to the query entered in the query field 506, and optionally one or more images 514. A user clicks on one of the returned results A-D, and the device displays received information related to the selected result. Additionally, the system retrieves supplemental content 516 (e.g., additional articles, hyperlinks, images, advertisements, etc.) related to the selected result and displays the supplemental content on the SERP interface 504.
[0078] It will be understood that the particular order of the query field 506, instant answer 508, entity description 510, results field 512, and supplemental content field 514 is not limited to that depicted in Figure 5, but rather, the elements may be arranged in any order. Furthermore, the elements depicted in the SERP interface 504 are not limited to the stacked arrangement shown in Figure 5, but rather may be arranged side-by-side, in a grid arrangement, etc.
[0079] The communication device 502 further includes a microphone 518 and one or more speakers 520 through which a user can input voice commands and receive audio from the communication device. For example, a user can initiate a query by saying the word "query" or "question" to activate the microphone, followed by a word or phrase that the user might otherwise manually enter into the query field 506. The results field 512 can be populated with results (e.g., hyperlinks, article titles, images 512, etc.) in response to the user's voice query. In another embodiment, the results can be read out as audio output by the speaker 520 and presented to the user.
[0080] In another embodiment, a voice-activated graphical icon (not shown) may be provided in the query field 506 or elsewhere in the SERP interface 504. When the user selects (e.g., taps or presses and holds) the voice-activated graphical icon, the user is prompted to begin speaking and may speak a natural language query into a microphone 518. One or more of the returned instant answers 508, entity descriptions 510, results 512, and / or supplemental content 514 may be presented to the user as audio output by one or more speakers 520.
[0081] The GUI further includes a conversation field 522 showing the conversation between the user and the GLM and / or search engine. Contained within the conversation field 522 or the user query and respective natural language responses, as well as one or more chat images 524 (i.e., images contained within the chat or conversation field).
[0082] In another embodiment, the surf 504 displayed on the GY 500 is a semantic SERP, such as the semantic SERP 406 described with respect to FIG.
[0083] 6, shown on a GUI 500 of a communication device 502 is displayed information 604, as well as a web page 602 including text, images, selectable graphical icons, selectable text, etc. Also shown is supplemental content 606 and a semantic SERP 608. The semantic SERP 608 may be similar to or identical to the semantic SERP 406 described with respect to FIG. 4. The supplemental content 606 may be obtained based on a user's expressed interest in the information provided in the semantic SERP.
[0084] Users are also permitted to engage in text or voice conversations that are displayed in the conversation field 610 of the semantic SERP. User queries and natural language responses generated by the system are displayed to the user as a dialogue in the conversation field 604. An example of a conversational dialogue that can be displayed in the conversation field 610 (or any of the conversation fields in the previous figures) is shown below:
[0085] Query 1: What state is Ann Arbor in? Response 1: Ann Arbor is located in the US state of Michigan. Query 2: Can you please tell me more? Response 2: Ann Arbor is a city in southeastern Michigan, located approximately 35 miles (56 km) west of Detroit. It is the county seat of Washtenaw County and is known as the home of the University of Michigan, one of the oldest and most prestigious public universities in the United States. Query 3: What SAT score does this university require? Response 3: The University of Michigan requires students to submit SAT scores when applying. The middle 50% of the Class of 2025 had an SAT score between 1340 and 1470.
[0086] As can be seen, the responses generated by the system take into account the context of the conversation. For example, when the user mentions "university" in Query 3, the system infers that the user is referring to the University of Michigan based on the context of Response 2. Communication device 502 also includes a microphone 518 and one or more speakers 520, which allow the user to speak queries and hear responses during a conversation, as described above with respect to FIG. 5.
[0087] The supplemental content 610 is obtained using the context of the conversation and may include additional links, images, selectable graphical icons, etc. that the user can click on to obtain additional information. For example, the content may include, without limitation, links to one or more hotels in the Ann Arbor area, links to restaurants in Ann Arbor, links to purchase tickets to University of Michigan sporting events, etc.
[0088] 7 illustrates a methodology related to providing streaming information within a SERP, according to one or more embodiments described herein. While the methodology is depicted and described as a series of acts performed in sequence, it should be understood and appreciated that the methodology is not limited by the order. For example, some acts may occur in a different order than described herein. Additionally, some acts may occur simultaneously with other acts. Furthermore, in some instances, not all acts may be required to implement the methodology described herein.
[0089] Furthermore, the acts described herein may be computer-executable instructions that may be implemented by one or more processors and / or stored on a computer-readable medium. Computer-executable instructions may include routines, subroutines, programs, threads of execution, etc. Furthermore, the results of the acts of the methodologies may be stored in a computer-readable medium, displayed on a display device, etc.
[0090] 7, a flow diagram illustrates a method 700 for providing a streaming experience in a SERP in a computing system, according to one or more aspects described herein. At 702, user input is received at a search system. At 704, search results are generated for a user query. At 706, the query and generated search results are sent to a GLM. At 708, narrative results from the GLM based on the query and provided search results are received. At 710, the search results and narrative results from the GLM are streamed to a client device for presentation in the SERP.
[0091] At 712, an indication of user interaction with the SERP is received. The user interaction may be a user click on a selectable icon, word, or phrase presented within the SERP, or a user voice command such as "tell me more about that" during a particular portion of the streaming data when it is first streamed into the SERP. At 714, additional search results are obtained based on the user's interests determined by the indication of user interaction with the SERP. At 716, a new query is generated for the GLM based on the user interaction. At 718, an updated GLM response narrative is received. At 720, the updated search results and the updated GLM response narrative are streamed to the SERP on the client device.
[0092] With continued reference to Figures 1-7, various additional contemplated features and aspects are described below. In one embodiment, the streaming HTML can include multimedia (MM) content. The MM content may itself be generated by the GLM and may stream in as it is generated, starting with, for example, a white noise image and progressing to a detailed image. The SERP may be an entire page with many sections, and / or multiple elements may be integrated into the interactive response as described above. Some initial content displayed may be generated, but as the user scrolls down or otherwise interacts with the UI, additional content may continue to stream in and be displayed.
[0093] Results from multiple paragraphs can be streamed in in parallel. When the user scrolls a paragraph out of view, the streaming can be stopped to conserve GPU resources. The user can continue to quickly skim the first sentence of multiple results. Priority signals for various streams can be dynamically sent to the model to prioritize GPU capacity to more important streams. Certain streams can be paused, deprioritized, or canceled.
[0094] The user can switch between different streams using certain widgets, e.g. "next page" for text and image results. The system can also do automatic switching so that the user can preview the first row of each page. The system caches results for faster search and retrieval. There can be an animated icon / avatar that the system is "thinking" while retrieving results. Subtask progress can be streamed in to show the user the progress. Certain parts of the page can be loaded quickly while other parts are streamed in. A whole page optimizer controls the display of streaming results to avoid a disjointed UX of constantly moving results for the user.
[0095] You can use a classifier to decide which model to use. Some responses may not be worth using an expensive model, and a low latency model will suffice.
[0096] In another embodiment, a workflow engine (not shown) is provided to parallelize tasks, such as Map Reduce, and a job optimizer (not shown) that determines what tasks or operations can be performed in parallel.
[0097] Further possibilities include prompts that can generate extensions, reach checkpoints, insert newly arrived streaming content, continue generating, etc. This functionality allows inputs to a prompt template to arrive out of order.
[0098] The described systems and methods can also provide adaptive refinement from fixed models, such as iteratively generated images, frame-by-frame generated video, music generated on the fly, voice generated on the fly, Excel-like tables where each aggregate cell is asynchronously populated, dependency graphs between cells from which fixed cells can be extracted, summarized, etc., Map-Reduce dependency graphs, workflow engines, etc.
[0099] In another example, the model itself can be used to plan / optimize / schedule a workflow, including executing tasks in parallel and asynchronously, and to determine which sub-prompts to invoke. In another example, the system can generate voice conversations using k different voices. This capability facilitates providing users with certain experiences, such as hosting simulated dueling news anchors or podcast interviews, providing two points of view read by different bot speakers, narrating a story with k different characters speaking differently, or providing multimedia video content with multiple speakers. This capability also facilitates providing adaptive capabilities, such as Bot 2 interrupting Bot 1 mid-sentence with a fully context-sensitive interjection.
[0100] The system may also allow the voice to speak only a brief summary of the content (which may not be visible on-screen) while making the displayed text more verbose / detailed, resulting in a mixed mode that includes, for example, text (full page) + voice summary of the content.
[0101] In another example, the system allows results to stream in a sidebar as the user navigates, facilitating zero-shot suggestions, commentary, recommendations, etc.
[0102] The system provides the ability to simultaneously attend k virtual meetings and can use summarization to abbreviate real-time content (as text or voice summaries) so that users can attend multiple meetings in parallel. This functionality can be extended to watching k TV channels, k sporting events, and seamlessly stitch together narration when switching between live commentary of different event streams.
[0103] According to another aspect, the system can synchronize the current streaming content from the model with the next concurrent command issued by the user. For example, if a user begins talking about something that is still streaming, when the GLM 112 generates the next response, it can use the synchronization information to infer what "it" meant, i.e., what the user was watching in the stream at the time the user said "it." This functionality can be applied to any streaming content, such as text, video, etc. In one embodiment, this functionality is provided by the synchronization module 208 (FIG. 2).
[0104] In another embodiment, the system provides adaptive streaming. For example, a user hovers the cursor over an image. The summary narration can capture the event and adapt the subsequent near real-time commentary of the page to focus on what the user is attending to, like a salesperson or teacher. The user can attend to things by any human-computer interface method.
[0105] Certain features also apply to augmented reality commentary. SERPs can be overlaid within VR or AR. There can be a continuous stream of "queries" so the model is constantly accumulating context and adapting its generated content / commentary. The contextual Edge sidebar can also be extended to include VR / AR content.
[0106] In another embodiment, the system provides streaming "summaries" of event streams, such as updating log plots of financial tickers. The model can perform streaming log calculations on the raw input stream. Running tokens generated from prompts can be checkpointed to generate a histogram, with the last non-full bucket representing the most recent data point. When the bucket window passes, the most recent data point is aggregated and the generation / prompt is set back. In this respect, the operation is similar to a read / write prompt. For example, [1,2,3,4],5,6⇒[1,2,3,4],[5,6,7,8],9,10⇒etc. Checkpoints can be stored on the prompt / generated sequence.
[0107] In another example, when listening to k conferences, a user can select which streams they want to actively listen to via voice or text and which streams they don't want to attend. The combined summary stream can give the user more details about the subset that most interests them and occasionally give them a high-level summary of what's happening in the other streams. This functionality can also be applied to asynchronous information requests the system is processing. For example, if a user asks the system to look up k things and the system generates k' subtasks, the system can give the user a summary of the k things it's doing. That is, the system can tell the user how things are going, what it plans to do next, etc. The user can guide the system in where to direct its attention and search / processing.
[0108] When generating a response, the system knows how long the user will have to look at the response (e.g., if the user is busy, if they are driving, what device the user is using, etc.) The system can generate responses of appropriate length, not just high-level summaries, ask the user where they would prefer further expansion, etc. Where the system anticipates the user may ask for more detail, the system can already start pre-generating responses.
[0109] The system can provide continuous commentary for a basketball game or soccer game. The system can listen to an AM radio station, convert speech to text, and then generate a more entertaining, stylized, abbreviated commentary, etc. The system can listen to k different stations talking about the same event and integrate the commentary. The system can continuously inform the user about the stock market, stocks the user is interested in, various industries, any other news articles the user is interested in, or any other world news or events that are happening, etc. The system can create a personalized Bloomberg TV station for the user. The system can generate a dynamic mix of information types of answers, such as charts, text, audio, etc. The system can dynamically watch / mix video content being streamed on k TV or media sources.
[0110] In another example, the user may be given the option to provide feedback regarding the quality or usefulness of the search results and / or chat interaction narrative, such as a "thumbs up" or "thumbs down" response, a rating (e.g., on a scale of 1 to 5), or the like.
[0111] In another embodiment, a user can submit a URL or similar as a query and ask the system to summarize its content. The system then parses the content of the URL and returns a streaming summary of the content to the user. Similarly, a user can initiate a chat function while reading an email and ask the system to summarize the content of the email or its associated document.
[0112] Other features include voice-driven search, visual question and answering for visual content, bot-labeled descriptions for objects on the page (icons plus descriptions), voice descriptions for user interactions with page elements, adaptive generation based on user interaction and attention, dialogue-driven interaction with content on SERP pages, wrapper or right-rail overlay, and interaction with other UIs for other content (email search, SharePoint search, etc.).
[0113] Additional features may include dialogue-driven interaction with dynamically generated web pages, dynamic layout rearrangement, dynamic editing of elements in the conversation history, full-page transitions from dialogue elements in a conversation, weather elements in a conversation, the ability to click on an answer card to return to a full-page portal / details / non-mini version of the element, the ability to switch between elements in a conversation and the full experience mode, etc. For example, a user may switch from a shopping answer to a full-page shopping page. Other functionality includes answer / explode views, the ability to place answers aside from the conversation, the ability to pin, the ability to place in a new browser tab, etc. Additionally, an "Explode" button may be provided for answers that have alternate expansion modes to switch between.
[0114] Additionally, options are provided for multi-page conversations, e.g. a given model state can contain dialogue / interactions across multiple tabs / pages / searches.
[0115] Other features include the ability to display the content of an entire news article as an answer that is added to the dialogue context. For example, the system can fetch a news article but display only the headline and image of the news article. In this scenario, the user can ask a next question that leverages the full content of the news article and / or prioritizes fetching the content to generate the next response.
[0116] In another embodiment, search results in conversation mode may be delivered as web result answer cards and / or semantic summary answer cards.
[0117] The described systems and methods also provide the ability to share conversations with others, enable multi-party conversations, save conversations to resume later, bookmark conversations, save the complete dialogue history to review later, timestamp dialogue conversations so that the next response can leverage recent conversation history across multiple windows / conversations, share conversation turns broadly (e.g., on social media), integrate mixed-mode external content / dialogue into enterprise chat applications (e.g., Teams, Skype), provide upsells to the Sapphire app experience, and more.
[0118] In another example, the described systems and methods facilitate providing a "new tab page" that includes one or more of a "what's new" summary, asynchronous updates about what's currently happening with user data, user interests, what's happening in the world, etc. Email can be another experience where mini-answer / expanded / full-page modes are provided. In email, sub-answers can be individual emails, information about people, people's answers / cards, etc. The "expanded" mode can launch a new window or tab for composing / sending an email reply. The answer can include a short list of related emails or SharePoint items. The expanded mode can also transition to a full-page Word document for document results. An option to return to the conversation in mini-answer mode is also contemplated.
[0119] Referring now to FIG. 8 , a high-level diagram of an exemplary computing device 800 usable in accordance with the systems and methodologies disclosed herein is shown. For example, computing device 800 may be a client computing device on which an operating system is stored, the operating system providing streaming functionality for presenting SERPs to a user on the client device. As another example, computing device 800 may be a server computing system providing streaming search presentation functionality. Computing device 800 includes at least one processor 802 that executes instructions stored in memory 804. The instructions may be, for example, instructions for implementing functionality described as being performed by one or more components discussed above or instructions for implementing one or more of the methods described above. Processor 802 can access memory 804 via system bus 806. In addition to storing executable instructions, memory 804 may also store content, graphical icons, profile information, and the like.
[0120] Computing device 800 additionally includes a data store 808 accessible by processor 802 via system bus 806. Data store 808 may include executable instructions, graphical icons, profile information, content, etc. Computing device 800 also includes an input interface 810 that allows external devices to communicate with computing device 800. For example, input interface 810 may be used to receive instructions from an external computing device, such as from a user. Computing device 800 also includes an output interface 812 that interfaces computing device 800 with one or more external devices. For example, computing device 800 may display text, images, etc. via output interface 812.
[0121] It is contemplated that external devices communicating with computing device 800 via input interface 810 and output interface 812 may be included within an environment that provides virtually any type of user interface with which a user can interact. Examples of types of user interfaces include graphical user interfaces, natural user interfaces, etc. For example, a graphical user interface may accept input from a user using an input device such as a keyboard, mouse, remote control, etc., and provide output on an output device such as a display. Furthermore, a natural user interface may allow a user to interact with computing device 800 in a manner free from the constraints imposed by input devices such as a keyboard, mouse, remote control, etc. Rather, a natural user interface may rely on speech recognition, touch and stylus recognition on as well as adjacent to the screen, gesture recognition, air gestures, head and eye tracking, voice and speech, vision, touch, gestures, machine intelligence, etc.
[0122] Additionally, although illustrated as a single system, it should be understood that computing device 800 may be a distributed system, such that several devices may communicate over network connections and collectively perform the tasks described as being performed by computing device 800.
[0123] The various functions described herein may be implemented by hardware, software, or any combination thereof. If implemented by software, the functions may be stored on or transmitted over as one or more instructions or code on a computer-readable medium. A computer-readable medium includes a computer-readable storage medium. A computer-readable storage medium may be any available storage medium accessible by a computer. By way of example and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where a disk typically reproduces data magnetically and a disc typically reproduces data optically with a laser. Furthermore, propagating signals are not included within the scope of computer-readable storage media. Computer-readable media also includes communication media, including any medium that facilitates transfer of a computer program from one place to another. For example, a connection may be a communication medium. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of communication media. Combinations of the above may also be included within the scope of computer-readable media.
[0124] Alternatively, or in addition, the functions described herein may be performed at least in part by one or more hardware logic components. For example and without limitation, exemplary types of hardware logic components that may be used include Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on a Chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0125] Various techniques are described herein, with at least the following examples:
[0126] (A1) In one aspect, a computing system is described herein. The computing system includes a processor and a memory storing instructions that, when executed by the processor, cause the processor to perform actions. The actions include receiving a query provided by a user and generating a prompt that is input to a generative language model, the prompt including instructions for the generative language model to generate a conversational output based on the query. The actions further include providing the prompt as input to the generative language model, the generative language model generating the conversational output based on the prompt. The actions also include streaming the conversational output to one of a search engine results page (SERP) or a web page to which the user navigates from the SERP, such that the conversational output appears to the user as if it were being typed in real time.
[0127] (A2) In some embodiments of the computing system of (A1), generating the prompt further includes providing a query to a search system, wherein the search system identifies search results based on the query and includes a portion of the search results in the prompt.
[0128] (A3) In some embodiments of the computing system of (A2), the acts further include receiving an indication of user interaction with the SERP in response to the conversational output. The acts also include obtaining additional search results in response to the user interaction. The acts further include generating and transmitting an updated prompt to the generative language model based in part on the additional search results and the user interaction. The acts further include receiving the updated conversational output from the generative language model. Additionally, the acts include streaming the updated search results and the updated conversational output to the SERP on the computing device.
[0129] (A4) In some embodiments of the computing system of at least one of (A1)-(A3), the acts further include providing supplemental content related to the interactive output for display on a dynamically updatable display pane on the computing device, and updating the supplemental content in response to at least one of additional user interactions and the updated interactive output.
[0130] (A5) In some embodiments of the computing system of (A4), the act further includes synchronizing streaming of the supplemental content with the interactive output by comparing metadata associated with the interactive output and the supplemental content.
[0131] (A6) In some embodiments of the computing system of at least one of (A1)-(A5), the actions further include identifying objectionable content in at least one of the interactive output and the supplemental content. The actions also include at least one of removing the objectionable content and removing a source of the objectionable content before rendering.
[0132] (A7) In some embodiments of the computing system of at least one of (A1)-(A6), the act further includes caching the response of the generative language model for retrieval in response to similar queries.
[0133] (B1) In another aspect, a computing system is described herein. The computing system includes a processor and a memory storing instructions that, when executed by the processor, cause the processor to perform actions. The actions include transmitting a user query to a search system. The actions further include receiving a conversational output from a generative language model, the generative language model generating the conversational output based on a prompt including the query. The actions also include streaming the conversational output to a search engine results page (SERP) or one of the web pages navigated to by the user from the SERP, such that the conversational output is presented to the user as if the conversational output were being typed in real time.
[0134] (B2) In some embodiments of the computing device of (B1), the interactive output includes information related to search results responsive to the query.
[0135] (B3) In some embodiments of at least one of the computing devices of (B1)-(B2), the acts further include transmitting an indication of user interaction with the SERP in response to the interactive output. The acts also include receiving updated search results and updated interactive output in response to the user interaction.
[0136] (B4) In some embodiments of the computing device of (B3), the act further includes displaying the streamed updated search results and the updated conversational output on the SERP.
[0137] (B5) In some embodiments of at least one of the computing devices of (B1) through (B4), the acts further include receiving and displaying supplemental content related to the interactive output on a dynamically updatable display pane on the SERP. The acts also include receiving and displaying updated supplemental content in response to at least one of additional user interactions and the updated interactive output.
[0138] (B6) In some embodiments of the computing device of (B5), the act further includes synchronizing streaming of the supplemental content with the interactive output by comparing metadata associated with the interactive output and the supplemental content.
[0139] (B7) In some embodiments of at least one of the computing devices of (B1) through (B6), the acts further include identifying objectionable content in at least one of the interactive output and the supplemental content, and additionally, the acts also include at least one of removing the objectionable content and removing a source of the objectionable content prior to rendering.
[0140] (C1) In another aspect, a method performed by a computing system is described herein. The method facilitates providing conversational interaction on a search engine results page presented on an interface of a computing device. The method includes receiving a query provided by a user. The method further includes generating a prompt that is input to a generative language model, the prompt including instructions for the generative language model to generate a conversational output based on the query. The method also includes providing the prompt as input to the generative language model, the generative language model generating the conversational output based on the prompt. Additionally, the method includes streaming the conversational output to a search engine results page (SERP) or one of the web pages navigated to by the user from the SERP, such that the conversational output appears to the user as if it were being typed in real time.
[0141] (C2) In some embodiments of the method of (C1), generating the prompt further includes providing a query to a search system, wherein the search system identifies search results based on the query and includes a portion of the search results in the prompt.
[0142] (C3) In some embodiments of the method of (C2), the method further includes receiving an indication of user interaction with the SERP in response to the conversational output. The method also includes obtaining additional search results in response to the user interaction. The method further includes generating and transmitting an updated prompt to the generative language model based in part on the additional search results and the user interaction. The method further includes receiving the updated conversational output from the generative language model. Additionally, the method includes streaming the updated search results and the updated conversational output to the SERP on the computing device.
[0143] (C4) In some embodiments of at least one of the methods of (C1)-(C3), the method further includes providing supplemental content related to the interactive output for display on a dynamically updatable display pane on the computing device, the method further including updating the supplemental content in response to at least one of additional user interactions and the updated interactive output.
[0144] (C5) In some embodiments of the method of (C4), the method further includes synchronizing streaming of the supplemental content with the interactive output by comparing metadata associated with the interactive output and the supplemental content.
[0145] (C6) In some embodiments of at least one of the methods of (C1) through (C5), the method further includes identifying objectionable content in at least one of the interactive output and the supplemental content. The method also includes at least one of removing the objectionable content and removing a source of the objectionable content before rendering.
[0146] (D1) In another aspect, a method is described herein that is performed by a computing device, the method including any of the acts described in embodiments (B1) through (B7).
[0147] What has been described above includes examples of one or more embodiments. Of course, it is not possible to describe every conceivable modification and variation of the above-described apparatus or methodology in order to describe the above-described aspects, but those skilled in the art will recognize that many further modifications and permutations of the various aspects are possible. Accordingly, the described aspects are intended to encompass all such changes, modifications, and alterations that fall within the spirit and scope of the appended claims. Furthermore, to the extent the term "include" is used in the detailed description or the claims, such term is intended to be inclusive in the same manner as the term "comprising" is interpreted when used as a transitional term in a claim.
Claims
1. a processor; When executed by the processor, receiving a query provided by a user; generating a prompt to be input to a generative language model, the prompt including instructions for the generative language model to generate a conversational output based on the query; providing the prompt as an input to the generative language model, the generative language model generating a conversational output based on the prompt; streaming the conversational output to one of a search engine results page (SERP) or a web page navigated to by the user from the SERP so that the conversational output appears to the user as if it were being typed in real time; a memory storing instructions that cause the processor to perform actions including:
1. A computing system comprising:
2. The computing system of claim 1 , wherein generating the prompt further comprises providing the query to a search system, the search system identifying search results based on the query and including a portion of the search results in the prompt.
3. The act is: receiving an indication of user interaction with the SERP in response to the interactive output; obtaining additional search results in response to the user interaction; and generating and transmitting updated prompts to the generative language model based in part on the additional search results and the user interaction; receiving an updated conversational output from the generative language model; streaming the updated search results and the updated interactive output to the SERP on the computing device; and The computing system of claim 2 further comprising:
4. The act is: providing supplemental content related to the interactive output for display on a dynamically updatable display pane on the computing device; updating the supplemental content in response to at least one of additional user interactions and updated interactive output; The computing system according to claim 1 , further comprising:
5. The act is: synchronizing streaming of the supplemental content with the interactive output by comparing metadata associated with the interactive output and the supplemental content; The computing system of claim 4 further comprising:
6. The act is: identifying objectionable content in at least one of the interactive output and the supplemental content; and Before rendering, removing the objectionable content; and Remove the source of the objectionable content; At least one of The computing system according to claim 1 , further comprising:
7. The act is: Caching the generative language model response for retrieval in response to similar queries The computer system according to claim 1 , further comprising:
8. 1. A method for facilitating providing conversational interaction on a search engine results page presented on an interface of a computing device, comprising: receiving a query provided by a user; generating a prompt to be input to a generative language model, the prompt including instructions for the generative language model to generate a conversational output based on the query; providing the prompt as an input to the generative language model, the generative language model generating a conversational output based on the prompt; streaming the conversational output to one of a search engine results page (SERP) or a web page navigated to by the user from the SERP so that the conversational output appears to the user as if it were being typed in real time; A method comprising:
9. 10. The method of claim 8, wherein generating the prompt further comprises providing the query to a search system, the search system identifying search results based on the query and including a portion of the search results in the prompt.
10. receiving an indication of user interaction with the SERP in response to the interactive output; obtaining additional search results in response to the user interaction; and generating and transmitting updated prompts to the generative language model based in part on the additional search results and the user interaction; receiving an updated conversational output from the generative language model; streaming the updated search results and the updated interactive output to the SERP on the computing device; and The method of claim 9 further comprising:
11. providing supplemental content related to the interactive output for display on a dynamically updatable display pane on the computing device; updating the supplemental content in response to at least one of additional user interactions and updated interactive output; The method according to claim 8 , further comprising:
12. synchronizing streaming of the supplemental content with the interactive output by comparing metadata associated with the interactive output and the supplemental content; The method of claim 11 further comprising:
13. identifying objectionable content in at least one of the interactive output and the supplemental content; and Before rendering, removing the objectionable content; and Remove the source of the objectionable content; At least one of 13. The method according to claim 8, further comprising: