Search engine integrated with a large language model

EP4732162A1Pending Publication Date: 2026-04-29ELASTIC TECHNOLOGIES (US) INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
ELASTIC TECHNOLOGIES (US) INC
Filing Date
2024-06-24
Publication Date
2026-04-29

AI Technical Summary

Technical Problem

Predefined large language models (LLMs) are not trained on domain-specific organizational documents, leading to inaccurate responses to user queries and requiring computationally expensive re-training or fine-tuning to address domain-specific content, while also posing security challenges in accessing sensitive information.

Method used

A search system integrates a predefined LLM with a search engine that retrieves domain-specific documents from a private database, generating prompts with context windows including responsive documents to enable accurate and personalized responses, reducing computational complexity and maintaining security through authorized access.

Benefits of technology

Enables LLMs to provide accurate and personalized responses to domain-specific queries without re-training, while ensuring secure access to sensitive information by leveraging authorized user access rights, thus reducing computational resources and enhancing security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024035282_26122024_PF_FP_ABST
    Figure US2024035282_26122024_PF_FP_ABST
Patent Text Reader

Abstract

A search system may receive a search query for documents stored in a database associated with an organization. A search system may retrieve, from the database, search results (including responsive documents) that are responsive to the search query. A search system may initiate display of the search results on the user interface. A search system may transmit a prompt to a large language model, where the prompt includes the search query and information from at least a portion of the responsive documents. A search system may receive a model response with textual data that responds to the search query, where the textual data is generated by the large language model using at least the portion of the responsive documents. A search system may initiate display of the model response in a user interface.
Need to check novelty before this filing date? Find Prior Art

Description

SEARCH ENGINE INTEGRATED WITH A LARGELANGUAGE MODELPRIORITY STATEMENT

[0001] This application claims priority to U.S. Provisional Application No. 63 / 509,807, filed June 23, 2023, the contents in which are herein incorporated by reference in its entirety.BACKGROUND

[0002] A predefined large language model (LLM) is typically configured (e.g., trained) using a large collection of publicly available information such as information found on the Internet. However, some organizational documents with domain-specific content may not be publicly available and / or used to train the LLM, and. therefore, the LLM may not be able to generate accurate responses from user queries about that domain-specific content. For example, an organization may store and / or use a service to store their vast collection of documents, which may include a work-from-home policy document or an intellectual property rights policy document. A system may enable a user to search for organization documents, but the user may have to read several documents to find the right answer. Further, a predefined LLM may not be trained to correctly answer queries about those organizational documents, and re-training and / or fine-tuning an LLM with domain-specific content may be computationally expensive.SUMMARY

[0003] This disclosure relates to a search system that communicates with a large language model (LLM) to generate textual data (e.g., an answer, response, a summary, etc.) that responds to a query (e g., a search query, a user query’) using responsive documents retrieved by a search engine in a manner that reduces the computational complexity of integrating an LLM with a search system. For example, an organization may store a plurality of documents (e.g., organization documents) in a database (e.g., a private database) on a serv er computer. In some examples, the LLM is a predefined LLM that was configured (e g., trained) using publicly available information. In some examples, the LLM is not configured(e.g., trained) using at least some of the information found in the documents stored on the organization’s private database.

[0004] However, the search system discussed herein enables a predefined LLM to formulate answers to queries using organization documents that are responsive to a search query , where the organization documents include information that may not have been used to configure (e.g.. train) the LLM. The search system includes a search engine that retrieves a set of responsive documents that are responsive to a user’s search query and provides search results on a user interface. The search system also includes a prompt generator that generates a prompt with the search query and a context window that includes information about the set of responsive documents. The LLM uses the context window for formulating a model response with textual data that answers the search query from the domain-specific content found in the set of responsive documents.

[0005] In this manner, the search system integrates an LLM in a manner that is computationally less expensive than re-training and / or fine-tuning the LLM to answer queries about domain-specific content. Furthermore, the content window may include other information that assists the LLM to generate an accurate and / or personalized answer such as conversation history data and / or personalization data about the user that submitted the query'. Furthermore, the search system maintains or increases a level of security7by retrieving and / or providing responsive documents to the LLM in which a user is authorized to access (e.g., a non-manager user may not have access rights to some documents that a manager user has access rights to).

[0006] In some aspects, the techniques described herein relate to a method including: receiving, via an input field on a user interface of a computing device, a search query for documents stored in a database associated with an organization; retrieving, from the database, search results that are responsive to the search query, the search results including responsive documents; initiating display of the search results on the user interface; transmitting a prompt to a large language model, the prompt including the search query and information from at least a portion of the responsive documents; receiving, from the large language model, a model response with textual data that responds to the search query, the textual data being generated by the large language model using at least the portion of the responsive documents; and initiating display of the model response in a user interface object associated with the user interface.

[0007] In some aspects, the techniques described herein relate to a non-transitory computer-readable medium storing instructions that cause at least one processor to executeoperations, the operations including: receiving, via an input field on a user interface of a computing device, a search query for documents stored in a database associated with an organization; retrieving, from the database, search results that are responsive to the search query, the search results including responsive documents; initiating display of the search results on the user interface; transmitting a prompt to a large language model, the prompt including the search query and information from at least a portion of the responsive documents; receiving, from the large language model, a model response with textual data that is responsive to the search uery, the textual data being generated by the large language model using at least the portion of the responsive documents; and initiating display of the model response in a user interface object associated with the user interface.

[0008] In some aspects, the techniques described herein relate to an apparatus including: at least one processor; and a non-transitory computer-readable medium storing executable instructions that when executed by the at least one processor cause the at least one processor to execute operations, the operations including: receiving, via an input field on a user interface of a computing device, a search query for documents stored in a database associated with an organization; retrieving, from the database, search results that are responsive to the search query, the search results including responsive documents; initiating display of the search results on the user interface; transmitting a prompt to a large language model, the prompt including the search query and information from at least a portion of the responsive documents; receiving, from the large language model, a model response with textual data that is responsive to the search query, the textual data being generated by the large language model using at least the portion of the responsive documents; and initiating display of the model response in a user interface object associated with the user interface.

[0009] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] FIG. 1 A illustrates a search system that integrates a large language model (LLM) with a search system according to an aspect.

[0011] FIG. IB illustrates an example user interface of the search system that depicts a model response with textual data that responds to a search query' using information about responsive documents that were retrieved by a search engine according to an aspect.

[0012] FIG. 1C illustrates an aspect of the search system for generating a prompt for an LLM according to an aspect.

[0013] FIG. ID illustrates example search strategies for retrieving semantically similar documents according to an aspect.

[0014] FIG. IE illustrates an example of personalization data and a document access control mechanism that maintains or increases the security of a search system according to an aspect.

[0015] FIG. IF illustrates an example of a search system that generates multiple prompts to formulate an answer using domain-specific content according to an aspect.

[0016] FIG. 2 illustrates a flow diagram for integrating an LLM with a search system according to an aspect.

[0017] FIG. 3 illustrates a flow diagram for integrating an LLM with a search system according to another aspect.

[0018] FIG. 4A illustrates an example user interface for displaying search results and an LLM response according to an aspect.

[0019] FIG. 4B illustrates an example user interface for displaying search results and an LLM response according to another aspect.

[0020] FIG. 4C illustrates an example user interface for displaying search results and an LLM response according to another aspect.

[0021] FIG. 5 illustrates an example user interface for displaying search results and an LLM response that was generated using personalization data according to an aspect.

[0022] FIGS. 6A to 6D illustrate example user interfaces for displaying search results and an LLM response according to various aspects.

[0023] FIG. 7 illustrates an example user interface for displaying search results and a domain specific LLM response according to an aspect.

[0024] FIG. 8A illustrates an example user interface for displaying a personalized LLM response that accounts for a first organization role of a user according to an aspect.

[0025] FIG. 8B illustrates an example user interface for displaying a personalized LLM response that accounts for a second organization role of a user according to an aspect.

[0026] FIG. 9 is a flowchart depicting example operations of a system for generating a model response with responsive documents that are responsive to a search query7according to an aspect.DETAILED DESCRIPTION

[0027] The search system includes a large language model (LLM) integrated search service configured to operate with an application (e.g., a chat application) to enable retrieval of documents that satisfy a search query and to initiate generation of an answer (e.g., a model response with textual data) by an LLM that responds to the search query7. The LLM integrated search service includes a search engine that receives a search query and retrieves documents responsive to the search query. The documents that are responsive to a search query7may be referred to as responsive documents. Examples of the search query may be “Does the company own my side hustles” or “how does compensation work” or “What is NASA”, etc. The search engine may retrieve and initiate a display of search results on a user interface of a computing device, where the search results include responsive documents stored in a database associated with an organization. For example, the database may store documents that have been received from one or more client devices associated with the organization, and the search sy stem may update an index structure with the ingested documents. The search engine uses the index structure to search the database and locate responsive documents that satisfy the search query.

[0028] The LLM integrated search service includes a prompt generator that generates a prompt, where the prompt includes the search query7and a context w indow with information about the responsive documents. In some examples, the context window includes the content of the responsive documents themselves. For example, in response to a search query, the search engine may retrieve a first responsive document and a second responsive document. The context window7includes the contents of the first responsive document and the contents of the second responsive document. In some examples, the context w indow includes a summary of one or more responsive documents, where the summary may be obtained from the LLM in a previous prompt. In some examples, the prompt generator may select a subset of the responsive documents and include that subset in the context window7. In some examples, the search engine ranks the responsive documents, and the prompt generator may select the top N number of ranked responsive documents for inclusion in the prompt.

[0029] The prompt generator transmits the prompt to the LLM. and the LLM uses the search query and the set of responsive documents to formulate a model response w ith textual data that responds to the search query. In some examples, the LLM is stored in a server computer that is separate from the server computer(s) that host the LLM integrated search service. The prompt generator receives the model response and provides the model response in a user interface (UI) object on the user interface. In some examples, the UI object ispositioned on the user interface in a location that is separate from the display of the search results. In some examples, the UI object includes a chat interface that enables a user to submit additional queries (also referred to as user queries) and receive LLM responses (e.g., model responses) from the LLM about the documents stored in the database associated with the organization. The documents stored in the database may be referred to as documents with domain-specific content, e.g., content that may be specific to a specific organization and not necessarily used for configuring (e.g.. training) the LLM.

[0030] Therefore, the search system discussed herein enables a predefined LLM to formulate answers to queries using documents with domain-specific content that are responsive to a search query, where the documents stored in the database include information that may not have been used to configure (e.g., train) the LLM. In some examples, the documents stored in the database may be referred to as organization documents. In this manner, the search system may integrate an LLM in a search system that manages domainspecific content in a manner that can reduce the amount of computing resources (e.g., central processing unit (CPU) power, memory’) for integrating a LLM in a search system.

[0031] In some examples, the search system may retrieve personalization data about the user that submitted the search query, and the prompt generator may' include the personalization data in the prompt so that the LLM can personalize the model response. In some examples, the personalization data may include document access control data about one or more access permissions (or restrictions) associated with a user for accessing the documents. In some examples, the search system may retrieve responsive documents from the database in which the user has access rights to. A document, a group of documents, or a database (or dataset) may include a document security setting, and the search system may retrieve documents responsive to the search query that satisfies the document security setting. For example, a manager may be allowed to access documents that are not accessible by nonmanagers. As such, model responses generated by the LLM may' use documents with access rights that correspond to the user that submitted the search query7.

[0032] In some examples, the personalization data may include a location of the user, organization group(s) associated with the user, and / or an organization role of the user. For example, if the user is located in Canada and the user submits a query7about ‘'how do I update my tax elections”, the search system may retrieve documents related to that query7, and the model response may include information about updating tax elections in Canada (versus tax elections in the United States). Also, a model response may7depend on other personalization data. For example, in response to the query '‘how does compensation work”, the modelresponse may depend on whether the user is a manager, a non-managing engineer, or other organization role or group associated with the user.

[0033] In addition, the search system may include one or more structures, techniques, or mechanisms that reduce the amount of tokens used in the context window, which can further reduce the computational cost of processing LLM queries (e.g., prompts) by an LLM. In some examples, the search system may retrieve, in response to a search query, a set of semantically similar documents according to one or more search strategies. Using the search strategies discussed herein, the set of semantically similar documents may be highly relevant to the search query', where a lesser number of documents can be included in the prompt, thereby reducing the token size of the prompt (thereby reducing the computational cost of processing an LLM query). The search strategies may include a vector database search, a natural language processing (NLP) enrichment search, a late interaction model search, and / or a regular token matching search. In some examples, the search engine uses a hybrid search that uses a combination of two or more of the above search strategies. These and other features are further described with reference to the figures.

[0034] FIGS. 1 A through IF illustrate a search system 100 that communicates with a large language model (LLM) 170 to generate a model response 130 with textual data 132 that responds to a search query' 126 using responsive documents 104a with domain-specific content 106. For example, an organization may use the search system 100 to store a plurality of documents 104 in a database 102 on a server computer 160. The database 102 may be associated with an organization, where the database 102 stores documents 104 received from one or more computing devices 152 associated with the organization.

[0035] The search system 100 includes an LLM integrated search service 115 configured to operate with an application 166 (e.g., a chat application 168) to enable retrieval of documents 104 from a database 102 that satisfy a search query 126 and to initiate generation of an answer (e.g., a model response 130 with textual data 132) by an LLM 170 that responds to the search query 126.

[0036] The LLM integrated search service 115 may include an ingestion engine 116 configured to receive, index, and store documents 104 in the database 102. The documents 104 may include private documents, e.g., documents that are only accessible by authorized users of the database 102. The documents 104 may include public documents, e.g., documents that are accessible by the general public. The documents 104 may include private and public documents. In some examples, the documents 104 are organizational documents of an organization associated with the database 102. The LLM 170 may be a predefinedLLM that was configured (e.g., trained) using publicly available information. The LLM 170 may not be initially trained to answer questions about information found in the documents 104 stored on the database 102. However, the search system 100 discussed herein enables the LLM 170 to formulate model responses 130 to queries (e.g., search queries 126, user queries 127) about content included in responsive documents 104a, where the responsive documents 104a include information that may not have been used to configure (e.g., train) the LLM 170.

[0037] During data ingestion, the ingestion engine 11 may persist data to storage (e.g., the database 102), where the data may include documents 104 and / or index structures 155. The database 102 may be stored on the server computer 160. As shown in FIG. 1C, the database 102 may be a searchable data store that includes a plurality of contextual data stores such as a contextual data store 102-1, a contextual data store 102-2, and a contextual data store 102-3. The contextual data store 102-1, the contextual data store 102-2, and the contextual data store 102-3 may include internal company documents from an internal knowledge base, technical documents, issues, or code from a version control system such as GitHub, structured data from an external database, sales organization data from a customer relationship database, and / or proprietary documents from online storage system.

[0038] The ingestion engine 116 may receive documents 104 to be stored, managed, and searched by the search system 100. A document 104 may be an instance of digital data. The documents 104 may cover a wide variety of information such as files, text documents, web documents, web pages, PDFs, files and / or records. The documents 104 may be associated with a wide variety of file formats. The documents 104 may also cover images and / or video files. The ingestion engine 116 may include one or more indexing engines configured to index the documents 104 received via one or more computing devices 152. The ingestion engine 1 16 may include a distributed computing system with a plurality of nodes (e.g., which may also be referred to indexing nodes). The ingestion engine 116 may receive (e.g., ingest) data (e.g., documents 104) from one or more computing devices 152. The ingestion engine 116 may generate one or more index structures 155 about the documents 104.

[0039] An index structure 155 may be a data structure that includes information about the documents 104 that have been indexed. In some examples, the index structure 155 includes metadata (e.g., cluster metadata, metadata structure, file metadata, etc ). In some examples, an index structure 155 is referred to as an index, a Lucene index (e.g., Lucene files) or segments (e.g., Lucene segments) or a stateless compound commit file. In someexamples, an index structure 155 is referred to as an index file. In some examples, an index structure 155 is referred to as an index. The index structure 155 may be used by the search engine 118 to efficiently find the responsive documents 104a that are relevant to a particular search query 126. The type of index structure 155 is dependent upon the type of documents 104 ingested by the ingestion engine 116, but may generally include a document identifier, document type, timestamp, index terms, ranking, etc.

[0040] In some examples, the search engine 118 may include a distributed computing system with a plurality of search nodes (e.g., which may also be referred to as nodes). In response to a search query 126 from a computing device 152, the search engine 118 may search the index structure(s) 155 for responsive documents 104a that are responsive to the search query 126. The search engine 118 receives a search query 126 and retrieves responsive documents 104a that are responsive to the search query 126 from the database 102. The search query 126 includes one or more search terms that are used to locate responsive documents 104a among the plurality of documents 104 that are stored in the database 102. In some examples, the search query 126 includes a natural language description about searching criteria.

[0041] As shown in FIGS. 1 A and IB, the search engine 118 may receive the search query’ 126 via an input field 165 on a user interface 156 associated with an application 166 executable by the computing device 152. The search engine 118 uses the index structure(s) 155 to identify search results 162 responsive to the search query 126. For example, the search engine 118 may receive the search term(s) specified by the search query' 126 and obtain the relevant search results 162 by searching the index stmcture(s) 155. In further detail, the search engine 118 may obtain responsive documents 104a from the index structure(s) 155, rank the responsive documents 104a. and generate search results 162 for at least some of the responsive documents 104a.

[0042] Ranking may' include applying a plurality of ranking signals to the responsive documents 104a. The ranking signals may include signals relating to quality, uniqueness of content, user experience, social signals (e.g., popularity), relevance, authoritative, the use of keywords, and / or freshness of content. For at least some of the responsive documents 104a. the search engine 118 generates a search result 162 for a responsive document 104a. The search result 162 may include the title of the responsive document 104a, a resource identifier (e.g., a source, a uniform resource location (URL)) of the responsive document 104a, a description (e.g.. a snippet obtained from the metadata or content of the responsive document 104a), or other data related to the content, and / or image(s) and / or video(s) related to theresponsive document 104a. A resource identifier may identify the location of where the document 104 is stored in the database 102.

[0043] The search engine 118 may retrieve and initiate a display of the search results 162 on the user interface 156, where the search results 162 include responsive documents 104a that are responsive to the search query 126. In some examples, initiating display of the search results 162 includes transmitting information to the computing device 152 that identifies the search results 162. where the computing device 152 displays the search results 162 on the user interface 156. In some examples, initiating display of the search results 162 includes transmitting information to an application 166 (e.g., the chat application 168) that identifies the search results 162, where the application 166 displays the search results 162 on the user interface 156.

[0044] The LLM integrated search service 115 includes a prompt generator 114 that generates a prompt 124 for the LLM 170, where the prompt 124 includes the search query 126 and a context window 128 with at least information about the responsive documents 104a identified by the search engine 118. The amount of information in the context window 128 may define a token size of the prompt 124. In some examples, the context window 128 includes the content of the responsive documents 104a themselves. In some examples, as shown in FIG. IF, the context window' 128 includes summaries 175 of the responsive documents 104a. A summary 175 may be a summarization of a responsive document 104a, where the summary 175 is generated by the LLM 170.

[0045] The responsive documents 104a may include one, two, three, four, or more than four documents 104 that are responsive to the search query 126. In some examples, the prompt generator 114 may select a subset of the responsive documents 104a identified by the search engine 118 and include that subset in the context window 128. In some examples, the search engine 118 ranks the responsive documents 104a, and the prompt generator 114 may select the top N number of ranked responsive documents 104a for inclusion in the prompt 124. The integer N may be a predefined threshold such as one, tw o, three, or four.

[0046] In some examples, the context window 128 also includes conversation history’ data 110 relating to a chat history between the user and the LLM 170 with respect to the search query 126. In other words, the conversation history data 110 may be the model responses 130 and the LLM queries submitted by the user for a particular search query' 126. In some examples, the search engine 118 may obtain a conversation identifier 120 associated with the search query 126. In some examples, the search engine 118 obtains the conversation identifier 120 from the search query 126 that is received via the input field 165. The searchengine 118 may retrieve, from the database 102, the conversation history data 110 relating to the search query 126 using the conversation identifier 120. In some examples, the prompt generator 114 may include the conversation history data 110 in the prompt 124.

[0047] In some examples, the context window 128 also includes personalization data 122 about the user that submitted the search query 126. The personalization data 122 may be obtained from a user profile 108 stored in the database 102, where the user profile includes information about the user. The user and / or the user profile may be associated with a user identifier 145. A user identifier 145 may be a string of values that uniquely represent a user. The search engine 118 may obtain a user identifier 145 associated with a user that submitted the search query 126. In some examples, the search engine 118 obtains the user identifier 145 from the search query 126 received at the search engine 118. The search engine 118 retrieves the personalization data 122 from the database 102 using the user identifier 145, where the personalization data 122 includes information about the user. In some examples, the prompt generator 114 may include the personalization data 122 in the prompt 124.

[0048] The prompt generator 114 transmits the prompt 124 to the LLM 170, and the LLM 170 uses the search query 126 and the responsive documents 104a to formulate a model response 130 that responds to the search query 126. In some examples, the LLM 170 also uses the conversation history’ data 110 and the personalization data 122 to generate the model response 130. The LLM integrated search service 115 (e.g., the prompt generator 114, the search engine 118) may communicate with the LLM 170 via one or more application programming interfaces (APIs) 1 12. The model response 130 includes textual data 132 that responds to the search query’ 126. The textual data 132 may be generative artificial intelligence (Al) content. The model response 130 may identify one or more responsive documents 104a that were used (e.g., primarily used) to generate the textual data 132. In some examples, the model response 130 identifies a single responsive document 104a that was primarily used to generate the textual data 132.

[0049] The prompt generator 114 or the search engine 118 receives the model response 130 from the LLM 170. The prompt generator 114 or the search engine 118 may store the search query 126 and the model response 130 as conversation history data 110 in the database 102. The prompt generator 114 or the search engine 118 may initiate display of the model response 130 on the UI object 158. Initiating display of the model response 130 may include transmitting information to the computing device 152 that causes the computing device 152 to display the model response 130 in the UI object 158. Initiating display of the model response 130 may include transmitting information to the application 166 (e.g., thechat application 168) that causes the application 166 to display the model response 130 in the UI object 158.

[0050] In some examples, as shown in FIG. IB, the UI object 158 or the model response 130 is positioned on the user interface 156 in a location that is separate from the display of the search results 162. The UI object 158 may be a subset of the user interface 156 that displays a chat history between the user and the LLM 170. In some examples, the UI object 158 is positioned between the input field 165 that receives a search query 126. and the search results 162. The UI object 158 includes a chat input field 125 that enables a user to submit additional user queries 127 and receive model responses 130 that are responsive to the user queries 127. For example, a user may submit a user query 127 via the chat input field 125, which causes the prompt generator 114 to generate another prompt 124 for the LLM 170, and another model response 130 is subsequently displayed in the UI object 158. The prompt generator 114 may store any user queries 127 and corresponding model responses 130 as conversation history' data 110 in the database 102.

[0051] The search query 126 may be referred to as a first query, and a user query 127 submitted via the chat input field 125 may be referred to as a second query. The prompt generator 114 may generate a new' prompt 124 with the second query, and a context window' 128 w ith the responsive documents 104a, the conversation history data 110 relating to the first query, and / or the personalization data 122. In some examples, the new prompt 124 includes the same responsive documents 104 that were retrieved from the first query. In some examples, in response to the second query, the search engine 1 18 may use the second query' as a new' search query 126, w'hich causes the search engine 118 to re-retrieve a set of new' responsive documents 104a that are responsive to the second query' and includes the new- responsive documents 104a in the new prompt 124.

[0052] In some examples, the search system 100 may include one or more structures, techniques, or mechanisms that reduce the amount of tokens used in the context window' 128, which can further reduce the computational cost of processing LLM queries (e.g., search queries 126, user queries 127) by an LLM 170. In some examples, as shown in FIG. ID, the search engine 118 may retrieve, in response to a search query 126, a set of semantically similar documents 113 for the responsive documents 104a according to one or more search strategies 111. Using the search strategies 111 discussed herein, the set of semantically similar documents 113 may be relevant (e g., highly relevant) to the search query 126, where a lesser number of documents 104a can be included in the prompt 124, thereby reducing thetoken size of the prompt 124 (thereby reducing the computational cost of processing a LLM query).

[0053] The search strategies 111 may include a vector database search 121, a natural language processing (NLP) enrichment search 123, a late interaction model search 129, and / or a regular token matching search 131. In some examples, the search engine 118 uses a hybrid search 133 that uses a combination of two or more of the following search strategies 111: a vector database search 121, an NLP enrichment search 123. a late interaction model search 129, and / or a regular token matching search 131 . With respect to a vector database search 121, a vector database may store numeric representation of documents that capture the context and meaning of those documents, including text, images and audio. The numeric representation, called vectors, may be obtained using a pre-trained machine learning (ML) model. A vector database search 121 may find the vectors (documents) that are the closest (in the vector space) to the vector representation of the search query. This may be used to implement semantic search, e.g., find text data that is the closest semantically, or find similar images. A regular token matching search 131 may be a technique where the search algorithm matches user queries against tokens within a dataset, and these tokens can be words or phrases.

[0054] In some examples, the context window 128 includes conversation history data 110 relating to a particular search query 126. The conversation history data 110 may refer to the LLM queries and LLM responses associated with a particular search query 126. The conversation history data 1 10 stored in the database 1 2 may refer to the collection of history data across query sessions or search queries 126. The search engine 118 may retrieve conversation history data 110 relating to the search query 126 from the database 102 and include the conversation history data 110 in the prompt 124. The conversation history data 110 may include the search query 126, the model response 130, any user queries 127 submitted via the chat input field 125, and / or any model responses 130 that were generated in response to a user query 127.

[0055] In some examples, a separate (new) conversation identifier 120 is assigned to different search queries 126. When a first search query (e.g.. a search query 126) is submitted, the search system 100 may assign a conversation identifier 120 to the uery session (or the search query 126), where the conversation identifier 120 can be used by the search engine 118 or the prompt generator 114 to retrieve the conversation history data 110 relating to the first search query, e.g., any LLM responses and LLM queries relating to the first search query. The prompt generator 114 may include the conversation history data 110in the prompt 124. In some examples, in response to receipt of a second search query (e.g., a subsequent search query 126), the search system 100 may assign a new conversation identifier 120, and the prompt generator 114 may use the new conversation identifier 120 to retrieve the conversation history data 110 associated with the second search query.

[0056] The prompt generator 114 may include personalization data 122 in the prompt 124 so that the LLM 170 can personalize the model response 130. In some examples, the search engine 118 may retrieve personalization data 122 about the user that submitted the search query 126 from the database 102. In some examples, the database 102 stores user profiles 108 for the users that are authorized to access the database 102. Each user profile 108 may provide information about a respective user. Each user profile may be associated with a user identifier 145. As shown in FIG. IE, a user profile 108 may include a location 180 of the user, one or more organization groups 182 associated with the user, and one or more organization roles 184 associated with the user. In some examples, the search engine 118 obtains the user identifier 145 from the search query7126 received at the search engine 118. The search engine 118 retrieves the personalization data 122 from the user profile 108 stored in the database 102 using the user identifier 145. where the personalization data 122 includes information about the user. In some examples, the prompt generator 114 may include the personalization data 122 in the prompt 124.

[0057] If the user is located in Canada and the user submits a search query 126 about “how do I update my tax elections7’, the search engine 118 may retrieve documents 104a related to that search query7126, and the model response 130 may include information about updating tax elections in Canada (versus tax elections in the United States). Also, a model response 130 may depend on other personalization data 122. For example, in response to the query “how does compensation work”, the model response 130 may depend on whether the user is a manager, a non-managing engineer, or other organization role or group associated with the user.

[0058] In some examples, the personalization data 122 may include document access control data 185 about one or more access permissions (or restrictions) associated with a user for accessing the documents 104. In some examples, the search engine 118 may retrieve responsive documents 104a from the database 102 in which the user has access rights to. A document 104, a group of documents 104, or a database (or dataset) may include a document security7setting 186, and the search engine 118 may retrieve documents 104a responsive to the search query 126 that satisfies the document security setting 186. For example, a manager may be allowed to access documents 104 that are not accessible by non-managers.As such, model responses 130 generated by the LLM 170 may use documents 104 with access rights that correspond to the user that submitted the search query’ 126.

[0059] Referring to FIG. IF, the search system 100 may generate an initial prompt (e.g., prompt 124-1) that causes the LLM 170 to generate summaries 175 of the responsive documents 104a, and the summaries 175 are included in a subsequent prompt (e.g., prompt 124-2) for generating a model response (e.g., model response 130-2) that responses to the search query 126. For example, the prompt generator 114 may generate a prompt 124-1. The prompt 124-1 includes the responsive documents 104a (e.g., the content themselves) with a request to generate a summary' 175 of each responsive document 104a. In response to the prompt 124-1, the LLM 170 generates a summary 175 of each responsive document 104a. The LLM 170 transmits a model response 130-1 with the summaries 175. In response to the model response 130-1, the prompt generator 114 may generate a prompt 124-2 with the search query 126 and the summaries 175. The prompt 124-2 may also include the conversation history data 110 and / or the personalization data 122. In response to the prompt 124-2, the LLM 170 may generate a model response 130-2 with textual data 132 that responds to the search query 126.

[0060] In some examples, the prompt generator 114 is configured to communicate with multiple LLMs 170 to respond to a search query' 126. By distributing the LLM processing, the computational cost of generating a response to a search query 126 using underlying documents from a database associated with an organization may be reduced. In some examples, the LLMs 170 include an LLM 170-1 and an LLM 170-2. The LLM 170-1 may be configured to generate summaries 175 from responsive documents 104a, and the LLM 170-2 may be configured to generate textual data 132 that responses to the search query- 126 using the summaries 175. The LLM 170-1 may receive the prompt 124-1 with the responsive documents 104a and generate a model response 130-1 with the summaries 175. The LLM 170-2 may receive the prompt 124-2 with the search query’ 126 and the context window 128 having the summaries 175 and generate a model response 130-2 with the textual data 132 that responds to the search query 126 using the summaries 175.

[0061] In some examples, the LLM 170-1 is a language model that is smaller than the LLM 170-2. For example, the LLM 170-1 may have less layers and / or weights than the LLM 170-2. In some examples, the LLM 170-1 may be configured to only generate summaries 175. In some examples, the LLM 170-1 is stored on the server computer 160, and the LLM 170-2 is stored on the server computer 160a. In some examples, the LLM 170-1 and theLLM 170-2 are stored on the server computer 160. In some examples, the LLM 170-1 and the LLM 170-2 are stored on the server computer 160a.

[0062] The computing device 152 may be any type of computing device that includes one or more processors 101, one or more memory devices 103, a display 154, and an operating system 105 configured to execute (or assist with executing) one or more applications 166. including a chat application 168. The chat application 168 may be a program configured to communicate with the search engine 1 18. In some examples, the chat application 168 is a native application installable on the operating system 105. In some examples, the chat application 1 8 is a web application executable by a browser application (e.g., one of the applications 166). In some examples, the chat application 168 is a web page executable by a browser application. In some examples, the user interface 156 is an interface of the chat application 168. In some examples, the computing device 152 is a laptop computer. In some examples, the computing device 152 is a desktop computer. In some examples, the computing device 152 is a tablet computer. In some examples, the computing device 152 is a smartphone. In some examples, the computing device 152 is a wearable device (e.g., a head-mounted display device such as an augmented reality (AR) or a virtual reality (VR) device).

[0063] A browser application is a web browser configured to access information on the Internet. The browser application may launch one or more browser tabs in the context of one or more browser windows on a display 154 of the computing device 152. A browser tab may display content (e.g., web content) associated with a web document (e.g., webpage, PDF, images, videos, etc.) and / or an application such as a web application, progressive web application (PWA), and / or extension. A web application may be an application program that is stored on a remote server (e.g., server computer 160) and delivered over the network through the browser application (e.g., a browser tab). In some examples, the user interface 156 is not an interface of a browser application.

[0064] The operating system 105 is a system software that manages computer hardware, software resources and provides common services for the applications 166. In some examples, the operating system 105 is an operating system designed for a larger display 154 such as a laptop or desktop (e.g., sometimes referred to as a desktop operating system). In some examples, the operating system 105 is an operating system for a smaller display 154 such as a tablet or a smartphone (e.g., sometimes referred to as a mobile operating system). In some examples, the chat application 168 is executable by the operating system 105. The chat application 168 may receive the search query 126 via the input field 165 of the userinterface 156, and the chat application 168 may transmit the search query 126 to the search engine 118.

[0065] The processor(s) 101 may be formed in a substrate configured to execute one or more machine executable instructions or pieces of software, firmware, or a combination thereof. The processor(s) 101 can be semiconductor-based - that is, the processors can include semiconductor material that can perform digital logic. The memory device(s) 103 may include a main memory that stores information in a format that can be read and / or executed by the processor(s) 101. The memory device(s) 103 may store the operating system 105, including the chat application 168 that, when executed by the processors 101, performs certain operations discussed with reference to the chat application 168 discussed herein. In some examples, the memory device(s) store one or more portions of the LLM integrated search service 115 that, when executed by the processors 101, performs certain operations discussed with reference to the LLM integrated search service 115. In some examples, the memory7device(s) 103 includes a non-transitory computer-readable medium that includes executable instructions that cause at least one processor (e.g., the processors 101) to execute the operations discussed herein.

[0066] The server computer 160 may be computing devices that take the form of a number of different devices, for example a standard server, a group of such servers, or a rack server system. The server computer 160 may represent a single server computer or multiple server computer. In some examples, the server computer 160 may represent multiple server computers that are in communication with each other. In some examples, the server computer 160 may be a single system sharing components such as processors and memories. In some examples, the server computer 160 may be multiple systems that do not share processors and memories. The network may include the Internet and / or other types of data networks, such as a local area network (LAN), a wide area network (WAN), a cellular network, satellite network, or other types of data networks. The network may also include any number of computing devices (e.g., computers, servers, routers, network switches, etc.) that are configured to receive and / or transmit data within the network. The network may further include any number of hardwired and / or wireless connections.

[0067] The server computer(s) 160 may include one or more processors 161 formed in a substrate, an operating system (not shown) and one or more memory' devices 163. The memory device(s) 163 may represent any kind of (or multiple kinds of) memory (e.g., RAM, flash, cache, disk, tape, etc.). In some examples (not shown), the memory devices may include external storage, e.g., memory physically remote from but accessible by the servercomputer(s) 160. The processor(s) 161 may be formed in a substrate configured to execute one or more machine executable instructions or pieces of software, firmware, or a combination thereof. The processor(s) 161 can be semiconductor-based - that is, the processors can include semiconductor material that can perform digital logic. The memory device(s) 163 may store information in a format that can be read and / or executed by the processor(s) 161. The memory device(s) 163 may store one or more portions of the LLM integrated search service 1 15, that, when executed by the processor(s) 161. perform certain operations discussed herein. In some examples, the memory device(s) 163 includes a non- transitory computer-readable medium that includes executable instructions that cause at least one processor (e.g., the processor(s) 161) to execute operations.

[0068] The LLM 170 may include any type of pre-trained LLM configured to generate a model response 130 in response to a prompt 124. In some examples, the LLM 170 is stored on a server computer 160a that is separate from the server computer 160 that hosts the LLM integrated search sendee 115. The server computer 160a may be server computing resources that are owned and / or managed by an entity that is separate from an entity that owns and / or manages the server computer 160. In some examples, the LLM 170 is not owned or managed by the organization associated with the database 102. In some examples, the LLM 170 is a third-party' LLM that is not managed or owned by the search system 100. In some examples, the LLM 170 is a predefined LLM that is managed or owned by the search system 100. In some examples, the LLM 170 is stored on the server computer 160 that hosts the LLM integrated search service 1 15.

[0069] The LLM 170 includes weights. The weights are numerical parameters that the LLM 170 leams during the training process. The weights are used to compute the output (e.g., the model response 130) of the LLM 170. The LLM 170 may receive the prompt 124 from the prompt generator 1 14. The LLM 170 includes a pre-processing engine configured to pre-process the information in the prompt 124. Pre-processing may include converting the textual input of the prompt 124 to individual tokens (e.g., words, phrases, or characters). Preprocessing may include other operations such as removing stop words (e.g., “the'’, “and", “of’) or other terms or syntax that do not impart any' meaning to the LLM 170. The LLM 170 includes an embedding engine configured to generate word embeddings from the pre- processed text input. The word embeddings may be vector representations that assist the LLM 170 to capture the semantic meaning of the input tokens and may assist the LLM 170 to better understand the relationships between the input tokens.

[0070] The LLM 170 includes neural network(s) configured to receive the word embeddings and generate an output. A neural network includes multiple layers of interconnected neurons (e.g., nodes). The neural network may include an input layer, one or more hidden layers, and an output later. The output may include a sequence of output word probability distributions, where each output distribution represents the probability of the next word in the sequence given the input sequence so far. In some examples, the output may be represented as a probability distribution over the vocabulary or a subset of the vocabulary. The neural network(s) is configured to receive the word embeddings and generate an output, and, in some examples, the query activity (e.g., previous natural language queries and textual responses). The output may represent a version of the model response 130. The output may include a sequence of output word probability distributions, where each output distribution represents the probability of the next word in the sequence given the input sequence so far. In some examples, the output may be represented as a probability' distribution over the vocabulary or a subset of the vocabulary . The decoder is configured to receive the output and generate the model response 130. In some examples, the decoder may select the most likely instruction, sampling from a probability distribution, or using other techniques to generate coherent and well written model response 130.

[0071] In some examples, the database 102 may be stored on a server computer 160 that also includes, or is associated with an entity that also manages, the LLM integrated search service 115 (e.g., the ingestion engine 116. the search engine 118. the prompt generator 1 14). In some examples, the database 102 is external to the server computer 160 that hosts the LLM integrated search sendee 115. In other words, in some examples, the database 102 is owned and / or managed by an entity that is different from the entity that owns and / or manages the search system 100. In some examples, the database 102 may be an external data store.

[0072] FIG. 2 illustrates a communication diagram depicting example operations of generating a model response with responsive documents that are responsive to a search uery according to an aspect. The example operations of FIG. 2 may be executed by the search system 100 of FIGS. 1A to IF. The communication diagram of FIG. 2 enables a predefined LLM (e.g., LLM 270) to formulate answers to queries using documents with domain-specific content that are responsive to a search query' (e.g., human chat input), where the documents stored in the database (e.g., database 202-1) include information that may not have been used to configure (e.g., train) the LLM 270. In this manner, the techniques discussed herein may integrate an LLM 270 in a search system that manages domain-specific content in a mannerthat can reduce the amount of computing resources (e.g., CPU power, memory) for integrating a LLM 270 in a search system. Further, the techniques discussed herein may maintain or increase the security of using organization documents for an LLM 270 to answer questions from users.

[0073] In operation 201, a chat application 268 receives a search query (e.g., human chat input) from a user. The chat application 268 may be an example of the chat application 168 of FIGS. 1 A to IF and may include any of the details discussed with reference to those figures. The chat application 268 may receive the search query (e g., the search query 126 of FIGS. 1 A to IF) via an input field (e.g., input field 165 of FIGS. 1A to IF) on a user interface (e.g., user interface 156 of FIGS. 1A to IF). In some examples, the chat application 268 may initiate storage of the search query in a database 202-2. In some examples, the database 202- 2 is a memory device that stores conversation history data (e.g., the conversation history data 110 of FIGS. lA to IF).

[0074] In operation 203, the chat application 268 initiates a search for semantically similar documents (e.g., responsive documents 104a, semantically similar documents 113 of FIGS. 1A to IF) in a database 202-1. The database 202-1 may be an example of the database 102 of FIGS. 1 A to IF and may include any of the details discussed with reference to those figures. A search engine (e.g., the search engine 118 of FIGS. 1A to IF) may receive the search query from the chat application 268, and the search engine may retrieve responsive documents from the database 202-1 that satisfy the term(s) of the search query.

[0075] In operation 205, the chat application 268 receives the search results (e.g., the search results 162 of FIGS. 1A to IF), where the search results 162 identify the responsive documents. For example, the search engine may generate the search results and provide the search results to the chat application 268. The chat application 268 may operate with a prompt generator (e.g., the prompt generator 114 of FIGS. lA to IF) to generate a prompt (e.g., the prompt 124 of FIGS. 1A to IF) with the search query and the responsive documents. In operation 207, the prompt is transmitted to an LLM 270. The LLM 270 may be an example of the LLM 170 of FIGS. 1A to IF and may include any of the details discussed herein. In response to the prompt, the LLM 270 may generate a model response (e g., the model response 130 of FIGS. lA to IF) (e g., the chat reply). The model response may include textual data that responds to the search query. In operation 209, the chat application 268 may receive the model response to be displayed on the user interface. In operation 211, the chat application 268 displays the responsive documents and the modelresponse. In some examples, the chat application 268 may initiate storage of the model response in the database 202-2.

[0076] FIG. 3 illustrates a communication diagram depicting example operations of generating a model response with responsive documents that are responsive to a search query according to another aspect. The example operations of FIG. 3 may be executed by the search system 100 of FIGS. lA to IF. The communication diagram of FIG. 3 enables a predefined LLM (e.g., LLM 370) to formulate answers to queries using documents with domain-specific content that are responsive to a search query (e.g., question), where the documents stored in the database include information that may not have been used to configure (e.g.. train) the LLM 370. In this manner, the techniques discussed herein may integrate an LLM 370 in a search system that manages domain-specific content in a manner that can reduce the amount of computing resources (e.g., CPU power, memory) for integrating a LLM 370 in a search system. Further, the techniques discussed herein may maintain or increase the security of using organization documents for an LLM 270 to answer questions from users.

[0077] In operation 301, a search UI 356 receives a search query (e.g., question) from a user. The search UI 356 may be an example of the user interface 156 of FIGS. 1A to IF and may include any of the details discussed with reference to those figures. In response to the search query, in operation 303, a chat API 312 may initiate a search to retrieve responsive documents (e.g., responsive documents 104a of FIGS. lA to IF), and, in some examples, to initiate retrieval of conversation history data (e g., conversation history data 110) associated with a conversation identifier (e.g., the conversation identifier 120 of FIGS. 1 A to IF). In operation 305, the chat API 312 may initiate a search engine 318 to retrieve the conversation history data.

[0078] In operation 307, the chat API 312 may initiate a search engine 318 to retrieve responsive documents and facets that satisfy the search query. In operation 309, the chat API 312 may initiate a search engine 318 to add conversation of responsive documents to summarize and persist conversation messages. In operation 311, the chat API 312 returns the responsive documents, facets, and stream identifier. In operation 313, the search UI 356 may render the responsive documents and facets. In operation 315, the search UI 356 initiates the chat API 312 with a request to perform a uery completion using a streaming identifier. In operation 317, in response to the request, the chat API 312 initiates a search engine 318 to retrieve conversation messages linked to the streaming identifier. Then, a prompt generator may transmit a prompt (e.g., the prompt 124) to the LLM 370 to receive a model response(e.g., the model response 130 of FIGS. 1 A to IF). In operation 319, the chat API 312 may receive the model response (e.g., a stream summary response) from the LLM 370. In operation 321, the chat API 312 returns the model response to the search UI 356. In operation 323, the search UI 356 displays the model response (e.g., the summarization). In operation 325, the search UI 356 may receive a further question or an indication of an adjustment to a facet.

[0079] FIGS. 4A to 4C illustrate an example user interface 456 for integrating an ULM into a search system according to an aspect. The user interface 456 may be an example of the user interface 156 of FIGS. 1A to IF and may include any of the details discussed with reference to those figures. The user interface 456 includes an input field 465 that receives a search query 426 (e.g.. ‘"Does the company own my side hustles”). Below the input field 465, the user interface 456 includes a UI object 458 that displays a model response with textual data 432 generated by an LUM (e.g., the LLM 170 of FIGS. 1 A to IF) that responds to or answers the search query 426. In some examples, the UI object 458 includes feedback user elements 490 that enables the user to provide feedback about the textual data 432 generated by the LLM. The UI object 458 may include a chat input field 425 that receives a user query 427. In some examples, the user query 427 may be a follow up question to be answered by the LLM. Below the UI object 458, the user interface 456 includes the search results 462.

[0080] As shown in FIG. 4A, in response to the submission of the search query 426, the user interface 456 may display the search results 462 relatively quickly with an indication that the answer (e.g., the textual data 432) is being processed by the LLM. Then, after the model response is generated, as shown in FIG. 4B, the answer is provided in the UI obj ect 458. In some examples, the user interface 456 first displays the search results 462. which is followed by the display of the textual data 432 generated by the LLM. In FIG. 4C, the UI object 458 may also identify a responsive document 404a that includes information that was used to generate the answer. In some examples, the UI object 458 includes a selectable element associated with the responsive document 404a, which, when selected, causes the display of the responsive document 404a.

[0081] FIG. 5 illustrates an example user interface 556 for integrating an LLM into a search system according to an aspect. The user interface 556 displays textual data 532 generated by the LLM using personalization data (e.g., the personalization data 122 of FIGS. 1A to IF). In some examples, the personalization data may include a location of the user, organization group(s) associated with the user, and / or an organization role of the user. Forexample, if the user is located in Canada and the user submits a search query 526 about "how do I update my tax elections”, the search system may retrieve documents related to that search query 526, and the model response may include textual data 532 about updating tax elections in Canada (versus tax elections in the United States).

[0082] The user interface 556 may be an example of the user interface 156 of FIGS. 1A to IF and may include any of the details discussed with reference to those figures. The user interface 556 includes an input field 565 that receives a search query 526 (e.g., ‘"How do I update my tax elections”). Below the input field 565, the user interface 556 includes a UI object 558 that displays a model response with textual data 532 generated by an LLM (e.g., the LLM 170 of FIGS. 1A to IF) that responds to or answ ers the search query' 526. In some examples, the UI object 558 includes feedback user elements 590 that enables the user to provide feedback about the textual data 532 generated by the LLM. The UI object 558 may also identify a responsive document 504a that includes information that was used to generate the answer. In some examples, the UI object 558 includes a selectable element associated with the responsive document 504a, which, when selected, causes the display of the responsive document 504a. The UI object 558 may include a chat input field 525 that receives a user query 527. In some examples, the user query 527 may be a follow up question to be answered by the LLM. Below the UI object 558, the user interface 556 includes the search results.

[0083] FIGS. 6 A to 6D illustrate an example user interface 656 for integrating an LLM into a search system according to an aspect. The user interface 656 displays textual data 632 generated by the LLM using personalization data (e.g., the personalization data 122 of FIGS. 1 A to IF). In some examples, the personalization data may include a location of the user, organization group(s) associated with the user, and / or an organization role 692 (e.g., engineer) of the user. In response to the search query 626 (“what is our work from home policy”), the answer (e.g., the textual data 632 generated by the LLM) may depend on the organization role 692 of the user.

[0084] The user interface 656 may be an example of the user interface 156 of FIGS. 1A to IF and may include any of the details discussed with reference to those figures. The user interface 656 includes an input field 665 that receives a search query 626 (e.g., “what is our work from home policy”). Below the input field 665, the user interface 656 includes a UI object 658 that displays a model response with textual data 632 generated by an LLM (e.g., the LLM 170 of FIGS. 1A to IF) that responds to or answers the search query 626. As shown in FIGS. 6C and 6D, the UI object 658 may also identify a responsive document 604athat includes information that was used to generate the answer. In some examples, the UI object 658 includes a selectable element associated with the responsive document 604a, which, when selected, causes the display of the responsive document 604a. The UI object 658 may include a chat input field 625 that receives a user query. In some examples, the user query may be a follow up question to be answered by the LLM. Below the UI object 658, the user interface 656 includes the search results 662.

[0085] As shown in FIG. 6A, a user may enter a search query 626 into the input field 665 of the user interface 656. In some examples, the user interface 656 may also display an organization role 692 (e.g., engineer) associated with the user. In some examples, the user interface 656 also displays one or more suggested search queries 694. The suggested search queries 694 may be previous search queries entered by the user and / or other users, which are related to the terms being entered in the input field 665.

[0086] In response to the submission of the search query 626, as show n in FIG. 6B, the search results 662 are displayed, where the search results 662 identify the responsive documents. Also, the user interface 656 begins to display the textual data 632 of the model response. In FIG. 6B, the answer is still being developed, and the textual data 632 includes the information that has been generated by the LLM. FIG. 6C illustrates the completed answer along w ith an indication of the responsive document 604a that was used to formulate the answer. A user may submit a user query 627 (e.g., “how many changes have been made to this policy in 2023”) via the chat input field 625. Then, as shown in FIG. 6D, the user interface 656 may display a subsequent answer (e.g., textual data 632) that responds to the user uery 627. In some examples, the LLM may generate the subsequent answer using the documents retrieved from the search query' 626 (e g., “what is our work from home policy”). In some examples, the LLM may generate the subsequent answer using updated documents retrieved from the user query' 627 (e.g., “how many changes have been made to this policy in 2023”).

[0087] FIG. 7 illustrates an example user interface 756 for integrating an LLM into a search system according to an aspect. The user interface 756 displays textual data 732 generated by the LLM that accounts for the domain-specific content of the documents that are stored in the database. For example, in response to a search query 726 (e.g., “What is NASA”), the user interface 756 displays an answer with textual data 732 that accounts for the domain-specific content of the documents (e.g., “NASA stands for North American South America Region.. . .”). In contrast, a conventional LLM that receives the same query maygenerate a response that indicates that NASA stands for National Aeronautics and Space Administration.

[0088] The user interface 756 may be an example of the user interface 156 of FIGS. 1 A to IF and may include any of the details discussed with reference to those figures. The user interface 756 includes an input field 765 that receives a search query 726 (e.g., “What is NASA”). Below the input field 765, the user interface 756 includes a UI object 758 that displays a model response with textual data 732 generated by an LLM (e.g., the LLM 170 of FIGS. 1 A to IF) that responds to or answers the search query 726. The UI object 758 may also identify a responsive document 704a that includes information that was used to generate the answer. In some examples, the UI object 758 includes a selectable element associated with the responsive document 704a, which, when selected, causes the display of the responsive document 704a. The UI object 758 may include a chat input field 725 that receives a user query 727. In some examples, the user query 727 may be a follow up question to be answered by the LLM. Below7the UI object 758, the user interface 756 includes the search results 762.

[0089] FIGS. 8A and 8B illustrate an example user interface 856 for integrating an LLM into a search system according to an aspect. The user interface 856 displays textual data 832 generated by the LLM using personalization data (e.g., the personalization data 122 of FIGS. 1 A to IF). In some examples, the personalization data may include a location of the user, organization group(s) associated with the user, and / or an organization role 892 (e.g., engineer in FIG. 8A, manager in FIG. 8B) of the user. In response to the search query 826 (“does compensation work”), the answer (e.g., the textual data 632 generated by the LLM) may depend on the organization role 692 of the user. For example, in FIG. 8 A, the organization role 892 of the user is an engineer, and therefore, the answer relates to compensation for an engineer. In FIG. 8B, the organization role 892 of the user is a manager, and therefore, the answer relates to compensation for a manager.

[0090] The user interface 856 may be an example of the user interface 156 of FIGS. 1A to IF and may include any of the details discussed with reference to those figures. The user interface 856 includes an input field 865 that receives a search query 826 (e.g., “how does compensation work”). Below the input field 865, the user interface 856 includes a UI object 858 that displays a model response with textual data 832 generated by an LLM (e.g., the LLM 170 of FIGS. 1A to IF) that responds to or answ ers the search query7826. The UI object 858 may include a chat input field 825 that receives a user query 827. In someexamples, the user query may be a follow up question to be answered by the LLM. Below the UI object 858, the user interface 856 includes the search results 862.

[0091] FIG. 9 is a flowchart 900 depicting example operations of a system for generating a model response with responsive documents that are responsive to a search query according to an aspect. The example operations of FIG. 9 may be executed by the search system 100 of FIGS. 1A to IF. The operations of FIG. 9 enable a predefined LLM to formulate answers to queries using documents with domain-specific content that are responsive to a search query, where the documents stored in the database include information that may not have been used to configure (e.g., train) the LLM. In this manner, the operations discussed herein may integrate an LLM in a search system that manages domainspecific content in a manner that can reduce the amount of computing resources (e.g., CPU power, memory) for integrating a LLM in a search system. Further, the techniques discussed herein may maintain or increase the security of using organization documents for an LLM to answer questions from users.

[0092] The flowchart 900 may depict operations of a computer-implemented method. Although the flowchart 900 is explained with respect to the search system 100 of FIGS. 1 A through IF, the flowchart 900 may be applicable to any of the implementations discussed herein, including FIGS. 2 to 8B. Although the flowchart 900 of FIG. 9 illustrates the operations in sequential order, it will be appreciated that this is merely an example, and that additional or alternative operations may be included. Further, operations of FIG. 9 and related operations may be executed in a different order than that shown, or in a parallel or overlapping fashion. The flowchart 900 may depict a computer-implemented method.

[0093] Operation 902 includes receiving, via an input field on a user interface of a computing device, a search query for documents stored in a database associated with an organization. Operation 904 includes retrieving, from the database, search results that are responsive to the search query, the search results including responsive documents. Operation 906 includes initiating display of the search results on the user interface.

[0094] Operation 908 includes transmitting a prompt to a large language model, the prompt including the search query and information from at least a portion of the responsive documents. Operation 910 includes receiving, from the large language model, a model response that is responsive to the search query', the model response being generated by the large language model using at least the portion of the responsive documents. Operation 912 includes initiating display of the model response in a user interface object on the user interface.

[0095] Clause 1. A method comprising: receiving, via an input field on a user interface of a computing device, a search query for documents stored in a database associated with an organization; retrieving, from the database, search results that are responsive to the search query, the search results including responsive documents; initiating display of the search results on the user interface; transmitting a prompt to a large language model, the prompt including the search query and information from at least a portion of the responsive documents; receiving, from the large language model, a model response with textual data that responds to the search query, the textual data being generated by the large language model using at least the portion of the responsive documents; and initiating display of the model response in a user interface object associated with the user interface.

[0096] Clause 2. The method of clause 1, further comprising: obtaining a conversation identifier associated with the search query; and retrieving, from the database, conversation history data relating to the search query using the conversation identifier, the prompt also including the conversation history data.

[0097] Clause 3. The method of clause 1 or 2, further comprising: obtaining a user identifier associated with a user that submitted the search query; and retrieving, from the database, personalization data having information about the user, the prompt also including the personalization data.

[0098] Clause 4. The method of clause 3, wherein the personalization data includes at least one of a location of the user, an organization group associated with the user, or an organization role of the user.

[0099] Clause 5. The method of clause 3, wherein the personalization data includes document access control data associated with the user, the method further comprising: retrieving, from the database, a responsive document with a document security setting that satisfies the document access control data associated with the user.

[0100] Clause 6. The method of any one of clauses 1 to 5, wherein the user interface object includes a chat input field, wherein the prompt is a first prompt, and the model response is a first model response, the method further comprising: receiving, via the chat input field, a user query; transmitting a second prompt to the large language model, the second prompt including the user query, at least a portion of the responsive documents, at least a portion of information included in the first prompt, and the first model response; receiving, from the large language model, a second model response that is responsive to the user query; and initiating display of the second model response in the user interface object.

[0101] Clause 7. The method of any one of clauses 1 to 6. wherein the model response identifies a responsive document that was used to generate the model response, the method further comprising: providing a selectable item on the user interface object, wherein the selectable item, when selected, displays the responsive document.

[0102] Clause 8. The method of any one of clauses 1 to 7, further comprising: transmitting an initial prompt to the large language model, the initial prompt including a request to generate summaries for the responsive documents; receiving, from the large language model, a response that includes the summaries; and transmitting the prompt to the large language model, the prompt including the summaries and the search query.

[0103] Clause 9. A non-transitory computer-readable medium storing instructions that cause at least one processor to execute operations, the operations comprising: receiving, via an input field on a user interface of a computing device, a search query for documents stored in a database associated with an organization; retrieving, from the database, search results that are responsive to the search query, the search results including responsive documents; initiating display of the search results on the user interface; transmitting a prompt to a large language model, the prompt including the search query and information from at least a portion of the responsive documents; receiving, from the large language model, a model response with textual data that is responsive to the search query', the textual data being generated by the large language model using at least the portion of the responsive documents; and initiating display of the model response in a user interface object associated with the user interface.

[0104] Clause 10. The non-transitory computer-readable medium of clause 9, wherein the operations further comprise: obtaining a conversation identifier associated with the search query; and retrieving, from the database, conversation history data relating to the search query using the conversation identifier, the prompt also including the conversation history data.

[0105] Clause 11. The non-transitory' computer-readable medium of clause 9 or 10, wherein the operations further comprise: obtaining a user identifier associated with a user that submitted the search query; and retrieving, from the database, personalization data having information about the user, the prompt also including the personalization data.

[0106] Clause 12. The non-transitory' computer-readable medium of clause 11, wherein the personalization data includes at least one of a location of the user, an organization group associated with the user, or an organization role of the user.

[0107] Clause 13. The non-transitory computer-readable medium of clause 11, wherein the personalization data includes document access control data associated with the user, wherein the operations further comprise: retrieving, from the database, a responsive document with a document security setting that satisfies the document access control data associated with the user.

[0108] Clause 14. The non-transitory’ computer-readable medium of any one of clauses 9 to 13, wherein the user interface object includes a chat input field, wherein the prompt is a first prompt, and the model response is a first model response, wherein the operations further comprise: receiving, via the chat input field, a user query; transmitting a second prompt to the large language model, the second prompt including the user query, at least a portion of the responsive documents, at least a portion of information from the first prompt, and the first model response; receiving, from the large language model, a second model response that is responsive to the user query; and initiating display of the second model response in the user interface object.

[0109] Clause 15. The non-transitory computer-readable medium of any one of clauses 9 to 14, wherein the model response identifies a responsive document that was used to generate the model response, wherein the operations further comprise: providing a selectable item on the user interface object, wherein the selectable item, when selected, displays the responsive document.

[0110] Clause 16. An apparatus comprising: at least one processor; and a non- transitory computer-readable medium storing executable instructions that when executed by the at least one processor cause the at least one processor to execute operations, the operations including: receiving, via an input field on a user interface of a computing device, a search query for documents stored in a database associated with an organization; retrieving, from the database, search results that are responsive to the search query, the search results including responsive documents; initiating display of the search results on the user interface; transmitting a prompt to a large language model, the prompt including the search query and information from at least a portion of the responsive documents; receiving, from the large language model, a model response with textual data that is responsive to the search query, the textual data being generated by the large language model using at least the portion of the responsive documents; and initiating display of the model response in a user interface object associated with the user interface.

[0111] Clause 17. The apparatus of clause 16. wherein the operations further comprise: obtaining a conversation identifier associated with the search query; and retrieving,from the database, conversation history data relating to the search query' using the conversation identifier, the prompt also including the conversation history data.

[0112] Clause 18. The apparatus of clause 16 or 17, wherein the operations further comprise: obtaining a user identifier associated with a user that submitted the search query7; and retrieving, from the database, personalization data having information about the user, the prompt also including the personalization data.

[0113] Clause 19. The apparatus of clause 18. wherein the personalization data includes at least one of a location of the user, an organization group associated with the user, or an organization role of the user, w herein the personalization data includes document access control data associated with the user, wherein the operations further comprise: retrieving, from the database, a responsive document with a document security setting that satisfies the document access control data associated with the user.

[0114] Clause 20. The apparatus of any one of clauses 16 to 19, w herein the user interface object includes a chat input field, wherein the prompt is a first prompt, and the model response is a first model response, wherein the operations further comprise: receiving, via the chat input field, a user query; transmitting a second prompt to the large language model, the second prompt including the user query, at least a portion of the responsive documents, at least a portion of information from the first prompt, and the first model response: receiving, from the large language model, a second model response that is responsive to the user query; and initiating display of the second model response in the user interface object.

[0115] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry7, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0116] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” “computer-readable medium” refers to any computer program product, apparatus and / ordevice (e.g., magnetic discs, optical disks, memory7, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0117] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0118] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (“LAN”), a wide area network (“WAN”), and the Internet.

[0119] The computing system can include clients and servers. A client and server are remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship with each other.

[0120] In this specification and the appended claims, the singular forms "a," "an" and "the" do not exclude the plural reference unless the context clearly dictates otherwise. Further, conjunctions such as “and,” “or,” and “and / or” are inclusive unless the context clearly dictates otherw ise. For example, “A and / or B” includes A alone, B alone, and A with B. Further, connecting lines or connectors shown in the various figures presented are intended to represent example functional relationships and / or physical or logical couplings between the various elements. Many alternative or additional functional relationships, physical connections or logical connections may be present in a practical device. Moreover,no item or component is essential to the practice of the implementations disclosed herein unless the element is specifically described as “essential” or “critical”.

[0121] Terms such as, but not limited to, approximately, substantially, generally, etc. are used herein to indicate that a precise value or range thereof is not required and need not be specified. As used herein, the terms discussed above will have ready and instant meaning to one of ordinary skill in the art.

[0122] Moreover, use of terms such as up. down, top. bottom, side, end, front, back, etc. herein are used with reference to a currently considered or illustrated orientation. If they are considered with respect to another orientation, it should be understood that such terms must be correspondingly modified.

[0123] Further, in this specification and the appended claims, the singular forms "a," "an" and "the" do not exclude the plural reference unless the context clearly dictates otherwise. Moreover, conjunctions such as “and,” “or,” and “and / or” are inclusive unless the context clearly dictates otherwise. For example, “A and / or B” includes A alone, B alone, and A with B.

[0124] Although certain example methods, apparatuses and articles of manufacture have been described herein, the scope of coverage of this patent is not limited thereto. It is to be understood that terminology' employed herein is for the purpose of describing particular aspects and is not intended to be limiting. On the contrary, this patent covers all methods, apparatus and articles of manufacture fairly falling within the scope of the claims of this patent.

Claims

WHAT IS CLAIMED IS:

1. A method comprising: receiving, via an input field on a user interface of a computing device, a search query for documents stored in a database associated with an organization; retrieving, from the database, search results that are responsive to the search query, the search results including responsive documents; initiating display of the search results on the user interface; transmitting a prompt to a large language model, the prompt including the search query and information from at least a portion of the responsive documents; receiving, from the large language model, a model response with textual data that responds to the search query, the textual data being generated by the large language model using at least the portion of the responsive documents; and initiating display of the model response in a user interface object associated with the user interface.

2. The method of claim 1 , further comprising: obtaining a conversation identifier associated with the search query7; and retrieving, from the database, conversation history data relating to the search query using the conversation identifier, the prompt also including the conversation history data.

3. The method of claim 1 or 2, further comprising: obtaining a user identifier associated with a user that submitted the search query; and retrieving, from the database, personalization data having information about the user. the prompt also including the personalization data.

4. The method of claim 3, wherein the personalization data includes at least one of a location of the user, an organization group associated with the user, or an organization role of the user.

5. The method of claim 3, wherein the personalization data includes document access control data associated with the user, the method further comprising: retrieving, from the database, a responsive document with a document security setting that satisfies the document access control data associated with the user.

6. The method of any of claims 1 to 5, wherein the user interface object includes a chat input field, wherein the prompt is a first prompt, and the model response is a first model response, the method further comprising: receiving, via the chat input field, a user query; transmitting a second prompt to the large language model, the second prompt including the user query, at least a portion of the responsive documents, at least a portion of information included in the first prompt, and the first model response; receiving, from the large language model, a second model response that is responsive to the user query; and initiating display of the second model response in the user interface object.

7. The method of any of claims 1 to 6, wherein the model response identifies a responsive document that was used to generate the model response, the method further comprising: providing a selectable item on the user interface object, wherein the selectable item, when selected, displays the responsive document.

8. The method of any of claims 1 to 7, further comprising: transmitting an initial prompt to the large language model, the initial prompt including a request to generate summaries for the responsive documents; receiving, from the large language model, a response that includes the summaries; and transmitting the prompt to the large language model, the prompt including the summaries and the search query.

9. A non-transitory computer-readable medium storing instructions that cause at least one processor to execute operations, the operations comprising: receiving, via an input field on a user interface of a computing device, a search query for documents stored in a database associated with an organization; retrieving, from the database, search results that are responsive to the search query, the search results including responsive documents; initiating display of the search results on the user interface; transmitting a prompt to a large language model, the prompt including the search query and information from at least a portion of the responsive documents;receiving, from the large language model, a model response with textual data that is responsive to the search query, the textual data being generated by the large language model using at least the portion of the responsive documents; and initiating display of the model response in a user interface object associated with the user interface.

10. The non-transitory computer-readable medium of claim 9, wherein the operations further comprise: obtaining a conversation identifier associated with the search query7; and retrieving, from the database, conversation history data relating to the search query using the conversation identifier, the prompt also including the conversation history data.

11. The non-transitory computer-readable medium of claim 9 or 10, wherein the operations further comprise: obtaining a user identifier associated with a user that submitted the search query; and retrieving, from the database, personalization data having information about the user, the prompt also including the personalization data.

12. The non-transitory computer-readable medium of claim 11, wherein the personalization data includes at least one of a location of the user, an organization group associated with the user, or an organization role of the user.

13. The non-transitory computer-readable medium of claim 11, wherein the personalization data includes document access control data associated with the user, wherein the operations further comprise: retrieving, from the database, a responsive document with a document security setting that satisfies the document access control data associated with the user.

14. The non-transitory computer-readable medium of any of claims 9 to 13, wherein the user interface object includes a chat input field, wherein the prompt is a first prompt, and the model response is a first model response, wherein the operations further comprise: receiving, via the chat input field, a user query;transmitting a second prompt to the large language model, the second prompt including the user query, at least a portion of the responsive documents, at least a portion of information from the first prompt, and the first model response; receiving, from the large language model, a second model response that is responsive to the user query; and initiating display of the second model response in the user interface object.

15. The non-transitory computer-readable medium of any of claims 9 to 14, wherein the model response identifies a responsive document that was used to generate the model response, wherein the operations further comprise: providing a selectable item on the user interface object, wherein the selectable item, when selected, displays the responsive document.

16. An apparatus comprising: at least one processor; and a non-transitory computer-readable medium storing executable instructions that when executed by the at least one processor cause the at least one processor to execute operations, the operations including: receiving, via an input field on a user interface of a computing device, a search query for documents stored in a database associated with an organization; retrieving, from the database, search results that are responsive to the search query', the search results including responsive documents; initiating display of the search results on the user interface; transmitting a prompt to a large language model, the prompt including the search query and information from at least a portion of the responsive documents; receiving, from the large language model, a model response with textual data that is responsive to the search query , the textual data being generated by the large language model using at least the portion of the responsive documents; and initiating display of the model response in a user interface object associated with the user interface.

17. The apparatus of claim 16, wherein the operations further comprise: obtaining a conversation identifier associated with the search query; andretrieving, from the database, conversation history data relating to the search query using the conversation identifier, the prompt also including the conversation history data.

18. The apparatus of claim 16 or 17, wherein the operations further comprise: obtaining a user identifier associated with a user that submitted the search query: and retrieving, from the database, personalization data having information about the user. the prompt also including the personalization data.

19. The apparatus of claim 18, wherein the personalization data includes at least one of a location of the user, an organization group associated with the user, or an organization role of the user, wherein the personalization data includes document access control data associated with the user, wherein the operations further comprise: retrieving, from the database, a responsive document with a document security setting that satisfies the document access control data associated with the user.

20. The apparatus of any of claims 16 to 19. wherein the user interface object includes a chat input field, wherein the prompt is a first prompt, and the model response is a first model response, wherein the operations further comprise: receiving, via the chat input field, a user query; transmitting a second prompt to the large language model, the second prompt including the user query, at least a portion of the responsive documents, at least a portion of information from the first prompt, and the first model response; receiving, from the large language model, a second model response that is responsive to the user query; and initiating display of the second model response in the user interface object.