Generating semantic search engine result pages
Through the semantic search engine system, the problem of insufficient information presentation in classic search engines is solved by using large language models to summarize and generate query answers, and the problem of inadequate presentation of information in classic search engines is achieved, and search results are easier to understand and navigate, providing detailed summary and direct answers.
Patent Information
- Application Number
- CN202480005182.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-30
- Filing Date
- 2024-01-31
- Publication Date
- 2025-07-11
AI Technical Summary
Classic search engines retrieve relevant content only based on user queries without providing additional information or analysis, resulting in users needing to navigate multiple results to determine relevant information, lacking simplified and structured presentation of information.
Using a semantic search engine system, the Big Language Model (LLM) is used to summarize and generate the answers to the query, provide the main facts and ambiguity removal related to the query, collect information from multiple data sources through machine learning models and format it into easy-to-understand summary, including direct answers or detailed summary, and optionally provide citation links.
Improves the comprehensibility and navigation efficiency of search results, and reduces the user's navigation needs between multiple results by generating detailed summary and direct answers, providing a more comprehensive information presentation.
Smart Images

Figure CN120303654A_ABST
Abstract
Description
Background Art
[0001] Classic search engines typically retrieve and rank relevant content based solely on a user's query, without providing additional information or analysis. Without additional information, users are required to navigate through multiple results to determine information relevant to their query. Embodiments have been described for these and other general considerations. Moreover, although relatively specific problems have been discussed, it should be understood that embodiments are not limited to solving the specific problems identified in the background art. Summary of the Invention
[0002] Aspects of the present disclosure relate to systems and methods for providing a semantic search engine capable of performing functions beyond the capabilities of classic search engines, such as, for example, summarizing and generating answers to queries, and providing a brief overview of key facts, aspects, or other disambiguation related to the query. Aspects of the present disclosure relate to organizing and summarizing information from a retrieval-based search engine into a semantically meaningful format, thus making the information easier for search engine users to understand and navigate.
[0003] The present summary is provided to introduce in a simplified form some concepts that will be further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of examples will be set forth in part in the description below, and in part will be apparent from the description, or may be learned by practice of the present disclosure. Brief Description of the Drawings
[0004] Non-limiting and non-exhaustive examples are described with reference to the following drawings.
[0005] Figure 1 An exemplary system including a semantic search engine is depicted.
[0006] Figure 2 is a block diagram illustrating an exemplary method for generating semantic search engine results.
[0007] Figure 3 is a block diagram illustrating a method for generating prompt words for a machine learning model utilized by a semantic search engine.
[0008] Figure 4 An exemplary user interface providing a summary of information generated by a semantic search engine is provided.
[0009] Figure 5A and Figure 5B illustrates an overview of an example generative machine learning model that may be used in accordance with aspects described herein.
[0010] Figure 6A block diagram showing example physical components of a computing device that can be utilized to practice aspects of the present disclosure is shown.
[0011] Figure 7 A simplified block diagram of a computing device that can be utilized to practice aspects of the present disclosure is shown.
[0012] Figure 8 A simplified block diagram of a distributed computing system in which aspects of the present disclosure can be practiced. Detailed Description
[0013] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, and in which are shown by way of illustration specific embodiments or examples. These aspects may be combined, other aspects may be utilized, and structural changes may be made without departing from the present disclosure. The embodiments may be practiced as a method, system, or apparatus. Accordingly, the embodiments may take the form of a hardware implementation, a fully software implementation, or an implementation combining software and hardware aspects. Thus, the following detailed description should not be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims and their equivalents.
[0014] Aspects of the present disclosure relate to organizing, synthesizing, and summarizing information from classical retrieval-based search engines into a semantically meaningful format such that the results are easier for a user to understand and navigate. That is, the aspects disclosed herein relate to synthesizing traditional search information in a manner that satisfies the intent associated with the received query. As part of the synthesis, aspects of the present disclosure can collect additional information from a variety of different data sources, such as local document stores, third-party platforms, applications, etc., in order to address the query intent. Aspects of the present disclosure can create a summary that provides an overview of the initial search results of the information, and then create subsections for disambiguation of different aspects of the original search query based on its intent. These subsections use reference links to attribute the summary information to its source to provide credibility. In some examples, in addition to or as an alternative to providing a summary, the aspects disclosed herein can provide the entire document, web page, dataset, etc., rather than creating a summary. Among other benefits, aspects of the present disclosure help users quickly find and understand the information they are looking for by providing a curated and structured view of the search engine results page (SERP).
[0015] Aspects of the present disclosure retrieve relevant information from a search engine based on a user's search query. The query can be a classic search query (keywords or phrases) or a conversational query (e.g., a chat message between a user and / or a chatbot), a query based on an email or other type of message, or a query generated based on a content item (e.g., a web page, an image, a video, a document, etc.). Aspects of the present disclosure utilize a large language model (LLM) (such as, for example, a generative model) to summarize content according to the intent detected from the query. In some cases, aspects of the present disclosure can generate a direct answer to the query and provide relevant references to support the information. Additionally, the aspects disclosed herein use information from reference documents to provide a brief overview of the main facts or aspects related to the user's query. The model can access data such as the date and location of the query and previous web results (e.g., the top five results, the top five results, the top ten results, etc.) and the surrounding information and / or context information of each result.
[0016] Among other technical benefits, aspects of the present disclosure provide capabilities beyond those of classic search engines by summarizing and generating answers to queries and providing a brief overview of the main facts, aspects, or other disambiguation related to the query. Classic search engines typically only retrieve and rank relevant content based on the user's query without providing additional information or analysis. Our system achieves new capabilities by leveraging large language models. Those skilled in the art will appreciate other technical benefits provided by the aspects disclosed herein.
[0017] Figure 1 An exemplary system 100 including a semantic search engine 120 system is depicted. System 100 includes a computing device 102, a semantic search engine 120, and one or more data repositories 106 that communicate via a network 115. The computing device 102 can be any of a variety of computing devices, including but not limited to mobile computing devices, laptop computing devices, tablet computing devices, desktop computing devices, and / or virtual reality computing devices. The computing device 102 can be configured to execute one or more applications 104 and / or services and / or manage hardware resources (e.g., processors, memory, etc.) that can be utilized by a user of the computing device 102. The (multiple) applications 104 can be native applications or web-based applications. For example, the (multiple) applications 104 can be a web browser, a digital personal assistant, a file browser, etc. The (multiple) applications 104 can be used for communication across the network 150 to submit queries to the semantic search engine 120. Although not shown, in an alternative example, an instance of the semantic search engine 120 can reside locally on the computing device 102.
[0018] In an example, the semantic search engine 120 receives a query from the computing device 102 and processes the query using the query processor 124. In one example, the query can be a query for information on a network such as the Internet. For example, the query can be a query provided to a search engine. In other aspects, the query can be generated based on user intent derived from user interactions (e.g., a user interacting with a chatbot, a user selecting a web page or other type of content) and / or from other content items (e.g., emails, documents, web pages, presentations, etc.). In yet another example, aspects of the present disclosure can generate additional queries related to the received query (e.g., disambiguation queries, alternative queries, etc.). In an example, the additional queries can be generated by an associated search engine, by a machine learning model, such as one or more of the models in the model repository 130. In an example, the query processor 124 processes the query (or queries) and generates an initial result set in response to receiving the query. For example, the query processor can be a search engine that generates a web search result set based on the received query. The query and the search result set can be provided to a machine learning model to process the initial result set. For example, one or more machine learning (ML) models can be stored in the model repository 130. The query processor 124 can provide the results to the models from the repository based on the type of content retrieved in the search results. In one example, a generative large language model (LLM) can be used to process the search results generated by the query processor 124. The generative models used in accordance with the aspects described herein (which are generally also referred to herein as types of ML models) can generate any variety of output types (and thus can be multimodal generative models in some examples), and in some examples can be generative transformer models and / or large language models (LLMs), generative image models. Example ML models include, but are not limited to, Generative Pretrained Transformer 3 (GPT-3), BigScience BLOOM (Large Open Science Open Access Multilingual Language Model), DALL-E, DALL-E2, Stable Diffusion, or Jukebox. Referring below to Figures 5A to 5BThe generative ML model shown discusses additional examples of these aspects. A generative large language model can process search results and determine whether an initial result set satisfies the intent or task associated with a query. If not, the generative LLM, which is part of the semantic search engine 120, can generate additional searches for information that can be used to satisfy the intent and / or task associated with the query. The generated searches can be provided to the query processor 124 and / or the data source search interface 126 to query one or more additional data sources based on the generated queries. In an example, different types of data sources 106 can be searched, such as, for example, web pages, application data repositories, document repositories, databases, etc. The data source search interface 126 helps process queries across different data sources. For example, the data source search interface 126 can include an API or library that can be utilized to access data from different data sources (such as weather information, stock information, third-party databases, etc.) to gather additional information relevant to the query and / or the intent determined based on the query and / or user interactions.
[0019] When retrieving data required to answer a query intent and / or task, the machine learning model employed by the semantic search engine 120 can summarize the content found in the results. As will also be discussed below, the machine learning model can be prompted to generate a summary in a specific format. The prompt generator 128 can be used to generate one or more prompts and provide the generated prompts to the ML model. The one or more provided prompts can be used to format the query results summary into a format suitable for result summarization. In an example, the prompts can include templates that can be used by the machine learning model to format information.
[0020] Figure 2 An exemplary method 200 for generating semantic search engine results is depicted. The process begins at operation 202 of receiving a query. In one example, the query can be a query for searching content on the web, such as a query received by a web search engine. For ease of explanation, the examples discussed herein are described with respect to web search queries; however, those skilled in the art will understand that the aspects disclosed herein can be used to process other types of queries, such as, for example, local directory searches, database searches, document repository queries, social media queries, audio and / or visual search queries, etc. In an example, the intent can be derived from the query. For example, a rule-based system, heuristics, and / or a machine learning model can be used to analyze the query to determine the intent or task associated with the user query. In addition to the query, the intent and / or task can be provided at operation 204.
[0021] The process continues to operation 204, where, in the example, a query is executed and the result of the query or a subset of the results (e.g., top results, top ten results, top one hundred results, data from relevant sources (e.g., information from news sources, weather sources, shopping sources, etc.) or other relevant data sources) is provided to a machine learning model together with the received query. For example, the results can be provided to a generative model, such as a generative LLM model. In one example, the underlying content of the search results (e.g., web page content, content from the database where the query was executed, documents, videos, audio files, etc. identified in response to the query) can be provided to a database. Alternatively or additionally, instead of providing the entire content (e.g., the entire web page), a summary of the content can be provided. Summary data related to the content can be pre-generated and retrieved from the database. Alternatively or additionally, one or more different machine learning models can be used to summarize the results beforehand, and the generated summary can be provided to the generative model. In other aspects, one or more different types of generative machine learning models can receive the search results and the query. The type of model that receives the query can be determined based on the type of the results (e.g., content, format, such as image, text, video, etc.).
[0022] The process continues to decision operation 206, where it is determined whether the initial search results are sufficient to respond to the query. For example, one or more machine learning models that receive the query and the initial query results can determine whether the results answer the query. For example, as described above, the query can be analyzed to determine the intent and / or task associated with the query. The intent can be analyzed when the query is received, or the intent can be determined by the generative model when the query and the results are processed at operation 204. Based on the determined intent and / or task, the search results can be analyzed to determine whether the intent and / or task associated with the query can be adequately resolved. If not, the process branches "no" to operation 208.
[0023] If the initial query results do not sufficiently satisfy the query (e.g., do not sufficiently satisfy the intent or task associated with the query), one or more machine learning models can generate additional search queries. The additional queries can relate to information not explicitly requested by the query. As an example, the received initial query can be: "Is February a good time to visit Japan?" The machine learning model can determine that the intent of the query is to plan a vacation to Japan in February. Although the initial search results generated, for example, by a web search engine can provide links to articles about Japan in February, one or more machine learning models can determine that the intent requires a more comprehensive answer, which may require additional information. In making this determination, the machine learning model can generate additional queries such as, for example, "Weather in Japan in February", "Things to do in Japan in February", "Things to do in Tokyo in February", "Flights to Japan", etc. These additional queries can be executed, for example, using a search engine to generate additional results, which can then be processed by one or more machine learning models.
[0024] At operation 208, a query for additional information can be performed, for example, by a search engine, a file system, a database, etc., and additional search results and optionally additional queries can be provided to one or more machine learning models. The additional information retrieved from these additional queries can be used to provide a comprehensive response to the initial query, thus satisfying the intent and / or task determined for the initial query, without a multi-step process of communicating with the user. The flow then continues to operation 210.
[0025] Briefly returning to decision operation 206, if one or more machine learning models determine that the initial set of query results satisfies the intent and / or task determined based on the query, the flow branches "Yes" to operation 210.
[0026] Although the examples provided herein relate to web search queries, the additional queries are not limited to web search. For example, the additional queries generated by one or more machine learning models can be queries for information searching a local device or data store (in the case where the user has given one or more ML models permission to search a local data repository), or can be directed to other data stores (e.g., API calls to the query application, database queries, calls to a specific data store such as stock data or weather data, etc.).
[0027] At operation 210, one or more prompt words can be provided to the ML model. The one or more provided prompt words can be used to format the query results summary into a format suitable for result summary. In an example, operation 210 can be optional. That is, one or more prompt words for formatting the results can be provided earlier, for example, with the initial query and result set, with additional search results generated at operation 208, etc. In an example, one or more prompt words can be templates that can be used to format or summarize information generated by one or more generative models. In an example, the template can be selected based on the type of data generated by the generative model, based on the task associated with the query, based on the intent associated with the query, etc.
[0028] Aspects of the present disclosure are operable to utilize a general ML model, i.e., a model not specifically trained to generate semantic search engine results. For example, method 200 can employ a generative large language model. Generally, an LLM is not trained to perform a specific task. Thus, the one or more prompt words generated and provided at operation 210 direct the generative large language model (or other type of generative machine learning model) to generate a summary of the results in a format suitable for the initially received query and / or based on the determined intent and / or task associated with the query.
[0029] The process continues to operation 212, where a summary generated by one or more machine learning models is received from the one or more machine learning models. While traditional search engines typically return links to web pages or files that match a search query, aspects of the present disclosure generate a summary that provides a detailed summary of the content related to the query. In an example, the summary of the content is formatted based on the one or more prompt words generated at operation 210. Additionally, in an example, the summary includes a reference to the underlying data source (e.g., web page, document, video, etc.) of the information included in the summary. The link can be selectable, enabling the user to be redirected to the source material by selecting the reference. Alternatively, depending on the determined intent, a direct answer can be generated. For example, if the query intent is related to specific information, such as the query "What is Abraham Lincoln's birthday?", a direct answer can be generated such that the answer is provided without a summary. In yet another aspect, both a direct answer and a summary can be generated and / or provided.
[0030] For example, Figure 4 An exemplary user interface depicting a summary of information generated by the semantic search engine is provided. Now turning to Figure 4 , an exemplary user interface 400 is provided in the form of a chat interface, where a query 402 regarding visiting Japan is received. In an example, the query can be received via a text box 401. In Figure 4In an alternative user interface not shown, a query can be received via different UI components, such as an address bar in an internet browser, via a search engine text box, via audio (e.g., verbal query, etc.). As Figure 4 shown, query 402 can be a natural language query. The semantic search engine generates a detailed summary that can be used to answer the user's query about information on Japan. In the example, the detailed summary can be divided into different parts based on the topic. For example, part 404 details general activities in Japan, part 406 shows things to do in Tokyo, part 408 details things to do in Kyoto, and part 410 details things to do in Niseko-mura. Additionally, the generative LLM can generate a summary that includes reference points, which, when activated by the user, can guide the user to the information source.
[0031] Although a specific user interface is shown in Figure 4 , alternative user interfaces can be employed, such as a search interface integrated into a browser or web page, a search interface as part of an operating system or application, etc. For example, in alternative aspects, the information summary can be included in a web page that includes traditional search results. For example, the summary of the search results can be included before, in the middle, or after the traditional list of results generated by the search engine. In another example, the summary can be displayed as part of another application user interface (e.g., within a mobile application, file browser, and operating system features, etc.). In yet another aspect, although Figure 4 not shown in, but in addition to or instead of the text summary, the summary can include various types of content. For example, the summary can include images, videos, animations, audio playback, other types of generated resources that can be displayed as part of the summary (e.g., documents, spreadsheets, presentations, etc.). Those skilled in the art will understand that the summary can be included in multiple different user interfaces capable of receiving and displaying the information generated by the aspects of the present disclosure.
[0032] Return Figure 2 , at operation 214, semantic search engine results are provided. In one example, the results can be provided to the application that received the initial query (e.g., a web browser, chat interface, etc.). In another example, providing the results can include displaying the results or causing the results to be displayed. In yet another example, the semantic search engine results can be stored for future use. That is, the semantic search engine results can be indexed and stored for retrieval when a subsequent query with a similar or related intent is received. By doing so, the response time for future queries can be reduced because the summary can be retrieved instead of generating a summary in response to the query.
[0033] Figure 3is a block diagram showing a method 300 for generating prompt words for a machine learning model utilized by a semantic search engine. The process begins at operation 302, where a query is analyzed to determine the intent and / or task associated with the query. In one example, a machine learning model can be used to determine the intent or task associated with the query. Alternatively or additionally, a rule-based process, query parser, or other type of process can be utilized to determine the intent associated with the application.
[0034] After determining the intent and / or task of the query, the process continues to operation 304, where the response format is determined. The response format can be based on a summary of the predicted content in response to the query. For example, the type of predicted content, content length, etc. can be used to determine the appropriate format to best present the response to the query.
[0035] At operation 306, one or more prompt words can be determined based on the determined response format. As described above, a generative LLM can be utilized by the semantic search engine to generate a response to the query. The generative LLM may not be trained to generate a specific type of response or response format required to satisfy the query intent and / or task. Instead of fine-tuning the generative LLM, which can be a long and expensive process, one or more prompt words can be generated and provided to guide the LLM to produce a response in the desired format. In one example, one or more prompt words can be generated by selecting an appropriate predefined prompt word template from a prompt word repository. In another example, a machine learning model trained to generate prompt words can be employed to generate one or more prompt words based on the received query (e.g., the query received at Figure 2 operation 202). After generating the prompt words, the process continues to operation 308, where the one or more prompt words are provided to the machine learning model utilized by the semantic search engine.
[0036] Figure 5A and Figure 5B shows an overview of an example generative machine learning model that can be used in accordance with aspects described herein. First referring to Figure 5A , the conceptual diagram 500 depicts an overview of a pre-trained generative model package 504 that processes an input and prompt words 502 to generate aspects of the model output 506 described herein.
[0037] In an example, the generative model package 504 is pre-trained based on various inputs (e.g., various human languages, various programming languages, and / or various content types), and thus does not need to be fine-tuned or trained for a specific scenario. Instead, the generative model package 504 can be more generally pre-trained such that the input 502 includes prompts that are generated, selected, or otherwise designed to cause the generative model package 504 to produce certain generative model outputs 506. It should be understood that the input 502 and the generative model outputs 506 can each include any number of content types, including but not limited to text output, image output, audio output, video output, programming output, and / or binary output, among other examples. In an example, the input 502 and the generative model outputs 506 can have different content types, as is the case when the generative model package 504 includes a generative multimodal machine learning model.
[0038] Accordingly, the generative model package 504 can be used in any of a variety of scenarios, and furthermore, different generative model packages can be used in place of the generative model package 504 with substantially no modification to other associated aspects (e.g., similar to those described herein with respect to Figure 1 , Figure 2 , Figure 3 and Figure 4 ). Thus, the generative model package 504 operates as a tool for performing machine learning processing, where certain inputs 502 to the generative model package 504 are programmatically generated or otherwise determined such that the generative model package 504 produces model outputs 506 that can subsequently be used for further processing.
[0039] The generative model package 504 can be provided or otherwise used according to any of a variety of paradigms. For example, the generative model package 504 can be used locally on a computing device (e.g., the computing device 102 in Figure 1 ), or can be accessed remotely from a machine learning service (e.g., the semantic search engine 120). In other examples, aspects of the generative model package 504 are distributed across multiple computing devices. In some cases, the generative model package 504 can be accessed via an application programming interface (API), such as can be provided by an operating system of a computing device and / or by a machine learning service, among other examples.
[0040] Referring now to the illustrated aspects of the generative model package 504, the generative model package 504 includes input tokenization 508, input embedding 510, model layers 512, output layer 514, and output decoding 516. In an example, the input tokenization 508 processes the input 502 to generate the input embedding 510 that includes a sequence of symbolic representations corresponding to the input 502. Accordingly, the input embedding 510 is processed by the model layers 512, the output layer 514, and the output decoding 516 to produce the model output 506.Figure 5B An example architecture corresponding to the generative model package 504 is depicted, which is also discussed in detail below. Even so, it should be understood that the architectures shown and described herein should not be considered restrictive, and in other examples, any of a variety of other architectures may be used.
[0041] Figure 5B FIG. is a conceptual diagram depicting an example architecture 550 of a pre-trained generative machine learning model that can be used in accordance with the aspects described herein. As described above, any of a variety of alternative architectures and corresponding ML models may be used in other examples without departing from the aspects described herein.
[0042] As shown, the architecture 550 processes the input 502 to produce a generative model output 506, aspects of which were discussed above with respect to Figure 5A The architecture 550 is depicted as a transformer model including an encoder 552 and a decoder 554. The encoder 552 processes an input embedding 558 (aspects of which may be similar to Figure 5A the input embedding 510 in ) that includes a sequence of symbolic representations corresponding to the input 556. In an example, the input 556 includes inputs and prompt words for generating 502 (e.g., corresponding to the skills of a skill chain).
[0043] Additionally, positional encoding 560 may introduce information about the relative and / or absolute positions of the tokens of the input embedding 558. Similarly, the output embedding 574 includes a sequence of symbolic representations corresponding to the output 572, and the positional encoding 576 may similarly introduce information about the relative and / or absolute positions of the tokens of the output embedding 574.
[0044] As shown, the encoder 552 includes an example layer 570. It should be understood that any number of such layers may be used, and the depicted architecture is simplified for illustrative purposes. The example layer 570 includes two sub-layers: a multi-head attention layer 562 and a feed-forward layer 566. In an example, residual connections are included around each of the layers 562, 566, followed by a normalization layer 564 and a normalization layer 568, respectively.
[0045] The encoder 554 includes an example layer 590. Similar to encoder 552, any number of such layers may be used in other examples, and the depicted architecture of decoder 554 is simplified for illustrative purposes. As shown, the example layer 590 includes three sub-layers: a masked multi-head attention layer 578, a multi-head attention layer 582, and a feed-forward layer 586. Aspects of the multi-head attention layer 582 and the feed-forward layer 586 may be similar to those discussed above with respect to the multi-head attention layer 562 and the feed-forward layer 566, respectively. Additionally, the masked multi-head attention layer 578 performs multi-head attention on the output of the encoder 552 (e.g., output 572). In an example, the masked multi-head attention layer 578 prevents a position from attending to subsequent positions. This masking, combined with the offset embedding (e.g., by one position, as shown in the multi-head attention layer 582), can ensure that the prediction for a given position depends on the known outputs of one or more positions less than the given position. As shown, residual connections are also included around layer 578, layer 582, and layer 586, followed by normalization layers 580, normalization layer 584, and normalization layer 588, respectively.
[0046] The multi-head attention layer 562, the multi-head attention layer 578, and the multi-head attention layer 582 may each linearly project queries, keys, and values to corresponding dimensions using a set of linear projections. An attention function (e.g., dot product or additive attention) may be used to process each linear projection, resulting in an n-dimensional output value for each linear projection. The resulting values may be concatenated and projected again such that the values are subsequently processed as Figure 5B shown (e.g., by the corresponding normalization layer 564, normalization layer 580, or normalization layer 584).
[0047] The feed-forward layer 566 and the feed-forward layer 586 may each be a fully connected feed-forward network applied to each position. In an example, the feed-forward layer 566 and the feed-forward layer 586 each include a plurality of linear transformations with rectified linear unit activation therebetween. In an example, each linear transformation is the same across different positions, while different parameters may be used compared to other linear transformations of the feed-forward network.
[0048] Additionally, aspects of the linear transformation 592 may be similar to the linear transformations discussed above with respect to the multi-head attention layer 562, the multi-head attention layer 578, the multi-head attention layer 582, and the feed-forward layer 566 and the feed-forward layer 586. The Softmax 594 may also convert the output of the linear transformation 592 into predicted next token probabilities, as indicated by the output probabilities 596. It should be understood that the architectures shown are provided as examples, and in other examples, any of a variety of other model architectures may be used in accordance with the disclosed aspects.
[0049] Accordingly, the output probability 596 can thus form a model output 506 according to the various aspects described herein, such that the output of the generative ML model is defined to correspond to the input. For example, the model output 506 can be associated with corresponding applications and / or data formats such that the model output is processed to display a semantic search engine page, among other examples.
[0050] Figures 6 to 8 and the associated description provides a discussion of various operating environments in which aspects of this description can be practiced. However, with regard to Figures 6 to 8 the devices and systems shown and discussed are for purposes of example and illustration and are not limiting of the numerous computing device configurations that can be used to practice the aspects of the present disclosure described herein.
[0051] Figure 6 is a block diagram that illustrates physical components (e.g., hardware) of a computing device 600 that can be utilized to practice aspects of the present disclosure. The computing device components described below can be applicable to the computing devices described above, including Figure 1 the computing device 102 in. In a basic configuration, the computing device 600 can include at least one processing unit 602 and a system memory 604. Depending on the configuration and type of the computing device, the system memory 604 can include, but is not limited to, volatile storage (e.g., random access memory), non-volatile storage (e.g., read only memory), flash memory, or any combination of these memories.
[0052] The system memory 604 can include an operating system 605 and one or more program modules 606 suitable for running software applications 620, such as one or more components supported by the systems described herein. As an example, the system memory 604 can store a semantic search engine 624 and / or (a) machine learning model(s) 626. The operating system 605, for example, can be suitable for controlling the operation of the computing device 600.
[0053] Furthermore, aspects of the present disclosure can be practiced in conjunction with a graphics library, other operating systems, or any other application programs and are not limited to any particular application or system. This basic configuration is shown in Figure 6 by those components within the dashed line 608. The computing device 600 can have additional features or functionality. For example, the computing device 600 can also include additional data storage devices (removable and / or non-removable) such as, for example, magnetic disks, optical disks, or magnetic tapes. Such additional storage is shown in Figure 6 by a removable storage device 609 and a non-removable storage device 610.
[0054] As described above, multiple program modules and data files may be stored in the system memory 604. When executed on the processing unit 602, the program modules 606 (e.g., applications 620) may perform processes including but not limited to aspects as described herein. Other program modules that may be used in accordance with aspects of the present disclosure may include, for example, email and contact applications, word processing applications, spreadsheet applications, database applications, slide presentation applications, drawing or computer-aided applications, and the like.
[0055] In addition, aspects of the present disclosure may be practiced in a circuit that includes discrete electronic elements, a packaged or integrated electronic chip that contains logic gates, a circuit that utilizes a microprocessor, or on a single chip that contains electronic elements or a microprocessor. For example, aspects of the present disclosure may be practiced via a system-on-a-chip (SOC), where Figure 6 each or many of the components shown in may be integrated onto a single integrated circuit. Such an SOC device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all of which are integrated (or “burned”) onto a chip substrate as a single integrated circuit. When operating via an SOC, the functionality described herein regarding the client handover protocol capabilities may be operated via specific application logic integrated on a single integrated circuit (chip) with other components of the computing device 600. Some aspects of the present disclosure may also be practiced using other technologies capable of performing logical operations such as, for example, AND, OR, and NOT, including but not limited to mechanical, optical, fluidic, and quantum technologies. Additionally, some aspects of the present disclosure may be practiced within a general purpose computer or in any other circuit or system.
[0056] The computing device 600 may also have one or more input devices 612 such as a keyboard, mouse, pen, voice or speech input device, touch or swipe input device, and the like. Output devices 614 such as a display, speaker, printer, etc. may also be included. The foregoing devices are examples, and other devices may be used. The computing device 600 may include one or more communication connections 616 that allow communication with other computing devices 650. Examples of suitable communication connections 616 include but are not limited to radio frequency (RF) transmitter, receiver, and / or transceiver circuitry; universal serial bus (USB), parallel, and / or serial ports.
[0057] As used herein, the term computer-readable medium may include computer storage media. Computer storage media can include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, or program modules. System memory 604, removable storage device 609, and non-removable storage device 610 are all examples of computer storage media (e.g., memory storage). Computer storage media can include RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other article which can be used to store the information and which can be accessed by computing device 600. Any such computer storage media can be part of computing device 600. Computer storage media does not include carrier waves or other propagated or modulated data signals.
[0058] Communication media can be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and includes any information delivery media. The term "modulated data signal" can describe a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media can include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0059] Figure 7 is a block diagram of an architecture showing one aspect of a computing device. That is, the computing device can implement some aspects in conjunction with a system (e.g., architecture) 702. In some examples, system 702 is implemented as a "smart phone" capable of running one or more applications (e.g., browser, email, calendar, contact manager, messaging client, games, and media client / player). In some aspects, system 702 is integrated as a computing device, such as an integrated personal digital assistant (PDA) and wireless phone.
[0060] One or more applications 766 may be loaded into the memory 762 and run on or in association with the operating system 764. Examples of applications include a phone dialer, an email program, a personal information management (PIM) program, a word processing program, a spreadsheet program, an Internet browser program, a messaging program, and the like. The system 702 also includes a non-volatile storage area 768 within the memory 762. The non-volatile storage area 768 can be used to store persistent information that should not be lost when the system 702 is powered down. The applications 766 may use and store information in the non-volatile storage area 768, such as emails or other messages used by an email application. A synchronization application (not shown) also resides on the system 702 and is programmed to interact with a corresponding synchronization application residing on a host computer to keep the information stored in the non-volatile storage area 768 synchronized with the corresponding information stored at the host computer. As should be understood, other applications may be loaded into the memory 762 and run on the mobile computing device 700 described herein (e.g., an embedded object memory insertion engine, an embedded object memory retrieval engine, etc.).
[0061] The system 702 has a power supply 770, which may be implemented as one or more batteries. The power supply 770 may also include an external power source, such as an AC adapter or a charging docking station that supplements or recharges the battery.
[0062] The system 702 may also include a radio interface layer 772, which performs the functions of sending and receiving radio frequency communications. The radio interface layer 772 facilitates a wireless connection between the system 702 and the "outside world" via a communication carrier or service provider. Transmissions to and from the radio interface layer 772 occur under the control of the operating system 764. In other words, communications received by the radio interface layer 772 may be propagated to the applications 766 via the operating system 764, and vice versa.
[0063] The visual indicator 720 can be used to provide visual notifications, and / or the audio interface 774 can be used to generate audible notifications via the audio transducer 725. In the example shown, the visual indicator 720 is a light-emitting diode (LED), and the audio transducer 725 is a speaker. These devices can be directly coupled to the power supply 770 such that when activated, they remain on for a duration indicated by the notification mechanism even if the processor 760 and / or the dedicated processor 761 and other components may be turned off to conserve battery power. The LED can be programmed to remain on indefinitely until the user takes an action to indicate the powered-on state of the device. The audio interface 774 is used to provide audible signals to the user and receive audible signals from the user. For example, in addition to being coupled to the audio transducer 725, the audio interface 774 can also be coupled to a microphone to receive audible input, such as to facilitate a phone conversation. According to aspects of the present disclosure, the microphone can also be used as an audio sensor to facilitate control of notifications, as will be described below. The system 702 can also include a video interface 776 that enables operation of the vehicle camera 730 to record still images, video streams, and the like.
[0064] The computing device implementing the system 702 can have additional features or functionality. For example, the computing device can also include additional data storage devices (removable and / or non-removable) such as, for example, magnetic disks, optical disks, or magnetic tapes. Such additional storage is Figure 7 illustrated by the non-volatile storage area 768.
[0065] As described above, data / information generated or captured by the computing device and stored via the system 702 can be stored locally on the computing device, or the data can be stored on any number of storage media that can be accessed by the device via the radio interface layer 772 or via a wired connection between the computing device and a separate computing device associated with the computing device (e.g., a server computer in a distributed computing network such as the Internet). As should be understood, such data / information can be accessed via the computing device, via the radio interface layer 772, or via a distributed computing network. Similarly, such data / information can be easily transferred between computing devices for storage and use according to well-known data / information transfer and storage components, including email and collaborative data / information sharing systems.
[0066] Figure 8Shows an aspect of the architecture of a system for processing data received at a computing system from remote sources such as a personal computer 804, a tablet computing device 806, or a mobile computing device 808, as described above. The content displayed at the server device 802 can be stored in different communication channels or other storage types. For example, a directory service 824, a web portal 825, a mailbox service 826, an instant message store 828, or a social networking site 830 can be used to store various documents.
[0067] An application 820 (e.g., similar to application 620) can be adopted by a client communicating with the server device 802. Additionally or alternatively, the server device 802 can adopt a machine learning model 821. The server device 802 can provide data to and from client computing devices such as a personal computer 804, a tablet computing device 806, and / or a mobile computing device 808 (e.g., a smart phone) via a network 815. As an example, the computer system described above can be embodied in a personal computer 804, a tablet computing device 806, and / or a mobile computing device 808 (e.g., a smart phone). In addition to receiving graphical data that can be used for preprocessing at a graphics generation system or postprocessing at a receiving computing system, these examples of any computing device can obtain content from a repository 816.
[0068] In one example, aspects of the present disclosure relate to a system including: at least one processor; and a memory storing instructions that, when executed by the at least one processor, cause the system to perform a set of operations including: receiving a query; generating an initial query result set; providing the query and the initial query result set to a generative large language model; receiving at least one additional query from the generative large language model; executing the at least one additional query; providing the results from the at least one additional query to the generative large language model; receiving semantic search engine results from the generative large language model; and providing the semantic search engine results.
[0069] In another example, the system further includes instructions for generating one or more alternative queries based on the query, and wherein generating the initial query result set includes: generating alternative query results based on the one or more additional queries.
[0070] In another example, the semantic search engine results are included in a summary generated by the generative large language model.
[0071] In another example, the format of the summary is determined based on the type of information included in the summary.
[0072] In another example, the format for the summary is determined based on a template provided to the generative large language model.
[0073] In yet another example, a summary includes one or more references, and one or more of the references link to one or more underlying data sources for the summary.
[0074] In more examples, the system includes: determining an intent or task based on a received query, where the intent or task is provided to a generative large language model.
[0075] In more examples, the system further includes an operation for using the generative large language model to determine whether additional information is needed, where the determination is based on the intent or task.
[0076] In yet another example, at least one additional query is generated by the generative large language model when it is determined that additional information is needed.
[0077] In another example, aspects of the present disclosure relate to a method for generating semantic search engine results, the method including: receiving a query; generating an initial query result set; providing the query and the initial query result set to a generative model; using the generative model to determine that additional information is needed; receiving at least one additional query from the generative model; executing at least one additional query; providing the results of at least one additional query to the generative large language model; receiving semantic search engine results from the generative model; and providing the semantic search engine results.
[0078] In an example, the method further includes analyzing the query to determine an intent or task based on the query, where analyzing the query includes providing the query to at least one of a generative model or an alternative machine learning model.
[0079] In yet another example, the method further includes determining a format of the semantic search engine results, where the format is determined based on the query or the task.
[0080] In more examples, the method further includes generating a prompt for the generative model, where the prompt is generated based on the format.
[0081] In yet another example, the prompt includes a template associated with the format, where the template defines the format for the semantic search engine results.
[0082] In another example, the generative model is a generative large language model.
[0083] In yet another example, the semantic search engine results are included in a summary generated by the generative large language model, and where the summary includes one or more references, and one or more of the references link to one or more underlying data sources for the summary.
[0084] In another aspect, aspects of the present disclosure relate to a computer storage medium including computer-executable instructions that, when executed by at least one processing unit, perform a method for generating semantic search engine results, the method including: receiving a query; generating an initial set of query results; providing the query and the initial set of query results to a generative large language model; using the generative model to determine a need for additional information; receiving at least one additional query from the generative large language model; performing the at least one additional query; providing the results from the at least one additional query to the generative large language model; receiving semantic search engine results from the generative large language model; and providing the semantic search engine results.
[0085] In an example, the method further includes: analyzing the query to determine an intent or task based on the query, wherein analyzing the query includes providing the query to at least one of a generative model or an alternative machine learning model; and determining a format for the semantic search engine results, wherein the format is determined based on the query or the task.
[0086] In yet another example, the method further includes generating a prompt for the generative model, wherein the prompt is generated based on the format.
[0087] In yet another example, the semantic search engine results are included in a summary generated by the generative large language model, and wherein the summary includes one or more references, and wherein the one or more references link to one or more underlying data sources for the summary.
[0088] For example, aspects of the present disclosure have been described above with reference to block diagrams and / or operational illustrations of methods, systems, and computer program products according to aspects of the present disclosure. The functions / actions recited in the blocks may occur out of the order shown in any flowchart. For example, two blocks shown in succession may in fact be executed subsequently simultaneously depending on the functionality / action involved, or the blocks may sometimes be executed in the reverse order.
[0089] The description and illustration of one or more aspects provided in this application are not intended to limit or restrict the scope of the claimed disclosure in any way. The aspects, examples, and details provided in this application are considered sufficient to convey possession and enable others to make and use the aspects of the claimed disclosure. The claimed disclosure should not be construed as limited to any aspect, example, or detail provided in this application. Various features (structural and methodical) are intended to be selectively included or omitted, whether shown and described in combination or separately, to produce embodiments having a particular set of features. Having provided the description and illustration of this application, those skilled in the art can envision variations, modifications, and alternative aspects that fall within the spirit of the broader aspects of the general inventive concept embodied in this application without departing from the broader scope of the claimed disclosure.
Claims
1. A system, comprising: At least one processor; And A memory that stores instructions which, when executed by the at least one processor, cause the system to perform a set of operations, the set of operations including: Receiving a query; Generating an initial query result set; Providing the query and the initial query result set to a generative large language model; Receiving at least one additional query from the generative large language model; Executing the at least one additional query; Providing the results from the at least one additional query to the generative large language model; Receiving semantic search engine results from the generative large language model; and Providing the semantic search engine results.
2. The system according to claim 1, further comprising instructions for generating one or more alternative queries based on the query, and wherein generating the initial query result set comprises: Generating alternative query results based on the one or more additional queries.
3. The system according to claim 1, wherein the semantic search engine results are included in a summary generated by the generative large language model.
4. The system according to claim 3, wherein the format of the summary is determined based on the type of information included in the summary.
5. The system according to claim 3, wherein the format for the summary is determined based on a template provided to the generative large language model.
6. The system according to claim 1, wherein the summary includes one or more references, and wherein the one or more references link to one or more underlying data sources for the summary.
7. The system according to claim 1, further comprising: Determining an intent or task based on the received query, wherein the intent or task is provided to the generative large language model.
8. The system according to claim 7, further comprising an operation for using the generative large language model to determine whether additional information is needed, wherein the determination is based on the intent or task.
9. The system according to claim 7, wherein the at least one additional query is generated by the generative large language model when it is determined that additional information is needed.
10. A method for generating semantic search engine results, the method comprising: Receiving a query; Generating an initial query result set; Providing the query and the initial query result set to a generative model; Using the generative model to determine that additional information is needed; Receiving at least one additional query from the generative model; Executing the at least one additional query; Providing the results from the at least one additional query to the generative large language model; Receiving semantic search engine results from the generative model; And Providing the semantic search engine results.
11. The method according to claim 10, further comprising: Analyzing the query to determine an intent or task based on the query, wherein analyzing the query includes: providing the query to at least one of the generative model or an alternative machine learning model.
12. The method according to claim 11 further comprises: Determining a format for the semantic search engine results, wherein the format is determined based on the query or the task.
13. The method according to claim 12 further comprises: Generating a prompt for the generative model, wherein the prompt is generated based on the format.
14. The method according to claim 13, wherein the prompt includes a template associated with the format, and the template defines the format for the semantic search engine results.
15. A computer storage medium comprising computer-executable instructions that, when executed by at least one processing unit, perform a method for generating semantic search engine results, the method comprising: Receiving a query; Generating an initial query result set; Providing the query and the initial query result set to a generative large language model; Using the generative model to determine that additional information is needed; Receiving at least one additional query from the generative large language model; Executing the at least one additional query; Providing the results from the at least one additional query to the generative large language model; Receiving semantic search engine results from the generative large language model; And Providing the semantic search engine results.