Systems and methods for real-time search-based generative artificial intelligence
The generative AI platform addresses the limitations of LLMs by integrating real-time search to generate timely and accurate responses, enhancing user interaction with up-to-date information and suggestions.
Patent Information
- Application Number
- JP2025502561
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-22
- Filing Date
- 2023-07-18
- Publication Date
- 2025-08-20
AI Technical Summary
Existing large-scale language models (LLMs) like ChatGPT and GPT-4 are limited by their reliance on pre-training and fine-tuning with curated datasets, failing to provide real-time knowledge in responses to user queries, especially for time-sensitive information.
A generative AI platform that integrates real-time search capabilities, allowing it to retrieve and aggregate up-to-date information from various data sources to generate responses to user inputs, incorporating visual and textual content, and provide citations and suggestions.
Enables the generation of concise, accurate, and timely responses to user queries, reducing the time and effort required to find and verify information, while offering suggestions for further exploration.
Smart Images

Figure 2025527146000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application is a non-provisional application of commonly assigned and co-pending U.S. Provisional Application No. 63 / 390,134, filed July 18, 2022, and U.S. Provisional Application No. 63 / 476,917, filed December 22, 2022, and claims priority under 35 U.S.C. § 119, each of which is expressly incorporated herein by reference in its entirety.
[0002] [Technical field] Embodiments relate generally to search engines and machine learning systems, and more particularly to real-time based generative artificial intelligence (AI) conversational applications. [Background technology]
[0003] Generative AI techniques have recently grown in assisting intelligent agents in conducting conversations with human users. For example, large-scale language models (LLMs) such as ChatGPT, GPT-4, and Google Bird provide conversational AI platforms for performing numerous natural language processing (NLP) tasks. However, these LLMs typically require multiple stages of pre-training, training, and fine-tuning using carefully curated training datasets to perform NLP tasks. In other words, these LLMs often only extract the underlying knowledge upon which their NLP output is delivered from the existing knowledge of the training data they are exposed to. [Brief explanation of the drawings]
[0004] [Figure 1A] FIG. 12 is a simplified diagram illustrating data flow between entities implementing the processes described in FIGS. 2-11 according to one embodiment described herein. [Figure 1B]FIG. 12 is a simplified diagram illustrating data flow between entities implementing the processes described in FIGS. 2-11 according to another embodiment described herein. [Figure 2] 1A and 1B, according to one embodiment described herein. FIG. 2B is a simplified diagram illustrating a computing device 200 that implements the text generation server 110 described in FIGS. [Figure 3] 3 is a simplified diagram illustrating a neural network structure implementing the text generation module 230 depicted in FIG. 2 according to one embodiment described herein. [Figure 4] 3 is a simplified diagram illustrating an example architecture of the NL pre-processing module 231 and the search module 232 shown in FIG. 2 according to embodiments described herein. [Figure 5] 3 is a simplified block diagram illustrating an example architecture of the generation sub-module 233 shown in FIG. 2 according to embodiments described herein. [Figure 6] FIG. 1C is a simplified block diagram of a networked system suitable for implementing the customized generative AI platform framework described in FIGS. 1A and 1B and other embodiments described herein. [Figure 7] FIG. 7 is an exemplary logic flow diagram illustrating a method for customized searching based on the framework and architecture illustrated in FIGS. 1-6, according to certain embodiments described herein. [Figure 8] FIG. 7 is an exemplary logic flow diagram illustrating a method for customized searching based on the framework and architecture shown in FIGS. 1-6, according to certain embodiments described herein. [Figure 9] FIG. 7 is an exemplary logic flow diagram illustrating a method of a generative AI system based on the framework shown in FIGS. 1-6, according to some embodiments described herein. [Figure 10A] FIG. 1 is an exemplary UI diagram illustrating an embodiment implementing a generative AI system integrated into a search system, according to embodiments described herein. [Figure 10B] FIG. 1 is an exemplary UI diagram illustrating an embodiment implementing a generative AI system integrated into a search system, according to embodiments described herein. [Figure 10C] FIG. 1 is an exemplary UI diagram illustrating an embodiment implementing a generative AI system integrated into a search system, according to embodiments described herein. [Figure 10D] FIG. 1 is an exemplary UI diagram illustrating an embodiment implementing a generative AI system integrated into a search system, according to embodiments described herein. [Figure 10E] FIG. 1 is an exemplary UI diagram illustrating an embodiment implementing a generative AI system integrated into a search system, according to embodiments described herein. [Figure 11] 10A-10C are exemplary UI diagrams illustrating an embodiment of implementing a text generation tool based on set user parameters, according to embodiments described herein.
[0005] Embodiments of the present disclosure and their advantages are best understood by referring to the following detailed description, in which like reference numerals are used to identify like elements shown in one or more of the figures, and it should be understood that the designations therein are for the purpose of illustrating, but not limiting, embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0006] [Detailed Description of the Invention] This application relates generally to search engines and machine learning systems, and more particularly to systems and methods for search assistance tools based on language models.
[0007] As used herein, the term "network" may comprise any hardware or software-based framework, including any artificial intelligence network or system, neural network or system, and / or any training or learning model implemented thereon or therewith.
[0008] As used herein, the term "module" may include a hardware or software-based framework that performs one or more functions. In some embodiments, a module may be implemented on one or more neural networks.
[0009] Generative AI techniques and LLMs have recently grown in assisting intelligent agents in conducting conversations with human users. However, existing LLMs can generally only provide generative output based on knowledge acquired from previous training data. For example, if a user inputs a question requiring real-time knowledge, such as "Which stocks should I buy today?", chat agents based on these LLMs, such as ChatGPT, cannot provide a satisfactory answer because they cannot obtain the latest time-varying information about the stock market.
[0010] Search engines, on the other hand, allow users to submit search queries and return search results in response. Users utilize search capabilities to learn about topics, stay on top of the news, conduct research, and accomplish other tasks. In this way, search results can provide significant utility when preparing written content, such as writing an essay or paper. For example, a user may want to write a summary highlighting Abraham Lincoln and the Civil War. However, it can be time-consuming for a user to review hundreds, thousands, millions, or even more search results returned from a search query. While traditional search engines provide utility in performing research on specific topics and aggregating a wide range of information, reviewing search results, identifying relevant information, and composing materials incorporating information from the search can be a time-consuming endeavor. In light of the need for improved generative AI systems that reflect time-changing information, such as real-time news events, the embodiments described herein provide systems and methods for a generative AI platform based on real-time search. Specifically, in response to user input to perform an NLP task, the generative AI platform may trigger a real-time search based on the user input and then aggregate the search results to generate an output. The search may be performed as an internet web search based on web indexing and / or a dedicated search from a small number of selected data sources. In one embodiment, the generative AI platform may analyze the content following the web links in the search results and use the analyzed content as part of the input to generate a text response in response to the user input.
[0011] In another embodiment, the generative AI platform may be equipped with a visual language model such that the visual language model can obtain and generate summaries of visual content (such as images or video content) from retrieved web links and use such summaries as input to generate text responses in response to user input.
[0012] For example, specialized output may be generated based on a user input query to summarize results, answer questions, create emails, essays, newsletters, posts, etc., and efficiently extract information for the user. In one embodiment, the generative AI system employs machine learning modules to receive plain text input, generate conversational output, and provide citations from search results that support the generated content. In this way, instead of having to visit and review many web pages, the user can receive a concise answer to their query with citations that support and confirm the output, thus reducing the time to find an answer and providing certainty in correctness. In some embodiments, the generative AI system may provide suggested follow-up responses for the user to generate additional conversational responses, as well as suggested search results, apps, or other information that may further assist the user.
[0013] Another example is a text generation-based search tool that provides text based on user-provided parameters. Specifically, the user can provide parameters such as the use case, desired tone, target audience, and subject matter of the generated text. The system employs a machine learning module to receive the user parameters, perform a search based on the subject, and generate text output incorporating the results and the user parameters. The generative AI system may further provide citations or related search results to allow the user to verify the information and conduct further research. In some embodiments, the generative AI system may further provide suggestions for additional topics, alternative parameters that may be useful to the user, or other information that may assist the user in using the generated text or generating further text.
[0014] In this way, the generative AI system can generate text output based on real-time search results that reflect the most up-to-date information according to the user query, thus improving generative AI technology.
[0015] 1A is a simplified diagram illustrating data flow between entities implementing the processes described in FIGS. 2-11 , according to one embodiment described herein. A user 130 interacts with a user device 120, for example, through providing natural language (NL) input 126, and the user device 102 interacts with a text generation server 110 hosting one or more natural language processing (NLP) models 115 through input 122. The input 122 may include user NL input 126, such as a user question, user parameters, user context, or other context. User parameters may include any user-configured parameters for a natural language task, such as the intended audience, the type of task (e.g., creating an email, abstract, research article, and / or the like), the tone of the output, and / or the like. User context may include any user preferences, such as the user's preferred search data sources, user activity on social media, and / or the like.
[0016] In one embodiment, the text generation server 110 interacts with various data sources 103a-n (collectively referred to as 103). For example, the data sources 103a-n may be any number of available databases, web pages, servers, blogs, content providers, cloud servers, and / or the like.
[0017] In one embodiment, upon receiving the input 122, the text generation server 110 may determine whether a real-time search is required and / or which data sources may be searched. In one embodiment, the text generation server 110 may employ at least the NLP model 115 to generate a search query; for example, in one implementation, the text generation server 110 may extract key terms from the text input 126 as a search query. As another example, in one implementation, the text generation server 110 may generate text embeddings and perform vector search based on the text embeddings. In another example, the text generation server 110 may perform a combined search based on both the text query and the vector embeddings.
[0018] In one embodiment, the text generation server 110 may employ at least an NLP model 115 to generate the NLP output.
[0019] In one embodiment, the text generation server 110 may engage a neural network-based AI model 115 to predict relevant data sources for a search based on user input. Further details of determining specific data sources based on a search query can be found in connection with FIG. 5 and in co-pending and commonly assigned U.S. Non-Provisional Application No. 17 / 981,102, filed November 4, 2022.
[0020] In another example, when the text generation server 110 receives the input 122 from the user device 120, the text generation server 110 may determine predefined data sources as being relevant to keywords, phrases, topics, parameters, or other elements of the search-related input 122. The determined data sources may further be influenced by previous user interactions, such as, for example, the user disapproving of search results from a particular data source, the user pre-setting preferred data sources, and / or the like.
[0021] In one embodiment, upon receiving input 122, text generation server 110 can determine what search queries to generate, how many search queries to generate, and other parameters related to performing a search. For example, as further described in connection with FIGS. 2-6 , text generation server 110 can host one or more neural network-based prediction modules. The prediction modules can generate queries based on keywords or phrases in input 122, user parameters, and / or other contextual information when the prediction modules determine that a search should be performed on a particular element from input 122.
[0022] In one embodiment, the text generation server 110 can then convert the search query into a customized search query 111 a-n that conforms to the specific formatting requirements of each data source 103 a-n. The customized search query 111 a-n is sent to each data source 103 a-n via its respective API 112 a-n. In response, the data sources 103 a-n can return query results 113 a-n to the text generation server 110 in the form of links to web pages and / or cloud files.
[0023] Instead of simply presenting links to search results (e.g., web pages) to the user device 120, the text generation server 110 can utilize one or more NLP models 115 to extract information from the search results, generate text based on the search results and any parameters specified in the input 122, generate a natural language response, and return it as NL output 125 for display on the user device 120.
[0024] For example, at least one NLP model 115 may be used to analyze web content according to links provided in search results 113a-n. In one embodiment, at least one NLP model 115 may generate a summary of the web content from at least one search result 113a and use such generated summary to generate final NL output 125.
[0025] In one embodiment, at least the NLP model 115 can be used to generate the output NL output 125 based on parsed content from the search results 113a-n and / or additional user-configured parameters. For example, the user-configured parameters can specify the type of NLP task (e.g., composing an email, creating a legal memorandum, conducting a conversation, and / or the like), the intended audience of the NL output 125, the tone of the NL output, and / or the like.
[0026] Note that the text generation server 110, NL output 125, and / or NLP model 115 are for illustrative purposes only. The framework 100 may be applied to any type of generation and / or generative model and / or to generating any type of output, such as, but not limited to, a code segment, an image, etc.
[0027] For example, input 122 can take various formats, such as text input, audio input, image input, video input, and / or the like. For example, input 122 can include two images and a text question such as, "Which photo was taken at the Berkeley Center in New York City on the BTS 2023 tour?" An image encoder can then be used in conjunction with NLP model 115 to encode the image input and facilitate searches based on the image encoding. In another implementation, a captioning model can be employed to generate captions for input images such that server 110 can perform searches using the text captions of the input images.
[0028] As another example, input 122 may include an audio clip (e.g., of a musical piece). Server 110 may generate an audio signature from the audio clip and perform a search, for example, in a data source storing a library of musical pieces. Upon receiving the name and / or title of the musical piece associated with the audio clip from the data source, server 110 may further generate a search based on the obtained name and / or title of the musical piece to obtain search results related to the musical piece. For example, a musical clip recorded by a user in Times Square, New York, may be uploaded to the production server, which may identify the musical clip as belonging to the Broadway musical The Phantom of the Opera and generate output (with both text and / or images related to the music) including a short description of the music and / or available schedules and tickets for the music, as well as links to purchase them.
[0029] For example, a visual language model may be used in conjunction with or in place of the NLP model 115 in the text generation server 110 to analyze image and / or video content from at least one search result 113b and generate a summary of the image and / or video content. Such summarized multimedia content may be used to generate the NL output 125. As another example, the visual language model and / or other multimodal models used in the text generation server 110 may generate text captions for images retrieved from a web page following a search result link. The text captions may be fed to the NLP model along with other text input from the search result to generate the NL output 125.
[0030] As another example, when user input 122 relates to a coding query such as "What are the differences for compiling a listing in Python and C#?", text generation server 110 can perform a code search and generate output based on the search results 113a-n. The output can include a Python code segment and a C# code segment, as well as a text portion explaining the differences between the two code segments. Additional details regarding performing a code search to obtain code segments can be found in co-pending and commonly assigned U.S. Non-Provisional Application No. 18 / 330,225.
[0031] As another example, the generation server 110 may further employ an NLP model 115 as a code generation model to generate code segments based on the search results 113a-n.
[0032] In another example, the generation server 110 may further insert one or more images, illustratively taken from a web page following a search result link, into the NLP output 125. For example, if the input 122 includes a request to "write a passage on the history of direct current versus alternating current," in addition to generating a text summary based on various search results, the NLP model 115 may further insert web images of Thomas Edison and Nikola Tesla into the output 125 for illustration.
[0033] As another example, the generation server 110 may further employ an image generation model that can generate image content based on textual and / or image content retrieved from the search results 113a-n to form the output 125. For example, if the input 122 includes the question "Why is Hillary Step on Everest famous?", the image generation and / or editing model may edit a photo of Hillary Step retrieved from the search results 113a-n by adding measurement labels indicating height, slope, temperature, wind speed, etc. overlaid on the photo, and use the edited photo in the generated output 125 for illustrative purposes.
[0034] In one embodiment, the NL output 125 may include references to the data sources 103 a-n in relevant portions based on the corresponding search results 113 a-n, respectively. In this way, the NL output 125 automatically includes the reference authority.
[0035] In one embodiment, a client component on the user device 120 may display the NL output 125 via a user interface. For example, the NL output 125 may be displayed in a side panel within a search browser, as shown in Figures 10A-10E. In another example, the NL output 125 may be displayed in a mobile UI on the mobile user device 120 in the form of a conversation with a cloud agent.
[0036] Figure 1B is a simplified diagram illustrating data flow between entities implementing the processes described in Figures 2-11, according to another embodiment described herein. As shown in Figure 1B, the text generation server 110 may communicate with several external LLMs 116a-n housed on external servers instead of and / or in addition to hosting its own NLP models 115 as shown in Figure 1A.
[0037] In one embodiment, the LLMs 116a-n may be housed on an external server accessible by the text generation server 110 over a network. In another embodiment, the LLMs 116a-n (or copies thereof) may be stored on the text generation server 110. The text generation server 110 may communicate with the LLMs 116a-n via their respective APIs 117a-n. In one embodiment, the text generation server 110 processes the input 122 and interacts with the data sources 103a-n to obtain search results 113a-n, similar to the description with respect to FIG. 1A . Once the text generation server 110 receives the results 113a-n from the data sources 103a-n, the text generation server 110 can generate NLP input including the search results 113a-n and send the NLP input to one or more LLMs 116a-n via their respective APIs 117a-n. For example, in one embodiment, the NLP input to one or more LLMs 116a-n may be a concatenation of search results 113a-n including their respective links, and / or a prompt indicating the type of NLP task.
[0038] In one embodiment, the text generation server 110 may select an LLM from the candidate LLMs 116a-n to forward the NLP request to depending on the type of NLP task. For example, the LLM 116a may be used to create an article, while another LLM 116b may be selected to generate system responses within a conversation. The text generation server 110 may then generate an NL output 125 based on the output from the LLMs and return it to the user device 120 for display. In one embodiment, the text generation server 110 may select an LLM from the candidate LLMs 116a-n depending on the type of data source from which the search results are obtained. For example, if the input 122 queries, "What are the latest tours for BTS?", the search module (e.g., 232 in FIG. 2) of the generation server 110 may determine to prioritize searching social media sources (e.g., for the query terms "tour," "BTS"). The generation server 110 may then determine that the search results 113a-n may contain a large amount of image and / or video content due to the nature of the data source. The production server 110 may then select an LLM capable of processing the visual data to forward the request to produce the output.
[0039] 2 is a simplified diagram illustrating a computing device 200 implementing the text generation server 110 depicted in FIGS. 1A and 1B , according to one embodiment described herein. As shown in FIG. 2, computing device 200 includes a processor 210 coupled to a memory 220. The operation of computing device 200 is controlled by processor 210. Also, while computing device 200 is shown as having only one processor 210, it will be understood that processor 210 may represent one or more central processing units, multi-core processors, microprocessors, microcontrollers, digital signal processors, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), graphics processing units (GPUs), etc. within computing device 200.
[0040] Computing device 200 may be implemented as a standalone subsystem, as a board added to a computing device, and / or as a virtual machine. In various embodiments, the communication device may comprise a personal computing device (e.g., a smartphone, a computing tablet, a personal computer, a laptop, a wearable computing device such as eyeglasses or a watch, a Bluetooth device, etc.) capable of communicating with a network. A service provider may utilize a network computing device (e.g., a network server) capable of communicating with a network. It should be understood that each of the devices utilized by users and service providers may be implemented as computer system 200 as follows.
[0041] Memory 220 may be used to store software executed by computing device 200 and / or one or more data structures used during operation of computing device 200. Memory 220 may include one or more types of machine-readable media. Some common forms of machine-readable media may include a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, and / or any other medium adapted to be read from by a processor or computer.
[0042] Processor 210 and / or memory 220 may be arranged in any suitable physical arrangement. In some embodiments, processor 210 and / or memory 220 may be implemented on the same substrate, in the same package (e.g., a system-in-package), on the same chip (e.g., a system-on-chip), etc. In some embodiments, processor 210 and / or memory 220 may comprise distributed, virtualized, and / or containerized computing resources. According to such embodiments, processor 210 and / or memory 220 may be located in one or more data centers and / or cloud computing facilities.
[0043] In some examples, memory 220 may include a non-transitory tangible computer-readable medium containing executable code that, when executed by one or more processors (e.g., processor 210), may cause the one or more processors to perform the methods described in further detail herein. For example, as shown, memory 220 includes instructions for search platform module 230, which may be used to implement and / or emulate the systems and models and / or to implement any of the methods described further herein. Search platform module 230 may receive input 240, such as an input search query (e.g., a word, sentence, or other input provided by a user), user parameters, and / or the like, via data interface 215, and generate output 250, which may be a conversational text-based response presented in elements, along with supporting links to different data sources, suggestions for future user input, suggestions for further user exploration, and / or the like. Examples of input data may include any of the inputs 122 in FIG. 1 and described in more detail with respect to FIG. 4, such as a user query 402, a user context 404, and / or other context 406 such as parameters, current events, and the like.
[0044] The data interface 215 may comprise a communications interface, a user interface (such as a voice input interface, a graphical user interface, etc.). For example, the computing device 200 may receive input 240 (such as a training data set) from a networked database via the communications interface. Alternatively, the computing device 200 may receive input 240 from a user via the user interface, such as a user-entered search query or parameters.
[0045] In some embodiments, the text generation module 230 is configured to generate text-based conversational responses to a user device (e.g., 120 in FIGS. 1A and 1B). The text generation module 230 may further include an NL preprocessing submodule 231, a search submodule 232, a generation submodule 233, and (optionally) an LLM interface submodule 234. The NL preprocessing submodule 231 may perform processing steps to perform searches, generate NL output, or assist other steps performed by the text generation module 230. For example, the NL preprocessing submodule 231 may tokenize and generate search queries based on received input, user parameters, or other information (e.g., 122 in FIGS. 1A and 1B). The NL preprocessing submodule 231 may also determine how many searches to perform and which words, topics, or phrases to search for. The search sub-module 232 may determine one or more data sources for the search based on, for example, a user's preference configuration, a user's past behavior indicating preferences, a search query, a source type, and / or the like. The search sub-module 232 may further generate a customized query according to each data source, send the customized query to a corresponding API (e.g., 112a-n in FIGS. 1A and 1B), and receive search results from the API. Further details of the operation of the search sub-module 232 can be found in FIG. 5 and in co-pending and commonly assigned U.S. Non-Provisional Application No. 17 / 981,102, filed November 4, 2022. The generation sub-module 233 may generate text-based output for the user based on the received and processed search results, user input, and information determined by the NL pre-processing sub-module 231. In some embodiments, the generation sub-module finalizes and generates output based on the data received and processed by the NL pre-processing sub-module 231 and the search sub-module 232. In other embodiments, the generation sub-module 233 generates a result as output 250 based on data received from the optional LLM interface sub-module 234 .The LLM interface sub-module 234 may interface the text generation module 230 with an external LLM, prepare input for an API associated with the external LLM, prepare prompts based on search results received by the search sub-module 232, prepare prompts based on input processed by the NL pre-processing sub-module 231, and / or the like. The LLM interface sub-module 234 may also process results received from the external LLM and provide them to the generation sub-module 233. Additional functionality of the search sub-module 232 may be further described in connection with FIG. 4.
[0046] Some examples of computing devices, such as computing device 200, may include non-transitory tangible computer-readable media containing executable code that, when executed by one or more processors (e.g., processor 210), may cause the one or more processors to perform the processes of a method. Some common forms of machine-readable media that may contain the processes of a method are, for example, a floppy disk, a flexible disk, a hard disk, magnetic tape, any other magnetic medium, a CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with a pattern of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, and / or any other medium adapted to be read by a processor or a computer.
[0047] 3 is a simplified diagram illustrating a neural network structure implementing the text generation module 230 depicted in FIG. 2, according to one embodiment described herein. In one embodiment, the text generation module 230 and / or one or more of its submodules 231-234 may be implemented via the artificial neural network structure shown in FIG. 3. A neural network comprises a computing system built on a collection of connected units or nodes called neurons (e.g., 244, 245, 246). The neurons are often connected by edges, and adjustable weights (e.g., 251, 252) are often associated with the edges. The neurons are often organized into layers, such that different layers may perform different transformations on their respective inputs and output the transformed input data onto the next layer.
[0048] For example, a neural network architecture may include an input layer 241, one or more hidden layers 242, and an output layer 243. Each layer may include multiple neurons, and the neurons between layers are interconnected according to a specific topology of the neural network topology. The input layer 241 receives input data (e.g., 210 in FIG. 2), such as a user input query (e.g., 122 in FIGS. 1A and 1B), user input preferences, and / or the like. The number of nodes (neurons) in the input layer 241 may be determined by the dimensionality of the input data (e.g., the length of a vector providing examples of the input). Each node in the input layer represents a feature or attribute of the input.
[0049] The hidden layer 242 is an intermediate layer between the input layer and the output layer of a neural network. Note that two hidden layers 242 are shown in FIG. 5 for illustrative purposes only, and any number of hidden layers may be utilized in a neural network structure. The hidden layer 242 may extract and transform input data through a series of weighted calculations and activation functions.
[0050] For example, as illustrated in FIG. 2 , the text generation module 230 receives input 210, including user input, parameters, and / or other information, and transforms the input into output 250 of a text-based response tailored to the input. To perform the transformation, each neuron receives the input signal, performs a weighted sum of the input according to the weight assigned to each connection (e.g., 251, 252), and then applies the activation function (e.g., 261, 262, etc.) associated with the respective neuron to the result. The output of the activation function is passed to the next layer of neurons or serves as the final output of the network. The activation functions may be the same or different across different layers. Exemplary activation functions include, but are not limited to, sigmoid, hyperbolic tangent, rectified linear unit (ReLU), leaky ReLU, Softmax, and / or the like. In this way, after several hidden layers, the input data received at the input layer 241 is transformed into somewhat different values that exhibit data characteristics corresponding to the task the neural network structure is designed to perform.
[0051] The output layer 243 is the final layer of the neural network structure. It generates the network's output or prediction based on the calculations performed in the preceding layers (e.g., 241, 242). The number of nodes in the output layer depends on the nature of the task being addressed. For example, in a binary classification problem, the output layer may consist of a single node representing the probability of belonging to one class. In a multi-class classification problem, the output layer may have multiple nodes, each representing the probability of belonging to a particular class.
[0052] Thus, the text generation module 230 and / or one or more of its submodules 231-234 may comprise a transformable neural network structure of layers of neurons, with weights and activation functions describing the nonlinear transformations in each neuron. Such neural network structures are often implemented on one or more hardware processors 210, such as graphics processing units (GPUs).
[0053] In one embodiment, the text generation module 230 and its submodules 231-234 may be implemented by hardware, software, and / or a combination thereof. For example, the text generation module 230 and its submodules 231-234 may include specific neural network structures implemented and executed on various hardware platforms 550, such as, but not limited to, CPUs (central processing units), GPUs (graphics processing units), FPGAs (field-programmable gate arrays), application-specific integrated circuits (ASICs), dedicated AI accelerators such as TPUs (tensor processing units), and dedicated hardware accelerators specifically designed for the neural network computations described herein. Exemplary specific hardware for neural network structures may include, but is not limited to, Google Edge TPUs, Deep Learning Accelerators (DLA), NVIDIA AI-focused GPUs, etc. The hardware 550 used to implement the neural network structures is specifically configured depending on factors such as the complexity of the neural network, the scale of the task (e.g., training time, input data scale, size of the training dataset, etc.), and the desired performance.
[0054] In one embodiment, the neural network-based text generation module 230 and one or more of its sub-modules 231-234 may be trained by iteratively updating the neural network's underlying parameters (e.g., bias parameters and / or coefficients in activation functions 261, 262 associated with neurons, such as weights 251, 252) based on a loss objective. For example, during forward propagation, training data, such as past encoding activity, is fed to the neural network. Data flows through the network's layers 241, 242, each performing calculations based on its weights, biases, and activation functions until the output layer 243 produces the network's output 250, such as generated text.
[0055] The output produced by the output layer 243 is compared to an expected output (e.g., a corresponding "ground truth" that provides examples of ground truth labels), e.g., actual text from the training data. For example, the loss function may be cross-entropy, mean squared error (MSE), and / or the like. Given the loss, the negative gradient of the loss function is calculated for each weight in each layer individually. Such negative gradients are calculated iteratively backward from the last layer 243 of the neural network to the input layer 241, one layer at a time. These gradients quantify the sensitivity of the network's output to parameter changes. A chained approach is applied to efficiently calculate these gradients by propagating gradients backward from the output layer 243 to the input layer 241.
[0056] The neural network parameters are updated backward (backpropagation) from the last layer to the input layer based on the calculated negative gradient using an optimization algorithm to minimize loss. Backpropagation from the last layer 243 to the input layer 241 may be performed for several training samples over several iterative training epochs. In this way, the neural network parameters may be gradually updated in a direction that produces less or minimized loss, indicating that the neural network has been trained to produce predicted output values closer to the target output values with improved prediction accuracy. Training may continue until a stopping criterion is met, such as reaching a maximum number of epochs or achieving satisfactory performance on validation data. At this point, the trained network can be used to make predictions on new, insightful data, such as specific parameters for user queries and responses.
[0057] Thus, the training process transforms the neural network into an "updated" trained neural network with updated parameters such as weights, activation functions, and biases. The trained neural network thus improves neural network technology in cloud-based generative AI systems.
[0058] Figure 4 is a simplified diagram illustrating an example architecture of the NL pre-processing module 231 and the search module 232 shown in Figure 2, according to an embodiment described herein. The NL pre-processing module 231 and / or the search module 232 may comprise software and hardware platforms implemented in the text generation server 110 of Figures 1A and 1B and / or the server 330 of Figure 6. For example, the text generation module 230 may be realized based on the neural network structure shown in Figure 3.
[0059] In one embodiment, the NL preprocessing module 231 receives input data. This input data may include one or more of natural language user input 402 (similar to NL input 126 in FIGS. 1A and 1B ), user context 404, and other context 406. The natural language user input 402 may include a word, multiple words, a sentence, or any other type of search query provided by a user performing a search using the platform 230. In some embodiments, the natural language user input 402 may be in the form of a conversational statement or question. For example, the natural language user input 402 may be search terms such as "what is a pseudo-recurrent neural network" or "who is Richard Socher." The user context 404 may include input describing the user, including a user ID, user preferences, user click logs, or other information collected or provided by the user, preferred data sources selected by the user, the user's past activity of "liking" or "disliking" search results or search sources, and / or the like. Other contexts 406 may include inputs representing other useful input information, including information about world events, searches made around the same time period, searches made in the same area or surrounding areas, searches that have increased in volume over a period of time, or other potential contextual information that may assist platform 410 in providing the user with an appropriate search response. Other contexts 406 may further include input selections that a user includes as typed input, such as desiring to have a response that is "professional," "causal," "comedic," etc., to indicate a desired output such as "paragraph," "single sentence," "paper," etc.
[0060] In one embodiment, other context 406 may include user-configured generation parameters such as type of output (e.g., paragraph, email, legal memorandum, article, conversational response, etc.), intended audience (e.g., elementary school students, professionals, friends, social media, etc.), tone (e.g., informational, persuasive, analytical, etc.), length, etc.
[0061] In one embodiment, the NL preprocessing module 231 can concatenate input information, such as natural language user input 402, user context 404, and other context 406, into an input sequence of tokens to generate one or more predictive text queries. Prediction can be performed when a new natural language user input 402 is received. In some embodiments, the NL preprocessing module 231 can further take into account previous natural language user input 402, user context 404, and other context 406 received from the same or different users to generate one or more predicted text queries. Furthermore, in some embodiments, the NL preprocessing module 231 may instead reuse a previous predicted text query based on similarity between the current input information and the previous input information.
[0062] The NL preprocessing module 231 may be trained on a dataset of previous natural language user input 402, previous user contexts 404, and / or previous other contexts 406, and corresponding ground truth queries.
[0063] The search module 232 can receive one or more search queries from the NL preprocessing module 231 and then determine a list of data sources for the search. In one implementation, the search module 232 can retrieve a predefined list of data sources pre-categorized based on the subject matter determined by the NL preprocessing module 231. In another implementation, the search module 232 can use a prediction module 431 to predict prioritized data sources for the search based on a concatenation of natural language user input 402, user context 404, and / or other contextual information 406, in a manner similar to that described in co-pending, commonly assigned U.S. Non-Provisional Application No. 17 / 981,102, filed November 4, 2022.
[0064] The search module 232 and its corresponding search module 433 can then send one or more search queries customized for each identified data source to the respective search API 422a-n and receive a list of search results from the respective search API 422a-n.
[0065] In some embodiments, the ranking module 434 may optionally rank the list of search apps 422a-n for performing a search. Each search app 422a-n corresponds to a particular data source 103a-n in FIGS. 1A and 1B. For example, upon receiving an input query related to coding, the ranking module 434 may select data sources such that the search app 422a corresponds to a search application configured to search in the “StackOverflow” database and the search app 422b corresponds to a search application configured to search in the “Tutorial Point” database. The ranking module 434 scores the multiple search apps 422a-n by running the input sequence through a neural network model once for each search app 422a-n using the input sequence processed by NL preprocessing 231, including the natural language user input 404, the user context 404, and other context 406. In this manner, the ranking module 434 can rank search results from the list of data sources via the search APIs 422a-n.
[0066] In some embodiments, if a user has indicated a prior preference for search results from particular sources, as reflected in user context 404, the ranking module 434 may rank search results from those particular sources higher than others. Additionally, other context 406 may indicate that other users value particular sources related to identified terms in the user query 402, and thus the ranking module 434 may incorporate this other context when ranking the search results. In some embodiments, the ranking module 434 may prioritize sources so as to reuse sources in subsequent follow-up responses related to a previous user query 402 from a previous related conversation with the same user.
[0067] Search results from the search APIs 422a-n are often in the form of links to web pages or cloud files within the respective data sources. A ranked list of search results may be passed from the ranking module 434 to the generation module 432.
[0068] The generation module 432 can follow links in the search results and extract information from the content of the web pages or cloud files. The information can then be processed by the generation module 432 and incorporated into a generated text response that is provided to the user as results 430. In one implementation, the generation module 432 can further incorporate links to web pages or cloud files, where the information is utilized in the results 430 as citations to support or provide further information to the user.
[0069] For example, the results 430 are sent to a user device for display via a graphical user interface or some other type of user output device. The results 430 are primarily text-based responses, but may further incorporate related search apps, such as to provide visual depictions of related data in addition to the text responses (e.g., to show weather information, stock charts, and / or the like). The results 430 may further include suggested follow-up inputs provided by the generation module 432, suggested URLs ranked by the ranking module 434 and provided by the generation module 432, or any other information of interest to the user. In other embodiments, the results 430 are transmitted to the generation sub-module 233 for processing and preparation for output to the user, as further described with respect to FIG. 5 .
[0070] Figure 5 is a simplified block diagram illustrating an example architecture of the generation sub-module 233 shown in Figure 2, according to an embodiment described herein. The generation sub-module 233 may comprise software and hardware platforms implemented in the text generation server 110 of Figures 1A and 1B and / or the server 330 of Figure 6. For example, the generation module 233 may be implemented based on the neural network structure shown in Figure 3.
[0071] In one embodiment, the generation sub-module 233 (similar to 233 in FIG. 2) receives one or more inputs 501 a-n, such as search input 501 a from a data source (e.g., search results 113 a-n in FIGS. 1A-1B), user input 501 b from the NL pre-processing sub-module 231 (e.g., the original user-asked question and / or user-entered parameters), and context input 501 n (e.g., previous conversation context). These inputs may further include search results, links to search results, web content, photos, videos, PDF documents, natural language user input, user context, other context, and / or other data or related information.
[0072] The generation sub-module 233 may process various data inputs and generate an output 503. For example, the output 503 may be similar to the NL output 125 of Figures 1A-1B.
[0073] In one implementation, the generation sub-module 233 may further use educational prompts including user-configured parameters such as type of output (e.g., email, legal memorandum, passage, news article, conversation, etc.), intended audience (e.g., expert, social media, friend, education, etc.), tone (e.g., informational, persuasive, warning, etc.) to guide content generation. In one implementation, the generation sub-module 233, which may comprise one or more NLP and / or multimodal models, may be trained on a corpus of text documents annotated with tone and / or intended audience. In one embodiment, during training, the prompts that guide the generation sub-module 233 to generate relevant text according to the user-configured parameters may be updated accordingly.
[0074] For example, the generation sub-module 233 may prepare a summary of the search results, generated text, images, sounds, or other content to be delivered to the user. In some embodiments, the generation sub-module 233 may insert a link to the search results within the generated context to serve as a citation to the generated content.
[0075] In some embodiments, the generation sub-module 233 operates in conjunction with the optional LLM sub-module 234 to request and process data from an external LLM based on the search results or generated text. In other embodiments, the functionality of the optional LLM sub-module 234 may be integrated into the generation sub-module 233.
[0076] 6 is a simplified block diagram of a networked system suitable for implementing the customized generative AI platform framework described in FIGS. 1A and 1B and other embodiments described herein. In one embodiment, block diagram 300 illustrates a system including a user device 310 that may be operated by a user 340, data vendor servers 345a and 345b-n, a server 330, and other forms of devices, servers, and / or software components that operate to perform various methodologies according to described embodiments. Exemplary devices and servers may include devices similar to computing device 200 described in FIG. 2, standalone, and enterprise-class servers that run an OS such as a MICROSOFT® OS, a UNIX® OS, a LINUX® OS, or other suitable device- and / or server-based OS. It should be understood that the devices and / or servers shown in Figure 3 may be deployed in other manners, and that the operations performed by, and / or services provided by, such devices and / or servers may be combined or separated for a given embodiment, and may be performed by more or fewer devices and / or servers. One or more devices and / or servers may be operated and / or maintained by the same or different entities.
[0077] User device 310, data vendor servers 345a and 345b-345n, and server platform 330 (e.g., similar to search server 110 of FIG. 1) may communicate with each other via network 360. User device 310 may be utilized by a user 340 (e.g., a driver, a system administrator, etc.) to access various features available to user device 310, which may include processes and / or applications associated with server 330 for receiving output data anomaly reports.
[0078] User device 310, data sources 345a and 345b-345n, and platform 330 may each include one or more processors, memory, and other suitable components for executing instructions, such as program code and / or data stored on one or more computer-readable media, to implement the various applications, data, and steps described herein. For example, such instructions may be stored on one or more computer-readable media, such as memory or data storage devices, internal and / or external to the various components of system 300 and / or accessible via network 360.
[0079] User device 310 may be implemented as a communications device that may utilize appropriate hardware and software configured for wired and / or wireless communications with data source 345 and / or platform 330. For example, in one embodiment, user device 310 may be implemented as an autonomous vehicle, a personal computer (PC), a smartphone, a laptop / tablet computer, a wristwatch with appropriate computing hardware resources, glasses with appropriate computing hardware (e.g., GOOGLE GLASS®), other types of wearable computing devices, embedded communications devices, and / or an IPAD® from APPLE®, or other types of computing devices capable of transmitting and / or receiving data. While only one communications device is shown, multiple communications devices may function similarly.
[0080] 3 includes a user interface (UI) application 312 and / or other applications 316, which may correspond to executable processes, procedures, and / or applications with associated hardware. For example, the user device 310 may receive a NL response (e.g., 125 in FIGS. 1A and 1B) from the server 330 in the form of a text-based output with corresponding information and display the message via the UI application 312 (see, e.g., FIGS. 10A-10E). In other embodiments, the user device 310 may include additional or different modules with dedicated hardware and / or software, as needed.
[0081] In various embodiments, the user device 310 includes other applications 316 as may be desired in particular embodiments to provide functionality to the user device 310. For example, the other applications 316 may include security applications for implementing client-side security features, programmatic client applications for interfacing with appropriate APIs over the network 360, or other types of applications. The other applications 316 may also include communication applications, such as email, text, voice, social networking, and instant messaging applications, that allow the user to send and receive email, phone, text, and other notifications over the network 360. For example, the other application 316 may be an email or instant messaging application that receives predicted result messages from the server 330. As described in further detail below with respect to FIG. 10 , the server 330 may provide email, text messages, or other use-case-specific text responses for the user 340 based on queries that may be directly incorporated into the other applications 316 and sent or processed without further input by the user 340. The other applications 316 may include device interfaces and other display modules that may receive input and / or output information. For example, other applications 316 may include software programs executable by a processor for asset management that include a graphical user interface (GUI) configured to provide a user 340 with an interface for viewing and interacting with user-interactive elements that display text or other elements based on search results.
[0082] The user device 310 may further include a database 318 stored in temporary and / or non-transitory memory of the user device 310, which stores various applications and data and may be utilized during the execution of various modules of the user device 310. The database 318 may store a user profile for the user 340, predictions previously viewed or saved by the user 340, historical data received from the server 330, and / or the like. In some embodiments, the database 318 may be local to the user device 310. However, in other embodiments, the database 318 may be external to the user device 310 and accessible by the user device 310, including a cloud storage system and / or database accessible via the network 360.
[0083] The user device 310 includes at least one network interface component 319 adapted to communicate with data sources 345a and 345b-345n and / or server 330. In various embodiments, the network interface component 319 may include a DSL (e.g., digital subscriber line) modem, a PSTN (public switched telephone network) modem, an Ethernet device, a broadband device, a satellite device, and / or various other types of wired and / or wireless network communication devices, including microwave, radio frequency, infrared, Bluetooth, and near field communication devices.
[0084] Data sources 345a and 345b-345n correspond to servers hosting one or more of search applications 303a-n (or collectively referred to as 303) and can provide search results to server 330 that include web pages, posts, or other online content hosted by data sources 345a and 345b-345n. Search application 303 may be implemented by one or more relational databases, distributed databases, cloud databases, and / or the like. Search application 303 may be comprised by platform 330, data sources 345, or some other collection.
[0085] In one embodiment, one or more data sources 345a-n may be similar to data sources 103a-n. In one embodiment, one or more additional external servers hosting LLMs (e.g., 116a-n in FIG. 1B) may communicate with server 330 via network 360.
[0086] In one embodiment, platform 330 may allow various data sources 345a and 345b-345n to partner with platform 330 as new data sources. The generative AI system provides an application programming interface (API) for each data source 345a and 345b-345n to plug into the generative AI system's services. For example, the California Bar Association may register with the generative AI system as a data source. In this manner, the data source "California Bar Association" may appear in a list of available data sources on the generative AI system. A user may select or deselect California Bar Association as a preferred data source for a search. Similarly, additional data sources 345 may partner with platform 330 to provide additional data sources for a search so that a user can understand where search results are aggregated.
[0087] Data sources 345a-n (collectively referred to as 345) include at least one network interface component 326 adapted to communicate with user device 310 and / or server 330. In various embodiments, network interface component 326 may include a DSL (e.g., digital subscriber line) modem, a PSTN (public switched telephone network) modem, an Ethernet device, a broadband device, a satellite device, and / or various other types of wired and / or wireless network communication devices, including microwave, radio frequency, infrared, Bluetooth, and near field communication devices. For example, in one embodiment, data source 345 may transmit asset information from search application 303 to server 330 via network interface 326.
[0088] The platform 330 may be housed with the text generation module 230 and its sub-modules described in Figure 2. In some embodiments, the platform 330 may receive data from the search application 303 and / or the network interface 326 at the data source 345 via the network 360 to generate user-interactive elements incorporating the search results. The generated user-interactive elements may also be transmitted via the network 360 to the user device 310 for review by the user 340.
[0089] The database 332 may be stored in temporary and / or non-transitory memory of the server 330. In one embodiment, the database 332 may store data obtained from the data vendor server 345. In one embodiment, the database 332 may store parameters of the search platform model 230. In one embodiment, the database 332 may store user input queries, user profile information, search application information, search API information, or other information related to a search being performed or a previously performed search.
[0090] In some embodiments, database 332 may be local to platform 330. However, in other embodiments, database 332 may be external to platform 330 and accessible by platform 330, including cloud storage systems and / or databases accessible via network 360.
[0091] Platform 330 includes at least one network interface component 333 adapted to communicate with user devices 310 and / or data sources 345a and 345b-345n over network 360. In various embodiments, network interface component 333 may comprise a DSL (e.g., digital subscriber line) modem, a PSTN (public switched telephone network) modem, an Ethernet device, a broadband device, a satellite device, and / or various other types of wired and / or wireless network communication devices, including microwave, radio frequency (RF), and infrared (IR) communication devices.
[0092] Network 360 may be implemented as a single network or a combination of multiple networks. For example, in various embodiments, network 360 may include the Internet, one or more intranets, a telephone line network, a wireless network, and / or other suitable types of networks. Thus, network 360 may correspond to a small communications network, such as a private or local area network, or a larger network, such as a wide area network or the Internet, accessible by various components of system 300.
[0093] [Workflow example] 7 is an exemplary logic flow diagram illustrating a method for customized search based on the framework and architecture shown in FIGS. 1-6, according to some embodiments described herein. One or more of the processes of method 700 may be implemented, at least in part, in the form of executable code stored on a non-transitory, tangible, computer-readable medium that, when executed by one or more processors, may cause the one or more processors to perform one or more of the processes. In some embodiments, method 800 corresponds to the operation of code text generation module 230 (e.g., FIGS. 2-6), which performs a search based on user input and provides natural language responses to conversational search queries using NL models.
[0094] At step 702, natural language input is received at a server from a user interface on a user device. As shown in FIG. 4 , according to some embodiments, this input may include one or more of natural language user input 402, user context 404, or other context 406. In some embodiments, natural language user input 402 is a conversational or natural language query provided by a user. In some embodiments, the natural language input may be one or more of text input, audio input, image input, and video input.
[0095] In step 704, the server generates one or more search queries based on the natural language input received in step 702. In some embodiments, this step may be performed by a parser network. The generation of the search queries may be performed, for example, as a preprocessing stage by the NL preprocessing sub-module 231 (described above in FIG. 2) or may be performed by the search sub-module 232 (described above in FIG. 2) as part of a search operation.
[0096] In some embodiments, the server may further determine one or more potential search objects within the generated search query. This may be used to determine particular data sources that would be beneficial to search based on the input, the particular API accessed, and / or the like. In some embodiments in which the server is performing more than one search, this step may be used to determine that previous search results are relevant to the search object of the current input. In this way, resources may be saved by avoiding the need to perform multiple overlapping searches when the previous search results may be sufficient to provide a natural language output for the most recent input.
[0097] In step 706, the server retrieves one or more search results through a real-time search on one or more data source servers based on the one or more search queries. Further information on how searches are performed is provided in co-pending, commonly assigned U.S. Non-Provisional Application No. 17 / 981,102, which is incorporated by reference. In some embodiments, the search is performed based at least in part on the search objects identified in step 704.
[0098] In step 708, the server generates natural language output based at least in part on the one or more search results and includes a reference to at least one data source server. In some embodiments, the output is generated entirely by the server receiving and processing the input and performing the search. In other embodiments, the output is generated at least in part through use of an external NL network or LLM interfaced through the LLM interface submodule 234 (discussed in FIG. 2).
[0099] In this manner, the server may provide one or more natural language outputs that include an indication of web pages, PDFs, images, videos, and other sources used to generate the natural language output, as described above with reference to Figures 1-6 and further described below with reference to Figures 8-9.
[0100] 8 is an exemplary logic flow diagram illustrating a method for customized search based on the framework and architecture shown in FIGS. 1-6, according to some embodiments described herein. One or more of the processes of method 800 may be implemented, at least in part, in the form of executable code stored on a non-transitory, tangible, computer-readable medium that, when executed by one or more processors, may cause the one or more processors to perform one or more of the processes. In some embodiments, method 800 corresponds to the operation of code text generation module 230 (e.g., FIGS. 2-6), which performs searches based on user input and provides natural language responses to conversational search queries using NL models.
[0101] In step 802, an input query is received by the search server via a data interface. As shown in Figure 4, according to some embodiments, the input query may include one or more of natural language user input 402, user context 404, or other context 406. In some embodiments, the natural language user input 402 is a conversational or natural language query provided by a user.
[0102] The generative AI system may also take user context into account when generating conversational responses to user queries. In some embodiments, user context 404 may include any combination of user profile information (e.g., user ID, user gender, user age, user location, zip code, device information, mobile application usage information, etc.), user configuration preferences or dislikes of one or more data sources (e.g., as illustrated in co-pending and commonly owned U.S. Nonprovisional Application No. 17 / 981,102), and the user's past activity approving or disapproving search results from particular data sources. For example, if a user did not previously approve of a particular website, the generative AI system may prioritize other websites for gathering information and providing conversational responses. Conversely, if a user previously approved of a particular website, the generative AI system may instead aim to find results from the approved website before searching other web pages to provide a conversational response.
[0103] In addition to context related to the user, the generative AI system may further consider other context, such as information related to previous user interactions with the generative AI system. In some embodiments, other context 406 may include one or more of previous user queries, previous responses by the generative AI system, previous searches performed by the generative AI system, and any other contextual information related to current or previous interactions between the generative AI system and the user or between the generative AI system and the Internet.
[0104] In step 804, the search server may convert the input query into a modified query. In some embodiments, the search server may convert the input query using an NL model implemented on one or more hardware processors in the search server. In this manner, the search server may convert the input query into an input sequence of tokens. In some embodiments, other contexts, such as user context and user preferences, may also be incorporated into the modified query and corresponding token sequence.
[0105] In step 806, the search server determines one or more potential search objects in the revised query, which may be keywords or phrases in the revised query that have been identified as being of particular interest.
[0106] In step 808, the search server performs a search based on the identified potential search objects from the revised query. In some embodiments, when multiple potential search objects are identified in step 806, multiple searches may be performed, such that a search is performed for each potential search object. In other embodiments, a single search may be performed based on two or more potential search objects. Before performing the search, the search server may identify potential data sources relevant to the revised query. The search server may then send search inputs based on the potential search objects to the identified potential data sources and incorporate the results into a set of search results. The relevant data sources may be identified based on one or more tokens generated from user input, user context, and other contexts, may be based on previously identified relevant data sources, or may be provided directly by the user.
[0107] In step 810, the search results from the search performed in step 808 are incorporated into the NL model and processed. For example, the NL model may rank the search results, or the search results may have been ranked when received by the generative AI system.
[0108] In step 812, the NL model generates a response to the user query. The NL model utilizes the search results and corresponding information, the user query, the user context, and other context-incorporating tokens, and other relevant information to generate a text-based response that corresponds to the desired output.
[0109] In step 814, the search server inserts a citation into the response generated in step 812. The search server may insert one or more citations at the end of the sentence to indicate where the information contained in the sentence can be found in the search results, provide additional links to search results of interest, and otherwise provide context to the user.
[0110] In step 816, the search server transmits a response to the user query and the set of search results obtained for the user query to the user device. In some embodiments, the results may further incorporate content based on one or more search apps, one or more interactive graphic elements, or other data or applications related to the generated response.
[0111] In some embodiments, the search server determines in steps 806-808 that one or more applications are relevant to the potential search object in the user query. Thus, when retrieving the first set of search results in step 808, or after a response is generated by the NL model in step 812, the search server may identify APIs corresponding to one or more applications relevant to the potential search object identified in step 806. These applications may then be included as output with the response to the first query, such that information retrieved via the APIs is provided to the user relevant to the first query.
[0112] The steps detailed above for method 800 may be repeated an indefinite number of times as the user responds to the output provided by the search server, asks additional questions, or otherwise engages in a conversation. In this manner, the generative AI system may maintain a history of the conversation. This allows the generative AI system to maintain context as the user asks questions or interacts with the generative AI system, enabling responses more precisely tailored to what the user is looking for. Additionally, this may allow the generative AI system to reduce resource usage by allowing searches performed on previous user inputs to be reused when the current user input is determined to be sufficiently similar to, or otherwise able to rely on, the same set of search results already assembled.
[0113] For example, when a second (or third, or subsequent) user input is provided, the generative AI system converts the input query into a revised query, as detailed in step 804. Once the search server has determined one or more potential search objects in the revised query, as detailed in step 806, the search server determines whether any of the potential search objects in the user input are related to one or more potential search objects in the previous input.
[0114] If the search server determines that a potential search object is related to a potential search object from a previous input, the search server may use the search results related to the potential search object from the previous input rather than initiating a new search based on the search object.
[0115] However, if the search server determines that the potential search objects are not related to the potential search objects from the previous input, the search server can instead perform a new search based on the potential search objects of the current user input. This ensures that the search server can maintain responses that are relevant to the current user input if the user asks the search server an unrelated question or changes the context of the user input. Although the previous input may be utilized to determine additional relevant context for the current user input, performing a new search may still yield additional relevant information.
[0116] 9 is an exemplary logic flow diagram illustrating a method of a generative AI system based on the framework shown in FIGS. 1-6 , according to some embodiments described herein. One or more of the processes of method 900 may be implemented, at least in part, in the form of executable code stored on a non-transitory, tangible, computer-readable medium that, when executed by one or more processors, may cause the one or more processors to perform one or more of the processes. In some embodiments, method 900 corresponds to the operation of text generation module 230 (e.g., FIGS. 2-6 ) that performs a search based on user input and provides a tailored written response to the user query using the NL model.
[0117] In step 902, an input query and one or more user constraints are received by the search server via a data interface. As shown in FIG. 4 , according to some embodiments, the input query may include one or more of natural language user input 402, user context 404, or other context 406. In some embodiments, the user query 402 is a conversational or natural language query provided by a user. The user constraints may include one or more parameters such as a use case, a desired tone, and a target audience.
[0118] In step 904, the search server may convert the input query and user constraints into a modified query. In some embodiments, the search server may convert the input query and user constraints using an NL model implemented on one or more hardware processors of the search server. In this manner, the search server may convert the input query and user constraints into an input sequence of tokens. In some embodiments, other contexts, such as user context and user preferences, may also be incorporated into the modified query and corresponding sequence of tokens.
[0119] In step 906, the search server determines one or more potential search objects in the revised query, which may be keywords or phrases in the revised query that are identified as being of particular interest. The search server may also determine key user constraints to use in step 914 when modifying the output satisfies the provided user constraints.
[0120] In step 908, the search server performs a search based on the identified potential search objects from the revised query. In some embodiments, when multiple potential search objects are identified in step 906, multiple searches may be performed, such that a search is performed for each potential search object. In other embodiments, a single search may be performed based on two or more potential search objects. Before performing the search, the search server may identify potential data sources relevant to the revised query. In some embodiments, specific data sources may be identified due to the provided user constraints determined in step 906. The search server may then send search inputs based on the potential search objects to the identified potential data sources and incorporate the results into a set of search results. The relevant data sources may be identified based on one or more tokens generated from the user input, user context, and other contexts, may be based on previously identified relevant data sources, or may be provided directly by the user.
[0121] In step 910, the search results from the search performed in step 908 are incorporated into the NL model and processed. For example, the NL model may rank the search results, or the search results may have been ranked when received by the generative AI system.
[0122] In step 912, the NL model generates a response to the user query. The NL model utilizes the search results and corresponding information, the user query, the user context, and other context-incorporating tokens, and other relevant information to generate a text-based response that corresponds to the desired output. In step 914, the response is modified based on user constraints, such as to ensure that the response meets a desired tone or audience, although user constraints may also be taken into consideration when generating the response to the user query. For example, for a social media posting use case, the response may need to be shorter and more concise, while for an essay use case, it may be desirable to generate a longer response.
[0123] In step 914, the search server modifies the response based on user constraints, which may include ensuring that the language used in the response meets a desired tone or is at a desired reading level for the target audience.
[0124] In step 916, the search server transmits a response to the user query and the set of search results obtained for the user query to the user device. In some embodiments, the results may further incorporate content based on one or more search apps, one or more interactive graphic elements, or other data or applications related to the generated response.
[0125] 10A-10E provide exemplary UI diagrams illustrating an embodiment implementing a generative AI system incorporated into a search system. In an exemplary embodiment, the text generation tool may be implemented as a search assistance tool, such as an artificial intelligence (AI) chatbot, that converses with a user during a search and provides a summary of search results in response to search topics of interest to the user. For example, when a user performs a search by inputting search terms, a search engine may return a list of search results. The text generation tool may generate a summary in response to the list of search results and present the search summary to the user in the form of a conversational response. That is, the input user query and the generated search results may be input to the generator.
[0126] In one implementation, the data sources most relevant to the search query "best headphones" may be selected for the search, e.g., a shopping site such as Looria, an electronics rating site such as PCMag, etc. Search results within each data source may be presented to the user in respective horizontal panels. In one embodiment, each data source may perform its own independent search within its database in response to the search terms. For example, the data source "PCMag" may search for "best headphones" based on reviewing articles of different headphones. Similarly, the data source "Looria" may search for "best headphones" based on user ratings, etc.
[0127] In one embodiment, the search assistance tool can utilize the entered search query and returned search results to provide the user with conversational statements summarizing what the best headphones are. For example, as shown in FIG. 8A, the search tool can present a search summary in a conversational format that describes the "Sony WH-1000XM5" based on search results from multiple data sources related to the search term "best headphones."
[0128] In some embodiments, the search assistance tool may further provide citations in the search summary to specific results from the search, allowing the user quick access to view the results. For example, as shown in FIG. 10A, a user-clickable citation to "Sony WH-1000XM5" headphones may be provided in the summary, and the user may select to be redirected to specific search results for "Sony WH-1000XM5."
[0129] In this way, the search assistance tool improves the user search experience by providing a more understandable and readable search result output environment for the user and easy linking to related web pages based on the results.
[0130] In some embodiments, the search assistance tool may implement a pre-trained language model and utilize the pre-trained model to provide a conversational output to the user based on the input. However, the search assistance tool may also utilize search results as reference inputs, allowing the search assistance tool to be up-to-date and avoiding outdated and incorrect answers due to information updates after training is complete. For example, if a new set of headphones is released after training is complete, the search assistance tool may provide the user with output that takes the new headphones into account and is therefore relevant to the user despite changes in technology.
[0131] In further embodiments, the search assistance tool can progressively update the search and search summary as the user can provide additional conversational input, thereby enabling the search assistance tool to provide more refined conversational results. In response to the user's conversational input, the search assistance tool can further determine whether the search engine should refine previous search results, generate additional output based on previous search results, and / or conduct a new search.
[0132] For example, as shown in FIG. 10B , following the above example, the user may further input, "I only like over-ear headphones—what are the best ones?" Based on this input, the search assistance tool may decide to refine the previous search based on "over-ear headphones" and, accordingly, provide input that is more tailored to the user's input and preferences while maintaining a conversational element. In such instances, the search page may be slightly updated or remain unchanged while the conversation between the user and the AI chatbot is taking place, allowing the user to continue browsing the results while the conversation is ongoing.
[0133] In other embodiments, the search assistance tool can determine whether a refined search is necessary. For example, if a user's input begins asking about a particular brand of headphones, the initial search results may not have meaningful input about that brand of headphones. Thus, the AI chatbot can initiate a refined search directed toward that brand to provide further input information for consideration and output to the user along with corresponding citations. This allows the AI chatbot to provide contextually relevant, up-to-date information and output to the user without requiring the user to initiate an additional search.
[0134] In another embodiment, the search assistance tool can determine to start an entirely new search when the user input switches to a new topic. For example, if following a search query for "best headphones," the user enters another input, "revert git commit," as shown in FIG. 10C, the search assistance tool can determine that this relates to an entirely different topic and therefore requires a new search.
[0135] In this way, the generative AI system may not have to perform a new search every time an input is generated by a user, to conserve computational resources and avoid overloading the search system, but may instead determine the point at which the context of the conversation has either changed to require new search results or become sufficiently specific to benefit from additional results from a more detailed search.
[0136] In further embodiments, the search assistance tool can provide responses corresponding to the requested input. For example, as shown in Figures 10C-10E, a user can search for "revert git commit" for specific help in writing code, and the search assistance tool can generate output that includes instructions on how to write the code, including citations, as well as code blocks that implement example code to perform the task. The user may continue to ask follow-up questions (e.g., Figures 10D-10E), and the search assistance tool may refine its search results based on the user input and generate updated summaries.
[0137] In a further implementation, a user can ask for a picture of an animal, such as a cat, and the search assistance tool can generate the image directly in the output along with the text. Thus, the search assistance tool can utilize input information from both the user and the search results to provide tailored output that is relevant and timely for the user.
[0138] FIG. 11 provides an exemplary UI diagram illustrating an embodiment that implements a text generation tool based on set user parameters. In some exemplary embodiments, the text generation tool may be built on a search engine to automatically generate text based on user customizations. For example, the text generation tool may include a generative model provided via a web-based platform. The web-based platform interface allows a user to provide customized user input regarding the text to be generated, such as the type of text, intended audience, tone, content, and / or the like.
[0139] A user can interact with the web-based platform interface to provide a query in the form of a desired topic 1140 for which a response should be generated. The generative AI system can take as input user constraints such as use case 1110, tone 1120, and target audience 1130. This information is provided to the search server and processed there, as described in step 902 of Figure 9. Step 916 of Figure 9 can then generate a response 1150 that is provided to the user via the web-based platform.
[0140] Where applicable, various embodiments provided by the present disclosure may be implemented using hardware, software, or a combination of hardware and software. Also, where applicable, various hardware and / or software components described herein may be combined into composite components comprising software, hardware, and / or both without departing from the spirit of the present disclosure. Where applicable, various hardware and / or software components described herein may be separated into subcomponents comprising software, hardware, or both without departing from the scope of the present disclosure. Furthermore, where applicable, it is contemplated that software components may be implemented as hardware components, and vice versa.
[0141] Software according to the present disclosure, such as program code and / or data, may be stored on one or more computer-readable media. It is also contemplated that the software specified herein may be implemented using one or more networked and / or other general-purpose or special-purpose computers and / or computer systems. Where applicable, the order of the various steps described herein may be changed, combined into composite steps, and / or separated into substeps to provide the features described herein.
[0142] Various features and steps described herein may be implemented as a system comprising one or more memories storing various information described herein, and one or more processors coupled to the one or more memories and a network, the one or more processors operable to perform the steps described herein, as a non-transitory computer-readable medium comprising a plurality of computer-readable instructions adapted, when executed by the one or more processors, to cause the one or more processors to perform methods including the steps described herein, as well as methods executed by one or more devices, such as hardware processors, user devices, servers, and other devices described herein.
Claims
1. 1. A processor-implemented method for neural network based text generation using real-time search results, comprising: receiving, at a server, a natural language input from a user interface on a user device; generating, at the server, one or more search queries based on the natural language input; obtaining one or more search results through a real-time search in one or more data source servers based on the one or more search queries; generating, at the server implementing a generative artificial intelligence (AI) neural network, an output based at least in part on the one or more search results, wherein the output includes a reference to at least one data source server.
2. 10. The method of claim 1, wherein the natural language output is generated by the generative AI neural network of the server, the generative AI neural network may comprise a language model.
3. sending a text generation input including the one or more search results to an external server hosting a language model; obtaining the output from the external server; The method of claim 1 further comprising:
4. The output is generated further based on user configuration parameters, the user configuration parameters comprising: a type of the natural language output; and an intended audience for the natural language output; and tones of the natural language output; and a format for the natural language output; and and a length of the natural language output.
5. The method of claim 1 , wherein the obtaining one or more search results through real-time searching includes obtaining content from a web file via a link in the one or more search results.
6. 2. The method of claim 1, wherein the output includes a portion of text that is a summary of one or more search results, and the reference to at least one data source server indicates that the portion of text is associated with the at least one data source server.
7. The method of claim 1 , wherein the natural language input comprises one or more of a text input, an audio input, an image input, and a video input.
8. 1. A system for generating neural network based text using real-time search results, comprising: a communications interface configured to receive natural language input via a user interface implemented on a user device; a server implementing a generative artificial intelligence (AI) neural network and a plurality of processor-executable instructions; one or more processors that execute the instructions to perform operations; The operations include generating, at the server, one or more search queries based on the natural language input; obtaining one or more search results through real-time searching on one or more data source servers based on the one or more search queries; and generating, at the server implementing a generative artificial intelligence (AI) neural network, an output based at least in part on the one or more search results, wherein the output includes a reference to at least one data source server.
9. 9. The system of claim 8, wherein the natural language output is generated by the generative AI neural network of the server, the generative AI neural network may comprise a language model.
10. sending a text generation input including the one or more search results to an external server hosting a language model; obtaining the output from the external server; The system of claim 8 further comprising:
11. The output is generated further based on user configuration parameters, the user configuration parameters comprising: a type of the natural language output; and an intended audience for the natural language output; and tones of the natural language output; and a format for the natural language output; and and a length of the natural language output.
12. The system of claim 8 , wherein the obtaining one or more search results through real-time searching includes obtaining content from a web file via a link in the one or more search results.
13. 10. The system of claim 8, wherein the output includes a portion of text that is a summary of one or more search results, and the reference to at least one data source server indicates that the portion of text is associated with the at least one data source server.
14. The system of claim 8 , wherein the natural language input comprises one or more of a text input, an audio input, an image input, and a video input.
15. 1. A processor-readable non-transitory storage medium storing a plurality of processor-executable instructions for generating neural network based text using real-time search results, the instructions being executed by one or more processors to: receiving, at a server, a natural language input from a user interface on a user device; generating, at the server, one or more search queries based on the natural language input; obtaining one or more search results through a real-time search in one or more data source servers based on the one or more search queries; and generating, at the server implementing a generative artificial intelligence (AI) neural network, an output based at least in part on the one or more search results, wherein the output includes a reference to at least one data source server.
16. 16. The processor-readable non-transitory storage medium of claim 15, wherein the output is generated by the generative AI neural network of the server, the generative AI neural network may comprise a language model.
17. sending a text generation input including the one or more search results to an external server hosting a language model; obtaining the output from the external server; 16. The processor-readable non-transitory storage medium of claim 15, further comprising:
18. The output is generated further based on user configuration parameters, the user configuration parameters comprising: a type of the natural language output; and an intended audience for the natural language output; and tones of the natural language output; and a format for the natural language output; and and a length of the natural language output.
19. 16. The processor-readable non-transitory storage medium of claim 15, wherein the obtaining one or more search results through a real-time search includes obtaining content from a web file via a link in the one or more search results.
20. 16. The processor-readable non-transitory storage medium of claim 15, wherein the output includes a portion of text that is a summary of one or more search results, and the reference to at least one data source server indicates that the portion of text relates to the at least one data source server.
Citation Information
Patent Citations
Horizontal member coupled non-welded modular structure with slab floor
KR102307326B1
System and method for transferable natural language interface
US20220129450A1
Cited By
Information processing apparatus, information processing method, and information processing program
JP2026001571A
Business support systems, business support methods, and programs
JP2026136604A