System and method for prompt-based query generation for diverse retrieval
Through the prompt-based query generation method, a large language model and a dual encoder are combined to generate task-specific synthetic data, which solves the problem of poor performance of neural search systems in the case of scarcity of data, and achieves efficient training and performance improvement in diversified search tasks.
Patent Information
- Application Number
- CN202380068098.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-21
- Filing Date
- 2023-09-21
- Publication Date
- 2025-05-06
AI Technical Summary
Existing neural search systems perform poorly in diversified search tasks, especially in the case of scarcity of data, requiring a large number of annotated query-document pairs to train, resulting in high data demand.
Through a hint-based query generation method, task-specific synthetic data is generated using a combination of large language model (LLM) and standard dual encoders, and only a few examples (such as 2 to 8) are used to train the document retrieval model.
Significant improvements in training neural search systems on various tasks have been achieved, reducing the need for highly engineered search architectures, and improving search performance.
Smart Images

Figure CN119948474A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS / INCORPORATED BY REFERENCE
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 376,508, filed on September 21, 2022, which is hereby incorporated by reference in its entirety. Background Art
[0003] Natural language processing tasks such as question answering, claim verification, etc. involve retrieving knowledge from large document collections containing millions or billions of candidates. Training machine learning models to perform such tasks involves large datasets of annotated query-document pairs for each task. Summary of the invention
[0004] There are many diverse and unique retrieval tasks that target different scenarios. For example, different question answering (QA) tasks typically have different query distributions. In addition, for example, retrieval tasks are not limited to question answering tasks, although retrieval systems are generally trained on question answering data. For example, the task may involve retrieving entities mentioned in the query, or finding evidence to support or argue a given statement. Unlike other retrieval systems, neural retrieval systems are generally trained on question answering data. Due to data scarcity, neural retrieval systems may perform poorly on many tasks because a large number of annotated query-document pairs are required to train the neural retrieval system for each task. In addition, for example, dual encoder models generally rely on more data than their cross-attention counterparts, and typically require tens of thousands of labeled examples to make the intra-batch softmax loss function effective. Therefore, there are technical problems in training neural retrieval systems on various tasks using a small number of labeled examples.
[0005] As described herein, prompt-based query generation, filtering, and retriever training (also referred to as the PROMPTAGATOR system or PROMPTAGATOR) can leverage the power of large language models (LLMs) with the efficiency of standard dual encoders. PROMPTAGATOR can create task-specific dual encoders by generating large amounts of task-specific synthetic data. PROMPTAGATOR is able to achieve significant improvements in retrieval by using only two to eight examples for prompts (e.g., without a reranker). PROMPTAGATOR also significantly reduces the need for highly engineered retrieval architectures.
[0006] In one aspect, a computer-implemented method for prompt-based query generation is provided. The method includes receiving, by a computing device, at least two prompts associated with a retrieval task to be performed on a document corpus associated with the task. The method also includes applying a large language model based on the at least two prompts and the document corpus to generate a synthetic training dataset including a plurality of query-document pairs, wherein each query-document pair includes a synthetically generated query and a document from the document corpus. The method additionally includes training a document retrieval model on the plurality of query-document pairs from the synthetic training dataset to take an input query associated with the retrieval task and predict an output document retrieved from the document corpus. The method further includes providing, by the computing device, the trained document retrieval model.
[0007] In a second aspect, a computing device for prompt-based query generation is provided. The computing device includes one or more processors and a data storage device. The data storage device has computer executable instructions stored thereon, which when executed by the one or more processors cause the computing device to perform functions. The functions include: receiving, by the computing device, at least two prompts associated with a retrieval task to be performed on a document corpus associated with the task; applying a large language model based on the at least two prompts and the document corpus to generate a synthetic training data set including a plurality of query-document pairs, wherein each query-document pair includes a synthetically generated query and a document from the document corpus; training a document retrieval model on the plurality of query-document pairs from the synthetic training data set to take an input query associated with the retrieval task and predict an output document retrieved from the document corpus; and providing, by the computing device, the trained document retrieval model.
[0008] In a third aspect, a computer program for prompt-based query generation is provided. The computer program includes instructions that, when executed by a computer, cause the computer to perform functions. The functions include: receiving, by a computing device, at least two prompts associated with a retrieval task to be performed on a document corpus associated with the task; applying a large language model based on the at least two prompts and the document corpus to generate a synthetic training dataset comprising a plurality of query-document pairs, wherein each query-document pair comprises a synthetically generated query and a document from the document corpus; training a document retrieval model on the plurality of query-document pairs from the synthetic training dataset to take an input query associated with the retrieval task and predict an output document retrieved from the document corpus; and providing, by the computing device, the trained document retrieval model.
[0009] In a fourth aspect, an article of manufacture for prompt-based query generation is provided. The article of manufacture includes one or more computer-readable media having computer-readable instructions stored thereon, which, when executed by one or more processors of a computing device, cause the computing device to perform functions. The functions include: receiving, by the computing device, at least two prompts associated with a retrieval task to be performed on a document corpus associated with the task; applying a large language model based on the at least two prompts and the document corpus to generate a synthetic training dataset comprising a plurality of query-document pairs, wherein each query-document pair comprises a synthetically generated query and a document from the document corpus; training a document retrieval model on the plurality of query-document pairs from the synthetic training dataset to take an input query associated with the retrieval task and predict an output document retrieved from the document corpus; and providing, by the computing device, the trained document retrieval model.
[0010] In a fifth aspect, a system for prompt-based query generation is provided. The system includes: means for receiving, by a computing device, at least two prompts associated with a retrieval task to be performed on a document corpus associated with the task; means for applying a large language model based on the at least two prompts and the document corpus to generate a synthetic training dataset comprising a plurality of query-document pairs, wherein each query-document pair comprises a synthetically generated query and a document from the document corpus; means for training a document retrieval model on the plurality of query-document pairs from the synthetic training dataset to take an input query associated with the retrieval task and predict an output document retrieved from the document corpus; and means for providing, by the computing device, the trained document retrieval model.
[0011] In a sixth aspect, a computer-implemented method of applying a trained document retrieval model is provided. The method includes receiving, by a computing device, an input query associated with a retrieval task to be performed on a document corpus associated with the task. The method also includes predicting, by the trained document retrieval model, a document from the document corpus, wherein the document has a high relevance to the input query, the document retrieval model having been trained on a plurality of query-document pairs from a synthetic training data set that has been generated by a large language model based on at least two prompts associated with the retrieval task. The method further includes providing, by the computing device and in response to the input query, the predicted document.
[0012] In a seventh aspect, a computing device for applying a trained document retrieval model is provided. The computing device includes one or more processors and a data storage device. The data storage device stores computer executable instructions, which, when executed by one or more processors, cause the computing device to perform functions. The functions include: receiving an input query associated with a retrieval task by the computing device, the retrieval task to be performed on a document corpus associated with the task; predicting a document from the document corpus by a trained document retrieval model, wherein the document has a high relevance to the input query, the document retrieval model has been trained on multiple query-document pairs from a synthetic training data set, the synthetic training data set has been generated by a large language model based on at least two prompts associated with the retrieval task; and providing the predicted document by the computing device and in response to the input query.
[0013] In an eighth aspect, a computer program for applying a trained document retrieval model is provided. The computer program includes instructions that, when executed by a computer, cause the computer to perform functions. The functions include: receiving, by a computing device, an input query associated with a retrieval task to be performed on a document corpus associated with the task; predicting, by a trained document retrieval model, a document from the document corpus, wherein the document has a high relevance to the input query, the document retrieval model having been trained on a plurality of query-document pairs from a synthetic training data set that has been generated by a large language model based on at least two prompts associated with the retrieval task; and providing, by a computing device and in response to the input query, the predicted document.
[0014] In a ninth aspect, a product for applying a trained document retrieval model is provided. The product includes one or more computer-readable media having computer-readable instructions stored thereon, which, when executed by one or more processors of a computing device, cause the computing device to perform functions. The functions include: receiving, by a computing device, an input query associated with a retrieval task to be performed on a document corpus associated with the task; predicting, by a trained document retrieval model, a document from the document corpus, wherein the document has a high relevance to the input query, the document retrieval model having been trained on a plurality of query-document pairs from a synthetic training data set, the synthetic training data set having been generated by a large language model based on at least two prompts associated with the retrieval task; and providing, by a computing device and in response to the input query, the predicted document.
[0015] In a tenth aspect, a system for applying a trained document retrieval model is provided. The system includes: means for receiving, by a computing device, an input query associated with a retrieval task to be performed on a document corpus associated with the task; means for predicting, by a trained document retrieval model, a document from the document corpus, wherein the document has a high relevance to the input query, the document retrieval model having been trained on a plurality of query-document pairs from a synthetic training dataset that has been generated by a large language model based on at least two prompts associated with the retrieval task; and means for providing, by the computing device and in response to the input query, the predicted document.
[0016] The foregoing summary is illustrative only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a diagram showing an overview of the training aspects of a document retrieval model according to an example embodiment.
[0018] Figure 2 is a diagram illustrating an example query generation model and an example document retrieval model according to an example embodiment.
[0019] Figure 3 Example prompt templates for various data sets are shown according to example embodiments.
[0020] Figure 4 An example comparison of different retriever frameworks according to example embodiments is shown.
[0021] Figure 5 An example comparison of different retriever frameworks and retriever and re-ranker frameworks according to example embodiments is shown.
[0022] Fig. 6A is a bar graph illustrating the impact of round-trip filtering according to an example embodiment.
[0023] Figure 6B is a diagram illustrating a human-annotated example according to an example embodiment.
[0024] Figure 6C Results of example ablations on a query generation model according to example embodiments are shown.
[0025] Figure 7 The impact of different Fine-tuned Language Network (FLAN) versions according to an example embodiment is shown.
[0026] Figure 8The impact of selected examples on model output according to an example embodiment is shown.
[0027] Fig. 9A Shown are example first-word distributions over queries generated from different models in the Argumentation Analysis (ArguAna) dataset, according to an example embodiment.
[0028] Fig. 9B Shown are example first word distributions over queries generated from different models in the ArguAna dataset, according to an example embodiment.
[0029] Fig. 9C Shown are example first word distributions over queries generated from different models in the ArguAna dataset, according to an example embodiment.
[0030] Fig.9D Shown are example first word distributions over queries generated from different models in the ArguAna dataset, according to an example embodiment.
[0031] Fig. 10A Example few-shot and zero-shot generated queries randomly sampled from various datasets are shown according to example embodiments.
[0032] Fig. 10B Example few-shot and zero-shot generated queries randomly sampled from various data sets are shown in accordance with example embodiments.
[0033] Fig. 10C Example few-shot and zero-shot generated queries randomly sampled from various data sets are shown in accordance with example embodiments.
[0034] Fig.11 A table with average query lengths for various data sets is shown in accordance with an example embodiment.
[0035] Fig.12 is a diagram illustrating the training phase and the inference phase of a machine learning model according to an example embodiment.
[0036] Fig.13 A distributed computing architecture according to an example embodiment is depicted.
[0037] Fig.14 is a block diagram of a computing device according to an example embodiment.
[0038] Fig.15 Depicted is a network of computing clusters arranged as a cloud-based server system according to an example embodiment.
[0039] Fig.16 is a flowchart of a method according to an example embodiment.
[0040] Fig.17 is a flow chart of another method according to an example embodiment. DETAILED DESCRIPTION
[0041] On the one hand, the present application relates to solving the data scarcity problem by utilizing the function of large language model (LLM) while maintaining the efficiency of small dual encoder. In order to solve the data scarcity problem, PROMPTAGATOR combines prompts with large language models as a query generator without fine-tuning. Due to the use of large language models, PROMPTAGATOR can generate good queries without any training data from natural problems. PROMPTAGATOR can use a small amount of prompts that realize few-sample retrieval, by inputting only two to eight examples for each task in the prompts, resulting in significant improvement under these new examples. PROMPTAGATOR can amplify the function of this small amount of examples by creating task-specific prompts, rather than using them to directly train dual encoders. A large amount of training data including query and document pairs can be generated. After query generation, task-specific noisy paired data can be used to train task-specific dual encoders. The second iteration for training can be performed by filtering noisy paired data to generate a clean data set and using a clean data set to train task-specific dual encoders.
[0042] Neural retrieval models such as dual encoders can search large document collections containing millions to billions of passages. However, the Benchmark Information Retrieval (BEIR) shows that it can be challenging for neural retrievers to perform well on a variety of retrieval tasks that lack specialized training data. To address this issue, some efforts focus on transferring knowledge from high-resource question answering (QA) datasets and propose architectures with good inductive biases, such as models that allow fine-grained token-level interactions (e.g., Contextualized Late Interaction Based on BERT (ColBERT) and Sparse Vocabulary and Extensions (SPLADE)), which are associated with higher inference costs.
[0043] Data augmentation via synthetic query generation generally involves question generators learned from high-resource QA datasets and usually does not generalize well to new retrieval tasks. Some approaches utilize synthetic question generation on a target corpus to alleviate the data scarcity problem. While effective in improving the performance of QA retrieval tasks, one of the limitations of such approaches is that the question generation model is trained on a question answering dataset, so the distribution of generated questions may not match the true query distribution in the target task.
[0044] The few-shot LLM query generator described herein can generate good queries without any fine-tuning of the model. In fact, synthetically generated data can be powerful enough to reduce and / or eliminate the use of annotated query-document pairs from traditional high-resource datasets such as natural questions.
[0045] To ensure the quality of the generated data, only the generated data is used to describe the filtering strategy. This filtering strategy can remove ambiguity, generality, and low-quality issues and can significantly improve retrieval performance. It should be noted that although LLM is used to generate training data, the retrieval model at inference time can be any standard size retriever.
[0046] Existing approaches using LLMs are not suitable for task-specific few-shot adaptation and often come with high inference costs. For example, some approaches involve using GPT-3 in a dual encoder; however, the embedding dimension of the dual encoder is 12,000, which can make the search index footprint and inference cost too high for many applications. Another existing approach involves hinting the LLM for question generation, but this approach does not use task-specific few-shot hints for fast task adaptation. This approach also focuses primarily on models that rerank top queries from existing retrievers, rather than directly adapting the bottom-level retriever, which can efficiently search millions or billions of documents.
[0047] The approach described herein involves accounting for differences across retrieval tasks (e.g., search intent and query distribution) and presents a few-shot retrieval evaluation for a dataset (e.g., the BEIR dataset). As described, PROMPTAGATOR is a simplified model for few-shot retrieval that prompts LLM to generate synthetic task-specific training data. As a result, neural retrievers and rerankers can be trained based on only a small number of supervised examples. As shown, PROMPTAGATOR with two to eight examples can produce significantly better retrievers than existing models trained on MS MARCO or NQ, which have over 500,000 human-annotated examples and use more expensive architectures. For example, PROMPTAGATOR can outperform ColBERT v2 and SPLADE v2 on several retrieval tasks that can be tested, while reranking can improve the results by another five (5) points on standard retrieval evaluation metrics.
[0048] Overview System
[0049] Figure 11 is a diagram illustrating an overview 100 of aspects of training a document retrieval model according to an example embodiment. Some embodiments involve receiving at least two prompts associated with a retrieval task, (e.g., a first prompt 110 and a second prompt 115), which is to be performed on a document corpus 105 associated with the task. Some embodiments involve applying an LLM 120 to generate (e.g., generate 125) a synthetic training dataset 130 including a plurality of query-document pairs 135 based on the at least two prompts (e.g., the first prompt 110 and the second prompt 115) and the document corpus 105. Each of the plurality of query-document pairs 135 may include a synthetically generated query and a document from the document corpus 105. Some embodiments involve training (e.g., train 140) a document retrieval model 145 on the plurality of query-document pairs 135 from the synthetic training dataset 130 to take an input query associated with the retrieval task and predict an output document retrieved from the document corpus 105. Additional details of each of these steps are provided below.
[0050] Task
[0051] This article describes a few-shot retrieval approach for diverse retrieval tasks, where each task is associated with a short description and a small number of annotated examples to clearly illustrate the search intent. As described herein, for each new retrieval task, the few-shot examples can be input to a large language model (e.g., LLM 120), such as a fine-tuned language network (FLAN). FLAN is an LLM that has not been trained on any document retrieval or document-to-query generation task. FLAN can be prompted to perform document-to-query generation (e.g., generation 125). In general, the few-shot examples ensure that the specific search intent of a particular task is captured. Using the query generator large language model (LLM) 120, a large number of relevant queries can be synthetically generated for any document, thereby generating rich data (e.g., multiple query-document pairs 135 from a synthetic training dataset 130) for training a retriever (e.g., a document retrieval model 145), including an efficient dual encoder model.
[0052] In general, different retrieval tasks may have different search intents, in other words, different definitions of "relevance". For example, consider the task of retrieving documents from Wikipedia. The task may be a first task of retrieving entities mentioned in the query, or a second task of finding evidence to support or refute a given statement. Even if the tasks share the same domain, determining which documents are relevant to the query may be very different for each task. Furthermore, even when the search intents of different tasks may be similar, different tasks may have different query distributions. For example, some queries may be long, compound questions, while other queries may be short financial questions. The retrieval domain may vary from general web pages to articles from Wikipedia, questions from Quora, or papers on specific diseases.
[0053] For illustration purposes, a few-shot retrieval setting is described for the BEIR benchmark. Given a large corpus (e.g., document corpus 105), a retrieval model (e.g., document retrieval model 145) can be configured to find documents that are highly relevant to a provided query q according to a predefined notion of relevance. Formally, the retrieval task can be formulated as:
[0054]
[0055] in is a large document corpus for retrieval (e.g., document corpus 105), Q is the query distribution, and I is the underlying search intent for the task. Depending on the task, D can be any collection of documents, such as the web or Wikipedia. Retrieval of Q also varies across tasks, such as short keyword search queries, questions, arguments, etc. If I(q, d) = 1, this indicates that the search intent of q has been satisfied by document d. For example, in a question answering task, I QA (q, d) = 1 indicates that document d answers query q. For the same (q, d) pair, the relevance can be 1 or 0, depending on the search intent. For example, some argument retrieval tasks may try to retrieve supporting arguments, while other argument retrieval tasks may try to retrieve opposing arguments.
[0056] Generally, the target retrieval corpus can be given (e.g., a document corpus 105 for a specific task), but the amount of annotated query-document pairs for a new task may be limited. Existing approaches focus on adapting retrievers to new corpora. But it does not take into account the query or intention I T As described in this article, search intent can be expressed with short descriptions and few examples.
[0057] Few-shot setting
[0058] Humans are generally able to understand retrieval tasks by reading short prompts and viewing a small number of examples. Therefore, it is natural to determine whether a small number (e.g., 8 or less) of examples may be sufficient for a machine learning model to learn a task-specific retriever. In this regard, a few-shot retrieval evaluation can be constructed on the BEIR heterogeneous retrieval benchmark.
[0059] BEIR includes 18 information retrieval datasets across 9 domains, including biomedicine, finance, news, Twitter, Wikipedia, StackExchange, Quora, science, and others. These datasets cover a diverse range of search intents: QA retrieval (question to document), duplicate question discovery (question to question), fact checking (claim to document), etc. The original BEIR evaluation used a zero-shot setting, where any query or related query-document pairs from the evaluation dataset could not be used for training.
[0060] In some aspects, as described herein, BEIR can be modified to a few-shot setting by randomly selecting a small number (e.g., 2 to 8) of in-domain relevant query-document examples as task-specific supervision. In general, this number of examples is available. When a development set is available, examples can be sampled from the development set. For BEIR tasks with only a test set, samples from the test data can be used.
[0061] PROMPTAGATOR MODEL
[0062] The PROMPTAGATOR can be configured to transform a small number of examples into many more examples by prompting an LLM (e.g., LLM 120) to generate more data (e.g., synthetic training dataset 130), rather than directly using these examples to train a retriever (e.g., document retrieval model 145). In some embodiments, the PROMPTAGATOR can include three components: prompt-based query generation, consistency filtering, and retriever training. During prompt-based query generation, task-specific prompts can be combined with a large language model to target The relevant queries are generated for all documents in the dataset. The filtering step can then clean the generated data based on round-trip consistency. In some embodiments, a retriever trained only on synthetic data (e.g., multiple query-document pairs 135 from the synthetic training dataset 130) can be used to filter the synthetic data. Finally, a retriever (e.g., a dual encoder) and a cross-attention re-ranker can be trained based on the filtered data.
[0063] Figure 2200 is a diagram illustrating an example query generation model and an example document retrieval model according to an example embodiment. As described herein, PROMPTAGATOR combines prompts with a large language model (e.g., LLM 205) tuned by instructions to form a query generation with much fewer restrictions than other models. It can rely on very few (e.g., 8 or less) examples to train the retriever (e.g., dual encoder 210) and can outperform custom designed models. In some embodiments, the cross-attention re-ranker 215 can be a list-by-list model based on T5. For example, it can be based on query-document pairing, receiving a list of documents from synthetic data 220 given a query as input. In addition, each query-document pair can be represented as "Query:{q}Document:{d}", and these pairs can be input into the encoder of the cross-attention re-ranker 215 (e.g., T5 model). In some embodiments, a projection layer can be applied on the output encoding of the first word unit, and the output can be used as a ranking score. In addition, for example, a softmax cross entropy loss can be used on a network consisting of positive (q, d + ) and the sampled negative (q, d - ) pairs (e.g., 31 sampled negative pairs). Unlike monoT5, which is a point-by-point reranker that uses an encoder-decoder model and is trained to generate relevance tags, the criss-cross attention reranker 215 can be a list-by-list reranker and can be directly optimized for ranking performance.
[0064] Prompt-based query generation
[0065] In some embodiments, task-specific few-shot examples 225 may be input into a large language model (e.g., LLM 205). LLM 205 may be prompted to perform document-to-query generation 230. More specifically, let {(q i , d i )} k represents k few-shot examples 225, where according to the target task T(I T (q i , d i )=1), each example is a query and documents related to the query
[0066] Following FLAN, instruction hints can be applied to LLM 205 with the string prefix as follows:
[0067] e doc (d i )|e 查询 (q1)|...|e doc (dk )|e 查询 (q k )|e doc (d), (Equation 2)
[0068] Where "|" indicates a delimiter word, e doc (d) and e 查询 (q) are task-specific documents and query descriptions, respectively, and d represents new documents presented at reasoning time. In some embodiments, LLM 205 can be prompted to doc (d)="Argument: {d}" and e 查询 ="Counter Argument:{q}" generates a counter argument.
[0069] Figure 3 Example prompt templates 300 for various data sets are shown according to an example embodiment. Figure 3 The full set of example descriptions used in the prompts is shown. The datasets include Argument Analysis (ArguAna), Financial Question Answering (FiQA), Natural Multi-Hop Question Answering (HotpotQA), Entity Search on the DBpedia Knowledge Base (DBpedia-Entity), Nutrition Information Corpus (NFCorpus), Touché 2020, Text Retrieval for Covid Conference Data (TREC-Covid), Scientific Claims (SciFact), Scientific Literature (SciDocs), and Fact Extraction and Verification (FEVER).
[0070] Reference again Figure 2 , although the LLM 205 can be configured to generate However, in some embodiments, if the query description does not precede the actual query, the generation failure may be recorded; otherwise, the query description may be accepted. And can generate synthetic related examples
[0071] In some embodiments, the 235 documents to generate the synthesis Examples 220 are large sets, thereby amplifying information from a few examples (e.g., task-specific few-shot examples 225) into a large synthetic dataset (e.g., synthetic data 220). In general, the query distribution of the synthetic dataset (e.g., synthetic data 220) can be designed to be similar to the true task distribution And the query-document pairs are designed to convey the true search intent. T .
[0072] In some embodiments, LLM 205 may be a FLAN model, and the query generator may be represented as p FLAN (q|d). FLAN can be trained on a set of tasks described via instructions, and generally achieves high zero-shot / few-shot performance on unseen tasks. Some embodiments involve using a checkpoint of 137 billion parameters. During prompt engineering, up to 8 examples can be used, and the number can be further reduced if the examples exceed FLAN's input length limit. In addition, for example, if individual queries and documents in the examples are longer than a threshold length, the individual queries and documents can be manually truncated. As another example, up to 1 million documents can be randomly sampled from each corpus, and sampling-based decoding can be used to generate 8 questions per document, and the temperature parameter is set to 0.7.
[0073] Round-trip filtering of generated data
[0074] In some embodiments, round-trip filtering 240 may be applied to improve the quality of synthetic data set 220. In general, for synthetic queries generated from paragraph d Desired synthetic query Its original paragraph d should also be retrieved. In other words, the original d should have a high probability under a certain retriever. (The reverse of query generation). When such conditions are not met, the generated query can be filtered out.
[0075] Round trip filtering 240 can be effective for synthetic question generation on QA tasks. However, the prior art generally relies on question answering models for reverse filtering. Since not all retrieval tasks are similar to question answering, such prior art approaches may not be sufficient. Instead, at step training 250, training can be applied to an initial retriever (e.g., a small dual encoder) from unfiltered synthetic data (e.g., synthetic data 220). The trained initial retriever (e.g., the trained small dual encoder) can then be used to filter the synthetic data. In general, this approach works very well for the different search intents observed in BEIR. More precisely, given a synthetic query-document pair The initial retriever (e.g., the trained small dual encoder 210) can be used to predict When d appears in the first K paragraphs returned by the retriever, the generated query is selected Such filtering can greatly reduce the number of synthetic queries and can significantly improve retrieval performance. (described later) FIG. 6A to FIG. 6C An example of queries from the few-shot PROMPTAGATOR that are removed by round-trip filtering is shown.
[0076] Modeling Round Trip Filtering as a Latent Variable
[0077] In some embodiments, round-trip filtering 240 may involve treating queries as latent variables to be estimated. Consider a hypothetical graphical model where each query q is a latent variable and the documents retrieved for that query follow some distribution p(d|q, θ * ) of the observed variables, where θ * represents the parameters of a hypothetical “optimal” retriever that selects the “best” document for a query. For synthetic data generation, we can get the parameters from the posterior p(q|d,θ * ) is used to sample queries, and the posterior can be given as follows according to the Bayesian rule:
[0078]
[0079] where p(q|θ * ) is a prior on the query. In some aspects, it can be assumed that p(q|θ* ) is a “non-informative prior” that is uniform for all q. In terms of representation, p(q|θ * ) can be expressed as p(q). When the parameters θ of the best retriever are known * When θ * The variables θ are generally unknown, and therefore, Expectation Maximization (EM) can be used to estimate the variables θ. In general, the round-trip filtering algorithm described herein is a version of EM. EM and other latent variable learning methods can be used to infer missing data.
[0080] In EM, a variable θ can be estimated by approximately maximizing the marginal likelihood of the observed variable, as described by the following relationship:
[0081]
[0082] For example, in EM, an initial estimate of p(q|d) can be made for each document d. In some embodiments, the FLAN query generator can be used as the initial estimate: p FLAN (q|d). The M-step of EM may involve computations described by the following relation:
[0083]
[0084] This can be seen as equivalent to Train an initial retriever on documents that have been paired with queries generated by an LLM (e.g., FLAN) As described herein. In some embodiments, the E-step of EM may involve estimating for each document d As described by the following relationship:
[0085]
[0086] As indicated by Equation 6, The highest probability query under d may be the query with the highest probability of retrieving d using the initial retriever, since p(q) is an uninformative prior that is uniform for all q.
[0087] In some embodiments, importance sampling may be used to In importance sampling, we can select a proposal distribution p from the selected prop (q|d) to sample the query q. Then we can sample the query q by the ratio In some embodiments, the proposed distribution can be selected as p FLAN (q|d). Therefore, instead of applying importance weighting to each sample, the filtering process described in this paper can be configured to discard samples with The sample has a low value of , which is proportional to the numerator in the above importance weight.
[0088] Although this EM-like process can be repeated until convergence, a single round of filtering generally produces sufficient quality gains. In some respects, the training of the final retriever model can be viewed as another M-step.
[0089] Few-shot PROMPTAGATOR retriever
[0090] As described herein, synthetically generated data (e.g., synthetic data 220) is configured to allow training of task-specific neutral retrievers for tasks where in-domain fine-tuning may be challenging due to data scarcity. In some embodiments, a dual encoder 210 retrieval architecture and a pre-training-fine-tuning scheme may be used.
[0091] For example, pre-training 245 can involve pre-training a retriever (e.g., dual encoder 210). In some embodiments, pre-training can be performed on C4 using an independent cropping task from Contriever, where two random crops from the same document can be considered as artificial positive (query, document) pairs, and training can involve a cross-entropy loss on random negative samples from the batch at step training 250. Subsequently, the random negative samples from the batch can be used to generate the query from the prompt. The dual encoder 210 is fine-tuned. After training for a set number of epochs (e.g., at step training 250), the initial dual encoder can be used to apply round-trip filtering 240 to synthetic data (e.g., in synthetic dataset 220), and the dual encoder 210 can continue to be fine-tuned on the filtered data.
[0092] Some embodiments may involve a modified PROMPTAGATOR, referred to herein as PROMPTAGATOR++. Generally, this can be a reranker (e.g., cross-attention reranker 215) trained on the same synthetic data generated from prompt-based QGen (e.g., in synthetic dataset 220). QGen can be configured to use the slower but more accurate cross-attention reranker 215 to refine the retrieved candidates. The cross-attention reranker 215 can be trained using a cross-entropy loss (e.g., using 31 sampled negative samples 255 from the top 200 segments retrieved by the PROMPTAGATOR retriever), which approximates the inference time distribution (e.g., reranking the top 200 segments from the retriever).
[0093] Zero-Shot PROMPTAGATOR
[0094] In some embodiments, prompt-based query generation can be configured to run in a zero-shot manner, where a zero-shot prompt 260 can be applied regardless of the target task. For example, a zero-shot prompt 260 can be "{d} Read the passage and generate a query." Here, {d} represents the document text. Zero-shot PROMPTAGATOR and zero-shot PROMPTAGATOR++ can be configured by training a retriever (e.g., document retrieval model 145, dual encoder 210) and a reranker (e.g., cross-attention reranker 215) on such data. These models can serve as a baseline to illustrate the benefits of adapting few-shot prompts to target tasks.
[0095] Comparison with existing methods
[0096] Figure 4An example comparison 400 of different retriever frameworks according to an example embodiment is shown. For example, the setting of PROMPTAGATOR can be compared with some other means. Several dimensions of the PROMPTAGATOR means are simpler than the comparison means. For example, the dual encoder 210 may not rely on difficult negative mining, distillation from cross-attention teachers, and / or word-level retrieval. In addition, a reranker, such as (e.g., a cross-attention reranker 215 with 125 million parameters), may be configured to be smaller than the reranker in other means. In some embodiments, a simpler and smaller architecture can be designed to achieve high performance results when training with synthetic data that has been adapted to a few samples, as described herein. Compared with the curious parrot (InPars) and unsupervised paragraph reranking (UPR) means for search, the PROMPTAGATOR means adopts task-specific few-sample adaptation. In contrast, the prompts in InPars and UPR are independent of the task and therefore suffer the same limitations as the previous query generation means. Another difference between these means is that when the existing means focus on reranking, the PROMPTAGATOR means implements few-sample learning for both reranking and retrieval.
[0097] Evaluate
[0098] As described herein, the PROMPTAGATOR approach can be evaluated against the BEIR benchmark. Additionally, for example, ablation studies and qualitative analysis can be performed. In some embodiments, the FLAN training set can overlap with two datasets in BEIR, NQ and Quora. Additionally, for example, some existing approaches do not report few-shot or zero-shot results on MS MARCO because they use that dataset for fully supervised learning. Therefore, the evaluation may not be based on the MS MARCO, NQ, and Quora datasets. Instead, the normalized discounted cumulative gain (nDCG@10) for 10 items, a retrieval evaluation metric on BEIR, can be evaluated.
[0099] For query generation in PROMPTAGATOR, questions can be sampled from FLAN at a given temperature (e.g., 0.7). In some embodiments, setting the filter threshold K for round-trip filtering to 1 can provide better results on MS MARCO. Therefore, for all BEIR datasets, the filter threshold K can be set to 1. In some embodiments, the dual encoder can be implemented based on the generalizable T5 retriever (GTR). To ensure efficiency, a T5-base encoder architecture consisting of 110 million parameters can be used. For the PROMPTAGATOR++ reranking model, a T5-base version 1.1 encoder checkpoint with 125 million parameters can be initialized. In some embodiments, the top 200 candidates retrieved from the PROMPTAGATOR dual encoder can be reranked at reasoning time.
[0100] In some embodiments, the hyperparameter may involve a batch size of 6k; however, some of the corpora in BEIR contain only a few thousand documents, so that multiple related documents appear in the same batch. This may sometimes interact adversely with the intra-batch softmax loss. Therefore, the dataset can be divided into groups based on the corpus size, such as, for example, a small group (corpus size <50k), a medium group (corpus size 50k-500k), and a large group (corpus size>500k). For dual encoder training, a batch size of 128 can be used for small datasets, and a batch size of 6k can be used for other datasets. In addition, for example, 5k steps can be fine-tuned for large datasets, and 1k steps can be fine-tuned for other datasets. For the ranking model, a batch size of 64 can be used for all datasets, and a large dataset can be fine-tuned up to 20k steps, and other datasets can be fine-tuned up to 5k steps.
[0101] result
[0102] Figure 5 An example comparison 500 of different retriever frameworks and retriever and reranker frameworks according to example embodiments is shown. For example, experimental results for such models are shown. As shown, the zero-shot PROMPTAGATOR can serve as a strong baseline, with a very good performance compared to the baseline from MS MARCO. In addition, for example, Few-Shot PROMPTAGATOR can significantly improve on Zero-Shot PROMPTAGATOR, improving nDCG@10 by more than 2 points on average. This evaluation highlights the impact of few-shot learning. For example, despite its simple training procedure and model architecture, Few-Shot PROMPTAGATOR can outperform strong baselines such as GenQ and GPL that also use query generation to augment training data, as well as ColBERT v2 and SPLADE v2 that rely on word-level interaction architectures and distillation schemes.
[0103] As shown, the resequencer PROMPTAGATOR++ can improve performance on nDCG@10 by another 5 points. It can produce significant improvements compared to UPR, whose resequencer uses T0 instructions, an LLM with similar instruction tuning as FLAN. It can also outperform monoT5-3B, which has been previously shown to achieve the previous state-of-the-art resequencer performance on BEIR. Existing resequencer approaches generally use large models with 3 billion parameters to achieve better generalization, while PROMPTAGATOR++ uses a resequencer with smaller million (e.g. 125 million) parameters.
[0104] By comparing the few-shot PROMPTAGATOR to the baseline, significant improvements can be observed on Webis-Touché 2020 (Touché), followed by ArguAna (arg). The goal of Webis-Touché 2020 is to retrieve documents on controversial topics such as, for example, "should felons who have completed their sentence be allowed to vote?" The goal of ArguAna is to find counter-arguments against the input argument, and the input argument may be several sentences long. These two tasks are generally different from the traditional QA retrieval data used by other models, which is dominated by factual questions. However, the few-shot PROMPTAGATOR can be successfully adapted to this task using a small number of examples.
[0105] The impact of round-trip filtering.
[0106] Fig. 6A 6 is a bar graph 600A illustrating the impact of round-trip filtering according to an example embodiment. Fig. 6AIn Figure 2, the quality difference between the Few-shot PROMPTAGATOR with filtering and the Few-shot PROMPTAGATOR without filtering is shown. In general, filtering improves the performance on 8 out of 11 datasets and leads to a 2.5 point improvement on average, demonstrating the effectiveness of the filtering strategy described in this paper. While filtering appears to have a negative impact on model quality on NF Corpus and SciFact, this may be due to the relatively small size of these datasets in terms of queries generated and may indicate overfitting of the retriever.
[0107] Generated queries vs human-annotated queries
[0108] Figure 6B 600B is a diagram showing a human annotated example according to an example embodiment. Figure 6B In
[15] , an 8-sample PROMPTAGATOR is evaluated on MSMARCO and compared to a dual encoder trained on supervised data from MS MARCO. For evaluation purposes, no additional components were added to maintain a simpler comparison. MS MARCO was chosen because there is enough labeled data for this task and neither FLAN nor the models described in this paper were trained on MS MARCO examples. The results show that eight examples plus LLM can replace a large portion of the supervised examples.
[0109] Comparison with other query generation methods
[0110] Figure 6C Results 600C of an example ablation on a query generation model according to an example embodiment are shown. Figure 6C The zero-shot PROMPTAGATOR is compared to two other query generation approaches: GenQ (a prior system that uses a T5 query generation model trained on MS MARCO) and Natural Question-QGen (NQ-QGen), a T5QGen model fine-tuned on NQ. Some advantages of the zero-shot PROMPTAGATOR are shown, such as, for example, outperforming both baseline models by a significant margin. In general, NQ-QGen uses the same filtering, dual encoder training, batch size, and training steps as PROMPTAGATOR, providing a fair comparison of the query generators. The results 600C indicate that the contributing factors to PROMPTAGATOR may be better queries from the hinted LLM, and may not be from a specific training scheme or not specific hyperparameters.
[0111] Few-shot vs. zero-shot
[0112] Reference again Figure 5In the example comparison of different retriever frameworks and retriever and reranker frameworks in 500, the few-shot PROMPTAGATOR generally seems to outperform the zero-shot PROMPTAGATOR, except for Climate-FEVER. The Climate-FEVER dataset uses one of three labels to annotate query-document pairs, namely "support", "refute", or "insufficient information". However, BEIR generally considers these three annotations to be correlated, which may lead to undesirable results. Using query-document pairs annotated "insufficient information" in prompts may have a negative impact on the generation quality. Using such pairs in few-shot prompts may have a negative impact on query generation. In some evaluations, few-shot prompts from FEVER can be used because the two datasets share the same corpus and similar search intent. Based on such modified annotated examples, the few-shot PROMPTAGATOR seems to outperform the zero-shot PROMPTAGATOR.
[0113] Impact of FLAN Version
[0114] FLAN was trained on a set of datasets with some overlap with BEIR; specifically, FLAN includes Natural Questions (NQ) and Quora. However, FLAN was not trained on query-document pairs from NQ or Quora. To determine whether the inclusion of this data can bias the results on the final retrieval evaluation, additional ablation experiments can be designed. For example, based on the FLAN scheme, additional LLMs can be trained by excluding both the NQ and Quora datasets.
[0115] Figure 7 The impact of different FLAN versions according to an example embodiment is shown. The results are divided into two tables 700a and 700b. The last column 705 in table 700b shows the average values from the corresponding rows of tables 700a and 700b. For example, the average value 46.5 in the last column 705 is the average value of the values in the first row of tables 700a and 700b. Similarly, the average value 47.0 in the last column 705 is the average value of the values in the second row of tables 700a and 700b.
[0116] Figure 8 The effects of selected examples on the model output according to an example embodiment are shown. For example, Figure 8The impact and / or improvement of fine-tuning an existing retrieval model (e.g., GTR) using eight (8) examples is shown. Although the accuracy may drop slightly, the overall performance is better than the existing retriever. The results are divided into two tables 800A and 800B. The last column 805 in Table 800B shows the averages of the corresponding rows from Tables 800A and 800B. For example, the average value 40.4 in the last column 805 is the average of the values in the first row of Tables 800A and 800B. Similarly, the average value 38.7 in the last column 805 is the average of the values in the second row of Tables 800A and 800B.
[0117] Qualitative analysis
[0118] Some advantages of the few-shot PROMPTAGATOR can be illustrated based on the query distributions generated by different query generation methods for ArguAna. For example, for each distribution, a histogram of the first word of each query can be shown.
[0119] 9A to 9D Shown are example first word distributions over queries generated from different models in the ArguAna dataset, according to an example embodiment. Fig. 9A The first word distribution for gold queries is shown. Fig. 9B shows the distribution of the first words of a few-shot PROMPTAGATOR generated query. As indicated, Fig. 9B The distribution shown is closer to Fig. 9A The true distribution shown. Fig. 9C The first word distribution of NQ-QGen generated queries is shown. Even though queries in this task are generally arguments rather than questions, NQ-QGen mainly generates questions. Fig.9D It shows that few-shot FLAN can generate diverse queries despite only 4 examples in the prompt. FIG. 10A to FIG. 10C Additional examples are presented.
[0120] FIG. 10A to FIG. 10C Example few-shot and zero-shot generated queries randomly sampled from various datasets according to example embodiments are shown. In general, queries generated by few-shot seem to be closer to the original query, while zero-shot queries seem to be mainly questions. For example, in the ArguAna dataset, few-shot queries are generally longer and more statement-like. In addition, for example, zero-shot queries are generally short queries that resemble questions. For the HotpotQA dataset, although both few-shot and zero-shot queries seem to be queries that generate similar questions, few-shot queries seem to generate multi-hop questions, while zero-shot queries seem to generate single-hop questions.
[0121] refer to Fig. 10A, columns 10A-C1 show example paragraphs, columns 10A-C2 show queries generated by the few-shot process, columns 10A-C3 show queries generated by the zero-shot process, and columns 10A-C4 show analysis indicating that in the ArguAna dataset, the few-shot generated queries are more statement-like and longer than the zero-shot generated queries.
[0122] See also Fig. 10B , columns 10B-C1 show example paragraphs, columns 10B-C2 show queries generated by the few-shot process, columns 10B-C3 show queries generated by the zero-shot process, and columns 10B-C4 show an analysis indicating that in the Touché-2020 dataset, few-shot generates more controversial argument-like queries, while zero-shot generates random statements that may be grammatically incorrect.
[0123] refer to Fig. 10C , columns 10C-C1 show example natural segments, columns 10C-C2 show queries generated by the few-shot process, columns 10C-C3 show queries generated by the zero-shot process, and columns 10C-C4 show analysis indicating that in the HotpotQA dataset, few-shot queries can be multi-hop questions, while zero-shot does not generate multi-hop questions.
[0124] Direct use of few-shot examples
[0125] Reference again Figure 8 , can be achieved by using Figure 3 The GTR dual encoder with few-shot examples (e.g., 110 million) shown in is fine-tuned on the task-specific model to analyze the effect of using few-shot examples directly. As expected, it does not provide a large impact on the GTR base, where the average nDCG is reduced from 40.4 to 38.7.
[0126] Query generation statistics
[0127] Fig.11 A table 1100 with average query lengths for various data sets is shown in accordance with an example embodiment. Fig.11In , the length of questions generated by different query generation systems is analyzed. Since the query generation model is fine-tuned on the NQ dataset, NQ-QGen seems to generate short queries, and the generated questions seem to have a similar length to those of NQ. The mean 1115 of NQ-QGen is shown as 9.6. However, compared to NQ-QGen, zero-shot PROMPTAGATOR seems to obtain a larger variance in length. The mean 1110 of zero-shot PROMPTAGATOR is shown as 13.5. In addition, for example, few-shot PROMPTAGATOR seems to provide significantly more variance in the length of the generated queries. The mean 1105 of few-shot PROMPTAGATOR is shown as 17.8.
[0128] Calculate usage and environmental impact
[0129] In some embodiments, 137 billion FLANs based on LaMDA can be used. LaMDA can be pre-trained on a large corpus consisting of 1.56 trillion words, consuming 451 megawatt-hours (MWh) of energy and a carbon footprint of 25.2tCO2e. In PROMPTAGATOR, 29.23 million queries * 2 prompts = 58.46 million queries can be generated, for a total of 610 million words.
[0130] Comparison with existing models
[0131] Several existing neural retrievers adopt a dual encoder architecture that encodes queries and documents independently into dense vectors and retrieves documents using Maximum Inner Product Search (MIPS). Existing approaches have mainly focused on developing better pre-training tasks, improving contrastive negative samples, and improving generalization across different domains.
[0132] Although dual encoders achieve fast retrieval, their expressiveness may be limited due to the fact that their scores are simply the dot product between the query vector and the document vector. Cross-attention models are generally used to re-rank the retrieved candidates, since the cross-attention re-ranker explicitly models the interaction between query tokens and document tokens. Distilling the cross-attention model into a dual encoder has effectively narrowed the gap between dense retrievers and cross-attention models.
[0133] An alternative approach to narrowing the gap between dense retrievers and criss-cross attention models is to allow a certain amount of fine-grained query-document interaction in the retriever. For example, queries and documents can be represented and retrieved using multiple vectors instead of a single vector. ColBER, COIL, and SPLADE use token-level interactions between queries and documents. Since these models do more than just model dot products, the MIPS algorithm cannot be used directly. As a result, these models typically have much higher inference / serving costs than dual encoders.
[0134] In general, the suggested LLM can be used for query generation to improve retrieval reranking. For example, UPR uses the suggested LLM to directly rerank paragraphs. InPars uses few-shot prompts with GPT-3 to generate synthetic data for training a T5-based reranker. Although InPars has been tested on multiple retrieval datasets, it is based on task-independent prompts built by MSMARCO and does not involve task-specific few-shot learning. InPars also focuses exclusively on reranking, while the PROMPTAGATOR approach addresses full-scale retrieval. Pre-trained LLMs have significantly advanced few-shot learning based on prompt strategies such as contextual learning and instruction prompts. Some approaches involve fine-tuning LLMs specifically for few-shot learning, while other approaches do not. As described in this article, the PROMPTAGATOR approach uses LLMs as a few-shot data generator.
[0135] Technical improvements
[0136] The techniques described herein allow for few-shot retrieval, and models can achieve significant improvements using only two (2) to eight (8) examples (referred to herein as prompts) for each task. Typically, a few prompts can train an LLM to automatically generate synthetic data, and these can be used to train a document retrieval model. Existing models attempt to generate training data directly. The term "prompt" as used herein can be general in the zero-shot case and can be task-specific in the few-shot case. The prompt will be combined with a large language model to generate queries for a target task using a corpus of documents.
[0137] PROMPTAGATOR achieves good performance using dual encoders without the need for highly engineered retrieval architectures. For example, no specialized architecture is required and existing large language models can be used.
[0138] The techniques described in this paper can generate queries for retrievers for many diverse target tasks.
[0139] Retrieval can be performed without a resequencer.
[0140] Large language models tuned with hints and instructions address the main limitations of other models that generally focus on question answering datasets.
[0141] The techniques described in this paper show that good query generation can significantly improve the performance of the dual encoder, and therefore, PROMPTAGATOR is significantly more efficient than other LLM applications for retrieval.
[0142] The PROMPTAGATOR-based dual encoder outperforms ColBERTv2 and SPLADEv2 on 11 (11) retrieval tasks.
[0143] PROMPTAGATOR guides LLM to generate text for the target task, further improving the retrieval performance.
[0144] Unlike existing few-shot retrieval approaches that are based on fine-tuning and typically require thousands of labeled examples, our technique uses a small number of task-specific cues to synthetically generate labeled examples.
[0145] The zero-shot PROMPTAGATOR generates more variations in terms of query length.
[0146] Comparison of the distribution of the first words of queries generated by different query generation methods indicates that the distribution of synthetic queries generated by the few-shot PROMPTAGATOR is much closer to the distribution of real queries generated by humans.
[0147] The model has reduced computational usage and environmental impact. For example, the energy cost and carbon footprint of the pre-trained model are 451MWh and 26tCO2e, respectively. In some embodiments, the additional instruction tuning gradient steps for fine-tuning FLAN can be less than 2% of the number of pre-training steps, and thus the estimated additional energy cost can be relatively small.
[0148] Although FLAN is used herein for illustration purposes, the LLM may be any model, including but not limited to T0 or PaLM variants.
[0149] Sample Application
[0150] The model can be deployed on a server.
[0151] The model can be targeted to a specific type of task, knowledge domain, dataset, organization, etc.
[0152] In some embodiments, the model can be applied to retrieve documents that support and counter arguments. For example, the query may be "companies should pay employees at least X", and the documents in the query-document pair may include all posts on the social networking site that support or oppose the query.
[0153] In some embodiments, the model may be applied to perform semantic searches in a database.
[0154] The retrieval task can be performed in different environments, such as in finance, healthcare, education, law, etc. Many such environments involve data with protected information, and organizations may prefer not to share protected data. The model can be provided to the organization (e.g., via an application programming interface (API), as software as a service (SaaS), machine learning as a service (MLaaS), etc.), and the organization can input a small number of prompts to train the model, and apply the trained model, while performing all operations, and the data remains behind the organization's firewall.
[0155] Retrieval tasks may include question-to-document retrieval, question-to-question retrieval, claim-to-document retrieval, argument-support document retrieval, or counter-argument-support document retrieval.
[0156] In some embodiments, the task may be to retrieve information from an email system.
[0157] The term "document" as used herein generally refers to any information retrieved from a document corpus in response to a query. The retrieved information may include a document, a collection of documents, an image, a portion of an image, a video, a sound clip, a text portion of one or more documents from a document corpus, an abstract and / or an extract based on a text portion of one or more documents from a document corpus, etc. In addition, for example, the term "document corpus" generally refers to a document associated with a specific task. For example, a query may be a legal query related to a litigation matter for a certain client, and the document corpus will correspond to client-specific documents, or more generally, to litigation-specific documents. As another example, a query may involve a medical problem related to a specific disease, and the document corpus will correspond to documents related to the disease or related diseases. In general, a user may indicate and / or select a relevant document corpus. A document corpus may include documents stored in one or more databases (e.g., mobile, desktop, cloud, etc.), mail folders, posts on social networking sites, web page collections, scientific databases, images in image libraries, videos in video libraries, audio, music libraries, etc. In general, the model can be applied to any collection of data that is configured to be searchable and from which information can be retrieved. This can include text documents, images, videos, audio files, etc. Additionally, for example, document retrieval can involve speech-to-text transcription.
[0158] In some embodiments, the query may be related to the issue being searched for. For example, a query to an email folder may state "find the information related to the dispute for baggage claim", and the retrieved document may be one or more emails that have been exchanged as part of the dispute process. In some embodiments, the document may be a summary outlining the history of the dispute, with relevant dates and issues. This document may be prepared based on a combination of sentiment analysis, natural language processing (NLP), and the like.
[0159] These and other example applications are considered within the scope of the present disclosure.
[0160] Train machine learning models to generate inferences / predictions
[0161] Fig.12 A diagram 1200 is shown illustrating a training phase 1202 and an inference phase 1204 of a trained machine learning model 1232, according to an example embodiment. Some machine learning techniques involve training one or more machine learning algorithms on an input set of training data to identify patterns in the training data and provide output inferences and / or predictions about (patterns in) the training data. The resulting trained machine learning algorithms may be referred to as trained machine learning models. For example, Fig.12 A training phase 1202 is shown in which one or more machine learning algorithms 1220 are trained on training data 1210 to become a trained machine learning model 1232. Then, during an inference phase 1204, the trained machine learning model 1232 may receive input data 1230 and one or more inference / prediction requests 1240 (possibly as part of the input data 1230), and in response provide one or more inferences and / or predictions 1250 as output.
[0162] Thus, the trained machine learning model 1232 may include one or more models of one or more machine learning algorithms 1220. The machine learning algorithm 1220 may include, but is not limited to, an artificial neural network (e.g., a convolutional neural network, a recurrent neural network, a Bayesian network, a hidden Markov model, a Markov decision process, a logistic regression function, a support vector machine, a large language model (LLM), a neural retrieval system, a suitable statistical machine learning algorithm, and / or a heuristic machine learning system as described herein). The machine learning algorithm 1220 may be supervised or unsupervised, and may implement any suitable combination of online learning and offline learning.
[0163] In some examples, the machine learning algorithm 1220 and / or the trained machine learning model 1232 may be accelerated using an on-device coprocessor, such as a graphics processing unit (GPU), a tensor processing unit (TPU), a digital signal processor (DSP), and / or an application-specific integrated circuit (ASIC). Such an on-device coprocessor may be used to accelerate the machine learning algorithm 1220 and / or the trained machine learning model 1232. In some examples, the trained machine learning model 1232 may be trained, resident, and executed to provide reasoning on a particular computing device, and / or may otherwise make reasoning for a particular computing device.
[0164] During the training phase 1202, the machine learning algorithm 1220 may be trained using unsupervised, supervised, semi-supervised, and / or reinforcement learning techniques by providing at least the training data 1210 as training input. Unsupervised learning involves providing a portion (or all) of the training data 1210 to the machine learning algorithm 1220, and the machine learning algorithm 1220 determining one or more output inferences based on the portion (or all) of the provided training data 1210. Supervised learning involves providing a portion of the training data 1210 to the machine learning algorithm 1220, wherein the machine learning algorithm 1220 determines one or more output inferences based on the portion of the provided training data 1210, and accepting or correcting the output inferences based on correct results associated with the training data 1210. In some examples, supervised learning of the machine learning algorithm 1220 may be governed by a set of rules and / or labels for the training input, and the set of rules and / or labels may be used to correct the inferences of the machine learning algorithm 1220.
[0165] Semi-supervised learning involves obtaining the correct result for a portion (but not all) of the training data 1210. During semi-supervised learning, supervised learning is used for a portion of the training data 1210 that has the correct result, and unsupervised learning is used for a portion of the training data 1210 that does not have the correct result. Reinforcement learning involves the machine learning algorithm 1220 receiving a reward signal about a previous reasoning, where the reward signal may be a numerical value. During reinforcement learning, the machine learning algorithm 1220 may output an inference and receive a reward signal in response, where the machine learning algorithm 1220 is configured to try to maximize the numerical value of the reward signal. In some examples, reinforcement learning also utilizes a value function that provides a numerical value representing the expected sum of the numerical values provided by the reward signal over time. In some examples, the machine learning algorithm 1220 and / or the trained machine learning model 1232 may be trained using other machine learning techniques, including but not limited to incremental learning and curriculum learning.
[0166] In some examples, the machine learning algorithm 1220 and / or the trained machine learning model 1232 may use transfer learning techniques. For example, the transfer learning technique may involve the trained machine learning model 1232 being pre-trained on a data set and additionally trained using the training data 1210. More specifically, the machine learning algorithm 1220 may be pre-trained on data from one or more computing devices, and the resulting trained machine learning model is provided to the computing device CD1, wherein CD1 intends to execute the trained machine learning model during the inference phase 1204. Then, during the training phase 1202, the pre-trained machine learning model may be additionally trained using the training data 1210, wherein the training data 1210 may be derived from the kernel data and non-kernel data of the computing device CD1. Such further training of the machine learning algorithm 1220 and / or the pre-trained machine learning model using the training data 1210 of the CD1 data may be performed using supervised learning or unsupervised learning. Once the machine learning algorithm 1220 and / or the pre-trained machine learning model has been trained on at least the training data 1210, the training phase 1202 may be completed. The resulting trained machine learning model can be used as at least one of the trained machine learning models 1232. The training data 1210 can be a plurality of query-document pairs, where each query-document pair includes a synthetically generated query and a document from a document corpus.
[0167] Specifically, once the training phase 1202 has been completed, the trained machine learning model 1232 may be provided to the computing device if it is not already on the computing device. After the trained machine learning model 1232 is provided to the computing device CD1, the reasoning phase 1204 may begin.
[0168] During the inference phase 1204, the trained machine learning model 1232 may receive input data 1230 and generate and output one or more corresponding inferences and / or predictions 1250 about the input data 1230. Thus, the input data 1230 may be used as an input to the trained machine learning model 1232 for providing corresponding inferences and / or predictions 1250 to kernel components and non-kernel components. For example, the trained machine learning model 1232 may generate inferences and / or predictions 1250 in response to one or more inference / prediction requests 1240. In some examples, the trained machine learning model 1232 may be executed by a portion of other software. For example, the trained machine learning model 1232 may be executed by an inference or prediction daemon so that it is readily available upon request for providing inferences and / or predictions. The input data 1230 may include data from the computing device CD1 executing the trained machine learning model 1232 and / or input data from one or more computing devices other than CD1. For example, input data 1230 may include at least two prompts associated with a retrieval task to be performed on a document corpus associated with the task. Additionally, for example, input data 1230 may include an input query associated with a retrieval task to be performed on a document corpus associated with the task. Other types of input data are also possible.
[0169] Inferences and / or predictions 1250 may include output documents, output query-document pairs, numerical values (e.g., retrieval scores), and / or other output data produced by a trained machine learning model 1232 operating on input data 1230 (and training data 1210). In some examples, the trained machine learning model 1232 may use the output inferences and / or predictions 1250 as input feedback 460. The trained machine learning model 1232 may also rely on past inferences as input for generating new inferences.
[0170] Large language models and / or document retrieval models may be examples of machine learning algorithms 1220. After training, the trained version of the neural network may be an example of trained machine learning model 1232. In this approach, an example of one or more inference / prediction requests 1240 may be a request to respond to an input query, and a corresponding example of an inference and / or prediction 1250 may be a predicted document from a corpus of documents in response to the input query.
[0171] In some examples, a computing device CD_SOLO may include a trained version of the document retrieval model, possibly after training. The computing device CD_SOLO may then receive a request to respond to an input query and use the trained version of the neural network to predict documents responsive to the input query from the document corpus.
[0172] In some examples, two or more computing devices CD_CLI and CD_SRV may be used to provide predicted documents; for example, a first computing device CD_CLI may generate and send a request to a second computing device CD_SRV to respond to an input query. CD_SRV may then use a trained version of a neural network to predict predicted documents responsive to the input query from a corpus of documents and respond to the request from CD_CLI for the predicted documents. CD_CLI may then, upon receiving a response to the request, provide the requested response to the input query (e.g., using a user interface and / or display, a printed copy, an electronic communication, etc.).
[0173] Sample Data Network
[0174] Fig.13 A distributed computing architecture 1300 is depicted according to an example embodiment. The distributed computing architecture 1300 includes server devices 1308, 1310 configured to communicate with programmable devices 1304a, 1304b, 1304c, 1304d, 1304e via a network 1306. The network 1306 may correspond to a local area network (LAN), a wide area network (WAN), a WLAN, a WWAN, an intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 1306 may also correspond to a combination of one or more LANs, WANs, intranets, and / or the public Internet.
[0175] although Fig.13 Only five programmable devices are shown, but the distributed application architecture can serve tens, hundreds, or thousands of programmable devices. In addition, programmable devices 1304a, 1304b, 1304c, 1304d, 1304e (or any additional programmable devices) can be any kind of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head mounted device (HMD), a network terminal, a mobile computing device, and the like. In some examples, such as shown by programmable devices 1304a, 1304b, 1304c, 1304e, the programmable device can be directly connected to the network 1306. In other examples, such as shown by programmable device 1304d, the programmable device can be indirectly connected to the network 1306 via an associated computing device (such as programmable device 1304c). In this example, programmable device 1304c can act as an associated computing device for passing electronic communications between programmable device 1304d and network 1306. In other examples, such as shown by programmable device 1304e, the computing device may be part of and / or located within a vehicle (such as a car, truck, bus, boat or ship, airplane, etc.). Fig.13In other examples not shown, a programmable device may be connected to network 1306 either directly or indirectly.
[0176] The server devices 1308, 1310 may be configured to perform one or more services as requested by the programmable devices 1304a-1304e. For example, the server devices 1308 and / or 1310 may provide content to the programmable devices 1304a-1304e. The content may include, but is not limited to, web pages, hypertext, scripts, binary data (such as compiled software), images, audio, and / or video. The content may include compressed and / or uncompressed content. The content may be encrypted and / or unencrypted. Other types of content are also possible.
[0177] As another example, server devices 1308 and / or 1310 may provide programmable devices 1304a-1304e with access to software for database, search, computing, graphics, audio, video, web / internet utilization, and / or other functions. Many other examples of server devices are possible.
[0178] Computing device architecture
[0179] Fig.14 is a block diagram of an example computing device 1400 according to an example embodiment. Specifically, Fig.14 The computing device 1400 shown in may be configured to perform at least one function for prompt-based query generation, and / or method 1600 , and / or method 1700 .
[0180] The computing device 1400 may include a user interface module 1401, a network communication module 1402, one or more processors 1403, a data storage device 1404, one or more cameras 1412, one or more sensors 1414 and a power system 1416, all of which may be linked together via a system bus, network or other connection mechanism 1405.
[0181] The user interface module 1401 may be operable to send data to and / or receive data from an external user input / output device. For example, the user interface module 1401 may be configured to send data to and / or receive data from a user input device such as a touch screen, a computer mouse, a keyboard, a keypad, a touch pad, a trackball, a joystick, a voice recognition module, and / or other similar devices. The user interface module 1401 may also be configured to provide output to a user display device such as one or more cathode ray tubes (CRTs), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices, whether now known or later developed. The user interface module 1401 may also be configured to generate audible output using devices such as speakers, speaker jacks, audio output ports, audio output devices, headphones, and / or other similar devices. The user interface module 1401 may further be configured with one or more tactile devices that may generate tactile outputs such as vibrations, and / or other outputs that may be detected by touch and / or physical contact with the computing device 1400. In some examples, user interface module 1401 may be used to provide a graphical user interface (GUI), such as, for example, a graphical user interface of a mobile phone device, for utilizing computing device 1400. In some examples, user interface module 1401 may be used to provide a user-adjustable control to receive a threshold distance.
[0182] The network communication module 1402 may include one or more devices that provide one or more wireless interfaces 1407 and / or one or more wired interfaces 1408 that may be configured to communicate via a network. The wireless interface 1407 may include one or more wireless transmitters, receivers, and / or transceivers, such as Bluetooth. TM Transceiver, Transceiver, Wi-Fi TM Transceiver, WiMAX TM Transceiver, LTE TM The wired interface 1408 may include one or more wired transmitters, receivers, and / or transceivers, such as an Ethernet transceiver, a Universal Serial Bus (USB) transceiver, or a similar transceiver that can be configured to communicate via a twisted pair, a coaxial cable, a fiber optic link, or a similar physical connection to a wired network.
[0183] In some examples, the network communication module 1402 can be configured to provide reliable, protected and / or authenticated communications. For each communication described herein, information for facilitating reliable communications (e.g., ensuring message delivery) can be provided, which may be part of a message header and / or message trailer (e.g., packet / message ordering information, encapsulation header and / or encapsulation trailer, size / time information, and transmission verification information (such as cyclic redundancy check (CRC) and / or parity check value)). One or more cryptographic protocols and / or algorithms can be used to ensure that the communication is secure (e.g., encoded or encrypted) and / or decrypted / decoded, such as, but not limited to, the Data Encryption Standard (DES), the Advanced Encryption Standard (AES), the Rivest-Shamir-Adelman (RSA) algorithm, the Diffie-Hellman algorithm, a secure socket protocol (e.g., a secure socket layer (SSL) or a transport layer security (TLS)), and / or a digital signature algorithm (DSA). Other cryptographic protocols and / or algorithms besides those listed herein may be used to protect (and then decrypt / decode) communications.
[0184] The one or more processors 1403 may include one or more general purpose processors and / or one or more special purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), application specific integrated circuits, etc.). The one or more processors 1403 may be configured to execute computer readable instructions 1406 contained in the data storage device 1404 and / or other instructions as described herein.
[0185] The data storage device 1404 may include one or more non-transitory computer-readable storage media that can be read and / or accessed by at least one of the one or more processors 1403. The one or more computer-readable storage media may include volatile and / or non-volatile storage components that may be fully or partially integrated with at least one of the one or more processors 1403, such as optical, magnetic, organic or other memory or disk storage devices. In some examples, the data storage device 1404 may be implemented using a single physical device (e.g., one optical, magnetic, organic or other memory or disk storage unit), while in other examples, the data storage device 1404 may be implemented using two or more physical devices.
[0186] Data storage 1404 may include computer-readable instructions 1406 and possibly additional data. In some examples, data storage 1404 may include storage required to perform at least a portion of the methods, scenarios, and techniques described herein and / or at least a portion of the functionality of the devices and networks described herein. In some examples, data storage 1404 may include storage for trained neural network models 1410 (e.g., models of trained neural networks, such as large language models, dual encoders, etc.). Specifically, in these examples, computer-readable instructions 1406 may include instructions that, when executed by one or more processors 1403, enable computing device 1400 to provide some or all of the functionality of trained neural network models 1410.
[0187] In some examples, computing device 1400 may include one or more cameras 1412. Cameras 1412 may include one or more image capture devices, such as still cameras and / or video cameras, that are configured to capture light and record the captured light in one or more images; that is, camera 1412 may generate images of the captured light. The one or more images may be one or more still images and / or one or more images used in video imagery. Camera 1412 may capture light and / or electromagnetic radiation emitted as visible light, infrared radiation, ultraviolet light, and / or as one or more other frequencies of light.
[0188] In some examples, computing device 1400 may include one or more sensors 1414. Sensors 1414 may be configured to measure conditions within computing device 1400 and / or conditions in the environment of computing device 1400 and provide data regarding these conditions. For example, sensors 1414 may include one or more of the following: (i) sensors for obtaining data about computing device 1400, such as, but not limited to, a thermometer for measuring the temperature of computing device 1400, a battery sensor for measuring the power of one or more batteries of power system 1416, and / or other sensors for measuring the condition of computing device 1400; (ii) identification sensors for identifying other objects and / or devices, such as, but not limited to, radio frequency identification (RFID) readers, proximity sensors, one-dimensional barcode readers, two-dimensional barcode (e.g., quick response (QR) code) readers, and laser trackers, where identification sensors may be configured to read identifiers (such as RFID tags, barcodes, QR codes) and / or other devices and / or objects that are configured to be read and provide at least identification information; (iii) sensors for measuring the location and (iv) an environmental sensor for obtaining data indicative of the environment of the computing device 1400, such as, but not limited to, an infrared sensor, an optical sensor, a light sensor, a biosensor, a capacitive sensor, a touch sensor, a temperature sensor, a wireless sensor, a radio sensor, a motion sensor, a microphone, a sound sensor, an ultrasonic sensor, and / or a smoke sensor; and / or (v) a force sensor for measuring one or more forces (e.g., inertial forces and / or gravity) acting on the computing device 1400, such as, but not limited to, one or more sensors measuring force, torque, ground force, friction in one or more dimensions, and / or a ZMP sensor that identifies a zero moment point (ZMP) and / or the location of the ZMP. Many other examples of sensors 1414 are also possible.
[0189] The power system 1416 may include one or more batteries 1418 and / or one or more external power interfaces 1420 for providing power to the computing device 1400. Each of the one or more batteries 1418 may serve as a source of stored power for the computing device 1400 when electrically coupled to the computing device 1400. The one or more batteries 1418 of the power system 1416 may be configured to be portable. Some or all of the one or more batteries 1418 may be easily removable from the computing device 1400. In other examples, some or all of the one or more batteries 1418 may be located inside the computing device 1400 and may therefore not be easily removed from the computing device 1400. Some or all of the one or more batteries 1418 may be rechargeable. For example, a rechargeable battery may be charged via a wired connection between the battery and another power supply device, such as by one or more power supply devices located outside the computing device 1400 and connected to the computing device 1400 via one or more external power interfaces. In other examples, some or all of the one or more batteries 1418 may be non-rechargeable batteries.
[0190] The one or more external power interfaces 1420 of the power system 1416 may include one or more wired power interfaces, such as USB cables and / or power lines, that implement a wired power connection to one or more power supply devices located outside the computing device 1400. The one or more external power interfaces 1420 may include one or more wireless power interfaces, such as a Qi wireless charger, that implement a wireless power connection to one or more external power supply devices. Once a power connection to an external power source is established using the one or more external power interfaces 1420, the computing device 1400 can draw power from the external power source through the established power connection. In some examples, the power system 1416 may include associated sensors, such as battery sensors associated with one or more batteries or other types of power sensors.
[0191] Cloud-based servers
[0192] Fig.15 A cloud-based server system is depicted according to an example embodiment. Fig.15, the functionality of the prompt-based query generation model, the large language model, and / or the document retrieval model may be distributed among the computing clusters 1509a, 1509b, 1509c. The computing cluster 1509a may include one or more computing devices 1500a, a cluster storage array 1510a, and a cluster router 1511a connected via a local cluster network 1512a. Similarly, the computing cluster 1509b may include one or more computing devices 1500b, a cluster storage array 1510b, and a cluster router 1511b connected via a local cluster network 1512b. Likewise, the computing cluster 1509c may include one or more computing devices 1500c, a cluster storage array 1510c, and a cluster router 1511c connected via a local cluster network 1512c.
[0193] In some embodiments, computing clusters 1509a, 1509b, and 1509c may be a single computing device residing in a single computing center. In other embodiments, computing clusters 1509a, 1509b, and 1509c may include multiple computing devices located in a single computing center, or even multiple computing devices located in multiple computing centers located in different geographical locations. For example, Fig.15 Each of the computing clusters 1509a, 1509b, and 1509c is depicted as residing in a different physical location.
[0194] In some embodiments, the data and services at the computing clusters 1509a, 1509b, 1509c may be encoded as computer-readable information stored in a non-transitory tangible computer-readable medium (or computer-readable storage medium) and accessible by other computing devices. In some embodiments, the computing clusters 1509a, 1509b, 1509c may be stored on a single disk drive or other tangible storage medium, or may be implemented on multiple disk drives or other tangible storage media located at one or more different geographic locations.
[0195] In some embodiments, each of computing clusters 1509a, 1509b, and 1509c may have an equal number of computing devices, an equal number of cluster storage arrays, and an equal number of cluster routers. However, in other embodiments, each computing cluster may have a different number of computing devices, a different number of cluster storage arrays, and a different number of cluster routers. The number of computing devices, cluster storage arrays, and cluster routers in each computing cluster may depend on one or more computing tasks assigned to each computing cluster.
[0196] For example, in computing cluster 1509a, computing device 1500a may be configured to perform various computing tasks of prompt-based query generation, large language models, and / or document retrieval models. In one embodiment, various functionalities of neural networks and / or computing devices may be distributed among one or more of computing devices 1500a, 1500b, and 1500c. Computing devices 1500b and 1500c in respective computing clusters 1509b and 1509c may be configured similarly to computing device 1500a in computing cluster 1509a. On the other hand, in some embodiments, computing devices 1500a, 1500b, and 1500c may be configured to perform different functions.
[0197] In some embodiments, computing tasks and stored data associated with various computing tasks of prompt-based query generation, large language models, and / or document retrieval models may be distributed across computing devices 1500a, 1500b, and 1500c based at least in part on processing requirements of the neural network and / or computing devices, processing capabilities of computing devices 1500a, 1500b, 1500c, latency of network links between computing devices in each computing cluster and between the computing clusters themselves, and / or other factors that may affect cost, speed, fault tolerance, resilience, efficiency, and / or other design goals of the overall system architecture.
[0198] The cluster storage arrays 1510a, 1510b, 1510c of the computing clusters 1509a, 1509b, and 1509c may be data storage arrays that include disk array controllers configured to manage read and write access to groups of hard disk drives. The disk array controllers, alone or in conjunction with their respective computing devices, may also be configured to manage backup or redundant copies of data stored in the cluster storage arrays to protect against disk drive or other cluster storage array failures and / or network failures that prevent one or more computing devices from accessing one or more cluster storage arrays.
[0199] Similar to the way that the functionality of various computational tasks of prompt-based query generation, large language models, and / or document retrieval models may be distributed across the computing devices 1500a, 1500b, 1500c of the computing clusters 1509a, 1509b, 1509c, various active and / or backup portions of these components may be distributed across the cluster storage arrays 1510a, 1510b, 1510c. For example, some cluster storage arrays may be configured to store a first layer of a neural network and / or a portion of the data of the computing device, while other cluster storage arrays may store a second layer of the neural network and / or other portions of the data of the computing device. Additionally, for example, some cluster storage arrays may be configured to store data for an encoder of a neural network, while other cluster storage arrays may store data for a decoder of the neural network. Additionally, some cluster storage arrays may be configured to store backup versions of data stored in other cluster storage arrays.
[0200] The cluster routers 1511a, 1511b, 1511c in the computing clusters 1509a, 1509b, and 1509c may include networking equipment configured to provide internal and external communications for the computing clusters. For example, the cluster router 1511a in the computing cluster 1509a may include one or more Internet switches and routing devices configured to (i) provide local area network communications between the computing device 1500a and the cluster storage array 1510a via a local cluster network 1512a, and (ii) provide wide area network communications between the computing cluster 1509a and the computing clusters 1509b and 1509c via a wide area network link 1513a to the network 1306. Cluster routers 1511b and 1511c may include similar network equipment as cluster router 1511a, and cluster routers 1511b and 1511c may perform networking functions for computing clusters 1509b and 1509b similar to the networking functions that cluster router 1511a performs for computing cluster 1509a.
[0201] In some embodiments, the configuration of the cluster routers 1511a, 1511b, 1511c may be based at least in part on the data communication requirements of the computing devices and the cluster storage arrays, the data communication capabilities of the network equipment in the cluster routers 1511a, 1511b, 1511c, the latency and throughput of the local cluster network 1512a, 1512b, 1512c, the latency, throughput and cost of the wide area network links 1513a, 1513b, 1513c, and / or other factors that may affect the cost, speed, fault tolerance, resilience, efficiency and / or other design criteria of the adjusted system architecture.
[0202] Example methods of operation
[0203] Fig.1616 is a flow diagram of a method 1600 according to an example embodiment. The method 1600 may be performed by a computing device, such as the computing device 1400. The method 1600 for prompt-based query generation may begin at block 1610, where the method involves receiving, by a computing device, at least two prompts associated with a retrieval task to be performed on a document corpus associated with the task.
[0204] At block 1620 , the method involves applying a large language model based on at least two prompts and a document corpus to generate a synthetic training dataset comprising a plurality of query-document pairs, wherein each query-document pair comprises a synthetically generated query and a document from the document corpus.
[0205] At block 1630 , the method involves training a document retrieval model on a plurality of query-document pairs from a synthetic training dataset to take an input query associated with a retrieval task and predict an output document retrieved from a corpus of documents.
[0206] At block 1640 , the method involves providing, by a computing device, a trained document retrieval model.
[0207] Some embodiments involve determining, by a document retrieval model and for each query-document pair, a retrieval score, wherein the retrieval score indicates a relevance of a document in the query-document pair to a query in the query-document pair.
[0208] In some embodiments, the document retrieval model may be a dual encoder model including a query encoder for encoding a query and a document encoder for encoding a document.
[0209] In some embodiments, the dual encoder model may be a joint embedding model where the query embedding for a given query is within a threshold distance of the document embedding for a given document when the given document has a high relevance to the given query.
[0210] In some embodiments, the plurality of query-document pairs may include noisy data. Such embodiments may involve filtering the plurality of query-document pairs to remove the noisy data. Such embodiments may involve fine-tuning a document retrieval model based on the plurality of filtered query-document pairs. In some embodiments, the fine-tuning of the document retrieval model may be based on a standard softmax loss using random negative samples within a batch.
[0211] In some embodiments, filtering includes one or more of length filtering, hint filtering, or round-trip filtering.
[0212] In some embodiments, the large language model may be a fine-tuned language network (FLAN).
[0213] In some embodiments, for each document in the corpus, queries may be sampled based on a distribution including an adjustable temperature hyperparameter. Such embodiments involve maintaining the temperature hyperparameter below a temperature threshold to generate a diverse array of queries.
[0214] In some embodiments, the retrieval task involves one or more of: question-to-document retrieval, question-to-question retrieval, claim-to-document retrieval, argument-supporting document retrieval, or counter-argument-supporting document retrieval.
[0215] Some embodiments involve pre-training a document retrieval model on unsupervised general-domain data.
[0216] In some embodiments, the retrieval task and document corpus can be maintained behind the firewall of the organization. Such embodiments involve providing large language models and document retrieval models to the organization. Receiving, applying, and training can be performed behind the firewall of the organization.
[0217] In some embodiments, at least two prompts are fewer than eight prompts.
[0218] Fig.17 17 is a flow diagram of a method 1700 according to an example embodiment. The method 1700 of applying a trained document retrieval model may be performed by a computing device, such as computing device 1400. The method 1700 may begin at block 1710, where the method involves receiving, by the computing device, an input query associated with a retrieval task to be performed on a document corpus associated with the task.
[0219] At box 1720, the method involves predicting a document from a document corpus by a trained document retrieval model, where the document has a high relevance to an input query, the document retrieval model having been trained on a plurality of query-document pairs from a synthetic training dataset that has been generated by a large language model based on at least two prompts associated with a retrieval task.
[0220] At block 1730 , the method involves providing, by the computing device and in response to the input query, the predicted documents.
[0221] In some embodiments, the document retrieval model may be a dual encoder model including a query encoder for encoding a query and a document encoder for encoding a document.
[0222] In some embodiments, the large language model may be a fine-tuned language network (FLAN).
[0223] In some embodiments, the retrieval task involves one or more of: question-to-document retrieval, question-to-question retrieval, claim-to-document retrieval, argument-supporting document retrieval, or counter-argument-supporting document retrieval.
[0224] In some embodiments, the input query may be a voice input, and wherein the retrieval task may be performed by a computer-implemented intelligent voice assistant.
[0225] As described in this paper, PROMPTAGATOR is a means for few-shot retrieval. It uses only a small number of annotated examples to train task-specific retrievers and rerankers. The few-shot examples amplified by prompt-based LLM query generation simplifies the complexity of training neural retrievers for new tasks and brings significant performance gains. PROMPTAGATOR can be viewed as a dual encoder that distills LLMs to a standard size via prompt-based query generation. Although the distillation process is computationally expensive, it significantly reduces the cost for inference.
[0226] The present disclosure is not limited in terms of the specific embodiments described in this application, which are intended to be illustrative of various aspects. As will be apparent to those skilled in the art, many modifications and variations may be made without departing from its essence and scope. In addition to those methods and devices listed herein, functionally equivalent methods and devices within the scope of the present disclosure will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.
[0227] The above detailed description describes various features and functions of the disclosed systems, devices, and methods with reference to the accompanying drawings. In the drawings, similar symbols generally identify similar components unless the context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized and other changes may be made without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure as generally described herein and shown in the accompanying drawings may be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are expressly contemplated herein.
[0228] With regard to any or all ladder diagrams, scenes and flow charts discussed in the accompanying drawings and herein, each frame and / or communication can represent information processing and / or information transmission according to example embodiments. Alternative embodiments are included in the scope of these example embodiments. In these alternative embodiments, for example, the functions described as frames, transmissions, communications, requests, responses and / or messages may not be performed in the order shown or discussed, including substantially performing at the same time or performing in reverse order, depending on the functions involved. In addition, more or less frames and / or functions can be used with any one of the ladder diagrams, scenes and flow charts discussed herein, and these ladder diagrams, scenes and flow charts can be combined with each other in part or in whole.
[0229] The box representing the processing of information may correspond to a circuit system that may be configured to perform a specific logical function of the method or technology described herein. Alternatively or additionally, the block representing the processing of information may correspond to a module, a segment, or a portion of a program code (including associated data). The program code may include one or more instructions that may be executed by a processor to implement a specific logical function or action in a method or technology. The program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device including a disk or hard drive or other storage medium.
[0230] Computer readable media may also include non-transitory computer readable media, such as non-transitory computer readable media that stores data for a short period of time, such as register memory, processor cache, and random access memory (RAM). Computer readable media may also include non-transitory computer readable media that stores program code and / or data for a longer period of time, such as auxiliary or persistent long-term storage devices, for example, such as read-only memory (ROM), optical disks or magnetic disks, compact disk read-only memory (CD-ROM). Computer readable media may also be any other volatile or non-volatile storage system. Computer readable media may be considered, for example, as a computer readable storage medium or a tangible storage device.
[0231] In addition, the boxes representing one or more information transfers may correspond to information transfers between software modules and / or hardware modules in the same physical device. However, other information transfers may be between software modules and / or hardware modules in different physical devices.
[0232] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are provided for purposes of illustration and are not intended to be limiting, with the true scope being indicated by the claims.
Claims
1. A computer-implemented method for prompt-based query generation, comprising: receiving, by a computing device, at least two prompts associated with a retrieval task to be performed on a corpus of documents associated with the task; applying a large language model based on the at least two prompts and the document corpus to generate a synthetic training dataset comprising a plurality of query-document pairs, wherein each query-document pair comprises a synthetically generated query and a document from the document corpus; training a document retrieval model on the plurality of query-document pairs from the synthetic training dataset to take an input query associated with the retrieval task and predict an output document retrieved from the document corpus; as well as A trained document retrieval model is provided by the computing device.
2. The computer-implemented method of claim 1 , further comprising: A retrieval score is determined by the document retrieval model and for each query-document pair, wherein the retrieval score indicates a relevance of the document in the query-document pair to the query in the query-document pair.
3. The computer-implemented method of claim 1, wherein the document retrieval model is a dual encoder model comprising a query encoder for encoding a query and a document encoder for encoding a document.
4. The computer-implemented method of claim 3, wherein the dual encoder model is a joint embedding model, wherein the query embedding of a given query is within a threshold distance of the document embedding of the given document when the given document has a high relevance to the given query.
5. The computer-implemented method of claim 1 , wherein the plurality of query-document pairs comprises noisy data, and the method further comprises: The plurality of query-document pairs are filtered to remove the noisy data.
6. The computer-implemented method of claim 5, further comprising: The document retrieval model is fine-tuned based on a plurality of filtered query-document pairs.
7. The computer-implemented method of claim 6, wherein the fine-tuning of the document retrieval model is based on a standard softmax loss using random negative samples within a batch.
8. The computer-implemented method of claim 5, wherein the filtering comprises one or more of length filtering, hint filtering, or round-trip filtering.
9. The computer-implemented method of claim 1, wherein the large language model is a Fine-Tuned Language Network (FLAN).
10. The computer-implemented method of claim 1, wherein for each document in the corpus, queries are sampled based on a distribution including a tunable temperature hyperparameter, and the method further comprises: The temperature hyperparameter is maintained below a temperature threshold to generate a diverse array of queries.
11. The computer-implemented method of claim 1, wherein the retrieval task comprises one or more of: question-to-document retrieval, question-to-question retrieval, claim-to-document retrieval, argument-support document retrieval, or counter-argument-support document retrieval.
12. The computer-implemented method of claim 1, further comprising: The document retrieval model is pre-trained on unsupervised general domain data.
13. The computer-implemented method of claim 1, wherein the retrieval task and the document corpus are maintained behind an organization's firewall, and the method further comprises: providing the large language model and the document retrieval model to the organization, and Wherein said receiving, said applying and said training are performed behind said firewall of said organization.
14. The computer implemented method of claim 1, wherein the at least two prompts are less than eight prompts.
15. A computer-implemented method of applying a trained document retrieval model, comprising: Receiving, by a computing device, an input query associated with a retrieval task to be performed on a corpus of documents associated with the task; predicting, by the trained document retrieval model, a document from the document corpus, wherein the document has a high relevance to the input query, the document retrieval model having been trained on a plurality of query-document pairs from a synthetic training dataset that has been generated by a large language model based on at least two prompts associated with the retrieval task; as well as The predicted documents are provided by the computing device and in response to the input query.
16. The computer-implemented method of claim 15, wherein the document retrieval model is a dual encoder model comprising a query encoder for encoding a query and a document encoder for encoding a document.
17. The computer-implemented method of claim 15, wherein the large language model is a Fine-Tuned Language Network (FLAN).
18. The computer-implemented method of claim 15, wherein the retrieval task comprises one or more of: question-to-document retrieval, question-to-question retrieval, claim-to-document retrieval, argument-support document retrieval, or counter-argument-support document retrieval.
19. The computer-implemented method of claim 15, wherein the input query is a voice input, and wherein the retrieval task is performed by a computer-implemented intelligent voice assistant.
20. The computer-implemented method of claim 15, further comprising: determining, by the computing device, a request to respond to the input query; sending the request from the computing device to a second computing device, the second computing device comprising a trained version of the document retrieval model; After sending the request, the computing device receives the predicted document from the second computing device, and Wherein said providing of the predicted document comprises providing the predicted document as received from the second computing device.
21. A computing device comprising: one or more processors; as well as A data storage device, wherein the data storage device has computer executable instructions stored thereon, which, when executed by the one or more processors, cause the computing device to perform functions including the computer-implemented method of any one of claims 1-20.
22. The computing device of claim 21, wherein the computing device is a mobile device.
23. A computer program comprising instructions which, when executed by a computer, cause the computer to perform the steps of the method according to any one of claims 1 to 20.
24. An article of manufacture comprising one or more non-transitory computer-readable media having computer-readable instructions stored thereon, the computer-readable instructions, when executed by one or more processors of a computing device, causing the computing device to perform functions including the computer-implemented method of any one of claims 1-20.
25. A computing system comprising: Means for performing the computer-implemented method of any one of claims 1-20.