Post-facto determination of contributions of source data to outputs of generative ai models
The system determines source data contributions to large language model outputs by extracting claims and assigning scores, addressing biases and hallucinations while ensuring fair compensation for content providers.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PRORATAAI INC
- Filing Date
- 2026-01-23
- Publication Date
- 2026-07-30
AI Technical Summary
Large language models do not track or record how individual pieces of training data influence their outputs, leading to biases and hallucinations, and there is a lack of transparency and compensation for content providers.
A system and method to determine the relative contributions of source data to a large language model's output by extracting claims, associating them with semantically similar source claims, and assigning contribution scores based on factors like source reputation and uniqueness, enabling compensation for content providers.
Enhances transparency and reduces biases and hallucinations by attributing outputs to trusted sources, providing fair compensation to content creators.
Smart Images

Figure US2026012432_30072026_PF_FP_ABST
Abstract
Description
POST-FACTO DETERMINATION OF CONTRIBUTIONS OF SOURCE DATA TO OUTPUTS OF GENERATIVE Al MODELSCROSS-REFERENCE TO RELATED APPLICATION(S)
[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 748,892, filed on January 23, 2025, the entire disclosure of which is incorporated by reference herein. This application is related to: U.S. Provisional Patent Application No. 63 / 617,898, titled “SYSTEMS AND METHODS FOR IMPROVING PERFORMANCE OF A LARGE LANGUAGE MODEL BY CONTROLLING TRAINING CONTENT” filed in the United States Patent and Trademark Office on January 5, 2024; U.S. Provisional Patent Application No.63 / 653,702, also titled “SYSTEMS AND METHODS FOR IMPROVING PERFORMANCE OF A LARGE LANGUAGE MODEL BY CONTROLLING TRAINING CONTENT” filed in the United States Patent and Trademark Office on May 30, 2024; U.S. Patent Application No. 19 / 009,697, titled “SYSTEMS AND METHODS FOR IMPROVING PERFORMANCE OF A LARGE LANGUAGE MODEL BY CONTROLLING TRAINING CONTENT,” filed on January 3, 2025, the entire disclosures of which are incorporated by reference herein.FIELD
[0002] The present disclosure pertains to natural language processing, including systems and methods for assessing the influence of source data on outputs produced by large language models.BACKGROUND
[0003] Large language models are trained on large amounts of source data or training data, where these source data include human-readable text such as books, articles, academic papers, computer source code, and the like. A trained large language model may be used to generate content based on the training data.Copyright infringement hurts authors, publishers, and the public at large.
[0004] The above information disclosed in this Background section is only for I enhancement of understanding of the present disclosure, and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art.SUMMARY
[0005] Aspects of embodiments of the present disclosure relate to systems and methods for determining relative contributions of pieces of training data that influenced content generated by a large language model. A trained large language model represents an accumulation or aggregation of the input training data, and the processes for training large language models generally do not track or otherwise record how individual pieces of training data influence the trained parameters of large language model, let alone which portions of the training data have an impact on a given output generated by the large language model.
[0006] A novel computer system is disclosed for determining contribution scores and assigning compensation to content providers, such as authors, publishers, musicians, song writers etc. The system generally has a neural network trained to extract claims from an output response generated by a large language model, a graph structure stored in a database, the graph structure including a plurality of claims, entities, events, concepts, or beliefs and relationship links between the plurality of claims, the claims being extracted from a plurality of source data, and a contribution engine configured to determine contribution scores of the source data used to train the large language model to generate the output response. A claim is a statement made by or about one or more of the entities, events, concepts or beliefs and each of the claims is connected by a relationship link to one or more other claims, entities, events, concepts or beliefs.
[0007] In an alternative embodiment, a method for determining a plurality of content sources’ relative contribution score to a large language model generated output response is disclosed. In some aspects, contribution and compensation are identified for a given content against 'authoritative' sources (e.g., source documents), for example, some embodiments relate to measuring the contributions of various source documents to a given output response generated by a large language model (e.g., in response to a question from a user).
[0008] The method provides for prompting the large language model to generate the output response, using a claims extraction language model to extract response claims from the output response, using the claims extraction language model to associate semantically similar fully qualified source claims for each response claim of the response (e.g., based on distances between embeddings of the source claims and embeddings of the extracted response claims), associating the semantically similar source claims from the content sources with the associated response claims, and determining the importance of each of the source claims to the associated response claims. The method further discloses the step of ranking the content sources that are more important with higher relative contribution scores, wherein thecontent source providing semantically similar source claims for the most response claims is ranked with a higher contribution score. Alternatively, the content source providing one or more source claims for the most important response claim is ranked with a higher contribution score. Alternatively, the content source providing the most unique source claim is ranked with a higher contribution score. Alternatively, the content source providing the earliest publication date for the source claims is ranked with a higher contribution score. Alternatively, the content source having brand reputation for providing source claims with journalist integrity is ranked with a higher contribution score. Alternatively, the content source with stronger brand pedigree is ranked with a higher contribution score.
[0009] In some embodiments, semantically similar source claims are stored in a vector database in association with embeddings (e.g., vectors representing the semantic meanings of the claims). An embedding for each of the response claims is computed and a search of the vector database for source claims that have embeddings that are similar to the embedding for a given response claim (e.g., based on cosine similarities) is determined.
[0010] The method also provides for assigning a credit score percentage for each content source based on ranking the content sources from highest to lowest contribution score. The content sources having higher contribution score rankings are assigned higher fractional compensation, wherein the highest fractional compensation is assigned to the content source having the overall highest aggregate contribution score for the output response.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings, together with the specification, illustrate exemplary embodiments of the present invention, and, together with the description, serve to explain the principles of the present invention.
[0012] FIG. 1 is a schematic depiction of a process for determining relative contributions of source data to a response generated by a large language model, according to one embodiment of the present disclosure.
[0013] FIG. 2 is a block diagram depicting the architecture of a system for determining relative contributions of source data to a response generated by a large language model, according to one embodiment of the present disclosure.
[0014] FIG. 3 is a flowchart depicting a workflow for generating a claims graph and for determining relative contributions of source data to a response generated by a large language model using the claims graph, according to one embodiment of the present disclosure.
[0015] FIG. 4 is a flowchart depicting a workflow for training a claims extraction language model, according to one embodiment of the present disclosure.
[0016] FIG. 5 is a schematic depiction of a portion of a claims graph and a knowledge graph, according to one embodiment of the present disclosure.
[0017] FIG. 6 is a flowchart depicting a workflow for determining contributions of source data to a response generated by a large language model, according to one embodiment of the present disclosure.
[0018] FIG. 7 provides an example of a workflow for extracting and verifying claims and determining percentage contributions of source documents, according to one embodiment of the present disclosure.
[0019] FIG. 8 is a block diagram illustrating components of a processing circuit or a processor, according to some example embodiments, configured to read instructions from a non-transitory computer-readable medium (e.g., a non-transitory machine-readable storage medium) and perform any one or more of the methods discussed herein.DETAILED DESCRIPTION
[0020] In the following detailed description, only certain exemplary embodiments of the present invention are shown and described, by way of illustration. As those skilled in the art would recognize, the invention may be embodied in many different forms and should not be construed as being limited to the embodiments set forth herein.
[0021] One of the biggest challenges for large language models or LLMs is the presence of biases and hallucinations in their output results and predictions. Biases can arise from several sources, such as biased training data or limitations of model architectures. On the other hand, hallucinations refer to the generation of information or content that appears plausible but is not based on facts or knowledge acquired during training. See, e.g., https: / / www.managementsolutions.com / sites / default / files / minisite / static / 72b0015f-39c9-4a52-ba63-872c115bfbd0 / llm / pdf / rise-of-llm.pdf
[0022] One reason that LLMs hallucinate is that they do not typically quote or cite to their sources. To avoid plagiarism and copyright infringement, LLMs are designed to approximate training data without directly quoting it. While generative models are trained to produce outputs that are not substantially similar to their training inputs, the process of training a generative model often involves making copies of copyrighted data. See, e.g., https: / / suchir.net / fair_use.html
[0023] Biases and hallucination may be combatted by determining which sources contribute to LLM produced output responses. Biased LLM responses and / orhallucinated responses may be reduced by ensuring the output response pulls from trusted sources. For instance, source clarity could determine that a statement is attributed to the New York Times instead of discussion forums like Reddit.Additionally, source transparency in LLM responses allows the model to pull directly from trusted sources instead of abstracting thoughts to avoid allegations of plagiarism and copying, which could potentially reduce bias and hallucination.Creating a model which identifies its sources and can pull directy from the sources not only improves model results, but it provides an opportunity to compensate the authors of the training data.
[0024] Thousands of published authors and creators are requesting payment from model provider companies for the use of their copyrighted works in training artificial intelligence tools, marking the latest intellectual property critique to target Al development. In an open letter (see, https: / / actionnetwork.org / petitions / authors-guild-open-letter-to-generative-ai-leaders) they signed, posted by the Authors Guild Tuesday, the writers accused Al companies of unfairly profiting from their work.“Millions of copyrighted books, articles, essays, and poetry provide the ‘food’ for Al systems, endless meals for which there has been no bill,” the letter said. “You’re spending billions of dollars to develop Al technology. It is only fair that you compensate us for using our writings, without which Al would be banal and extremely limited.” (See https: / / www.cnn.com / 2023 / 07 / 19 / tech / authors-demand-payment-ai / index. htm I)
[0025] LLM enabled search as Perplexity Al (perplexity.ai) incorporated citations into LLM responses. Perplexity Al debuted a revenue-sharing model for publishers after a period of plagiarism accusations. See, e.g., https: / / www.cnbc.com / 2024 / 07 / 30 / perplexity-ai-to-share-revenue-with-publishers-after-plagiarism-accusations.html
[0026] “Decontextualization” refers to a process of rewriting a sentence (or clause) appearing in a larger work such that the rewritten sentence conveys the meaning of the original sentence without the context of the larger work. Choi et al. (Choi, Eunsol, et al. "Decontextualization: Making sentences standalone." Transactions of the Association for Computational Linguistics 9 (2021 ): 447-461.) introduced the manual task of decontextualizing a sentence in context and rewriting it in a way that its meaning is preserved, and it can be interpreted out of context. See also, Rashkin, Hannah, et al. "Measuring attribution in natural language generation models." Computational Linguistics 49.4 (2023): 777-840. and Min, Sewon, et al. “FACTSCORE: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation” arXiv preprint arXiv:2305.14251 (2023).:
[0027] Large language models are trained on large amounts of source data or training data, where these source data include human-readable text such as books, articles, academic papers, computer source code, and the like. The large language model may be implemented using a neural network using, for example, a transformer deep neural network architecture as described in Vaswani, Ashish, et al . (2017). “Attention is All you Need” (PDF). Advances in Neural Information Processing Systems. 30. Curran Associates, Inc. Training such a large language model includes updating the parameters of the neural network weights and biases of the neural network in response to source data. Such large language models are thereby trained to compute or predict the likelihood that a given token (e.g., a word) will appear next in a sequence of tokens (or sequence of words). Language models may be further trained or fine-tuned to generate responses to instructions or messages (sometimes referred to an instruction tuning), such that users and developers can interact with a language model by providing natural language instructions (e.g., “write an essay about the main themes represented in Shakespeare’s Hamlet”).
[0028] The content of the source data includes statistical relationships between words or tokens representing facts, opinions, and the like in the real world. For example, in the corpus of newspaper articles, the words “Obama,” “Trump,” and “Biden” will, statistically, all be commonly found in association with the word “president,” and the word “Paris” will, statistically, be found in association with the word “France” more often than with the word “president.”
[0029] A large language model will therefore tend to generate responses in accordance with the content of the source data and based on the relationships between words or tokens found in the source data or training data. For example, a large language model trained only on text from cookbooks may tend to generate responses that look like recipes (and may have little to no statistics regarding the word “president”) and a large language model trained only on text from newspapers may tend to generate responses that look like newspaper articles.
[0030] In many cases, the training of large language models (e.g., the computation of the parameters of the neural networks) does not keep track of how the weights are updated based on each piece of source data. Furthermore, during inference (e.g., when the language model is generating a response to some input) it is difficult to determine what degree the generated response was influenced by any specific piece of source data (e.g., when generating a summary of a president’s term in office, it is difficult to determine which specific books, articles, encyclopedias, and the like influenced the statistics stored in the large language model to produce the given result).
[0031] Therefore, aspects of embodiments of the present disclosure relate to automatically, and in real-time, determining the significance of each source’s contribution to a response generated by a large language model.
[0032] FIG. 1 is a schematic depiction of a process for determining relative contributions of source data to a response generated by a large language model, according to one embodiment of the present disclosure. As shown in FIG. 1, a user may provide a prompt 101 (“why is education important?”) to an instruction-tuned large language model (LLM) or chatbot LLM to as and the LLM may generate a response 103 (‘Education is important for a multitude of reasons. Firstly, it has a significant impact on the economy. According to Investopedia, 'A country's economy becomes more productive as the proportion of educated workers increases since educated workers can more efficiently carry out tasks that require literacy and critical thinking" [1], This suggests that education is an investment in human capital, similar to an investment in better equipment. Furthermore, research from Georgetown University's Center on Education and the Workforce found that a bachelor's degree adds $2.8 million to a person's lifetime earnings [2], This highlights the financial benefits of pursuing higher education. . . .’).
[0033] Various claims 105 are then extracted from the generated response 103. FIG. 1 shows three example claims: “Education has a significant impact on the economy.” “Education is an investment in human capital.” and “Education is similar to an investment in better equipment.” (The response 103 includes many other claims, which are not explicitly shown in FIG. 1.)
[0034] A “claim” is a statement made by or about entities, events, concepts in given context(s), including beliefs. Claims may be classified based on their type and interpretation. Some examples of claims include:
[0035] Claim of Fact: Assert something is true or false based on evidence; something that can be verified, e.g., Water is essential for human life.
[0036] Claim of Value: Judgements based on values, ethics, etc.; opinion about something, e.g., I prefer oranges over apples.
[0037] Claim of Policy: A course of action that should be taken or avoided, problem-solving and calls for changes in laws, policies, etc. e.g., California state should invest more in renewable energy sources.
[0038] Claim of Definition: definition or categorization of a concept or term; identifying what something means and how it should be understood, e.g., Penguins are birds that don’t fly.
[0039] Claim of Cause and Effect: arguing one event or condition leads to another, often requiring evidence linking the cause to the effect, e.g., Smoking leads to lung cancer.
[0040] The response claims extracted from the response are then compared to source claims in a claims graph stored in a vector database 107 (other types of databases could be used such as relational databases, document stores, and the like) based on semantic similarity. Semantic similarity is determined by converting the response claims into embeddings (or vectors) and searching the vector database 107 for source claims (extracted from training data sources) that have vectors that are close (e.g., based on cosine distance) to the response claim vectors. The source claims may be indexed in the database 107 based on their vectors. The claims graph stores relationships between claims that were extracted from source data (e.g., the source data that was used to train the LLM that generated the response 103). In effect, the database 107 is searched for claims appearing in the source data that match (e.g., are semantically similar to) the claims extracted from the response 103. A contribution engine 111 determines the relative contributions of different sources (example relative contributions shown in FIG. 1 as “Microblog Q: 0.53; Article A: 0.23; Article C: 0.15; Book K: 0.09”).
[0041] FIG. 2 is a block diagram depicting the architecture of a system 200 for determining relative contributions of source data to a response generated by a large language model, according to one embodiment of the present disclosure. Various components of the system 200 shown in FIG. 2 are implemented using a computer system that includes one or more processors (e.g., central processing units, graphics processing units, neural network accelerator units, and the like) and one or more memories storing instructions that, when executed by the one or more processors, implement the system 200. The one or more memories include non-transitory computer readable media that store data in a persistent manner. Various methods and workflows of embodiments of the present disclosure described in more below are implemented using computer instructions stored in the one or more memories and executed by the one or more processors. For example, one or more large language models are defined by parameters (e.g., trained weights and biases) that are stored in the one or more memories, where generating an output of the large language model (e.g., executing or performing inference using the large language model) includes performing computations based on input values (e.g., vector representations of input text) and the parameters stored in the one or more memories. The one or more large language models may be further configured using customized prompts (e.g., system prompts) that are also stored in the one or more memories and provided as part of an input to a large language model (not shown in FIG. 2).
[0042] As shown in FIG. 2, an input user prompt 201 (e.g., the text “Why is education important?” 101 as shown in FIG. 1) may be provided from a user throughan API gateway 210 (e.g., an HTTP API providing an endpoint for a POST message). The prompt is provided to a retrieval augmented generation (RAG) inference module 220. The RAG inference module 220 may perform prompt processing to initially clean the user prompt 201 (e.g., screen for dangerous text, convert text to token embeddings in a vector space, and the like) and to extract information for retrieving and ranking relevant documents from data storage 230 (or blob store), where the retrieved documents, the prompt 201 , and a system prompt (e.g., “You are a helpful artificial intelligence assistant”) are provided to a large language model to generate a response.
[0043] In more detail, publisher content 241 from publisher sources 240 is ingested, processed, and stored in the data storage 230 and will generally be referred to as documents or source data. The source data may include, but are not limited to, books, articles, blog posts, microblog posts, and the like. These source data may then be indexed (e.g., including generating search indexes based on keywords and RAG vector indexed based on creating vector representations that summarize the contents of the documents), where the search index and / or the RAG vector index are used to identify relevant source data to be provided to the LLM to generate a response 203.
[0044] While FIG. 2 shows an approach based on RAG inference, embodiments of the present disclosure are not limited thereto and may also be applied in circumstances where there is no separate document retrieval and the LLM is provided with only, for example, the user prompt 201 and a system prompt to generate the response 203.
[0045] The LLM output response 203 (and, optionally, the user prompt 201 ) are provided to a contribution engine 250 to compute the relative contributions of different source data to the generated response 203. In some embodiments of the present disclosure, the contribution engine 250 determines these relative contributions based on a claims graph 260 stored in a database, where the claims graph 260 includes claims extracted from the source documents stored in the data store 230.
[0046] FIG. 3 is a flowchart depicting a workflow 300 for generating a claims graph and for determining relative contributions of source data to a response generated by a large language model using the claims graph, according to one embodiment of the present disclosure. At 301 , source content or source data are received from publisher sources 240 in the form of the publisher content 241 (e.g., books, articles, blog posts, and the like), where the source content is stored (at 303) in a data store 230 (e.g., a blob store) after, in some cases, performing pre-processing of the data (e.g., extraction of the data into a standardized format such as plain text, normalization of character encodings, and the like).
[0047] At 305, source claims are extracted from source data and indexed by a claim extraction and indexing module 270 (e.g., implemented by a computer system including a processor and memory, where the memory stores instructions that, when executed by the processor, cause the processor to perform operations to implement actions taken by the claim extraction and indexing module 270).
[0048] In order to attribute ideas to sources, snippets or chunks of the LLM output responses need to be broken down into individual claims. In some cases, this involves a process called decontextualization, where a chunk of text (e.g., a sentence or clause) appearing in a source document is rewritten such that the rewritten text conveys the meaning of the original chunk without the context of the full source document.
[0049] Some aspects of embodiments of the present disclosure relate to performing claims extraction using a claims extraction language model.
[0050] FIG. 4 is a flowchart depicting a workflow 400 for training a claims extraction language model, according to one embodiment of the present disclosure. FIG. 4 also shows portions of a source data ingestion process, including generating the RAG vector index described above based on the input source data. As shown in FIG. 4, a claims extraction and indexing module 410 extracts claims from the input source data after splitting the input source data into paragraphs and sentences (or chunks) and stores the extracted claims in a claims index 420.
[0051] In some embodiments, the claims extraction and decontextualization is performed using a large language model (e.g., the Llama 3.1 405B language model from Meta® or general-purpose large language model) that is prompted to extract claims from some given text.
[0052] In some embodiments, the prompt for extracting claims includes instructions to the LLM to use a summary and a context of a portion of document to extract a claim from each input sentence, where each claim is understandable on its own without additional context and provides examples of claims extracted from some input sentences.
[0053] One example of a claims extraction prompt is shown in Table 1 below, where the claims extraction prompt may be used with a generic instruction tuned large language model (e.g., Llama 3.1 405B) or a fine-tuned language model (e.g., fine-tuned starting with Llama 3.1 7B, where the fine tuning of a large language model is described in more detail below with respect to FIG. 4):Table 1Consider the summary and context provided below which are, respectively, a summary and an excerpt of a document with title: {{ title }}Use the summary and context to interpret each of the sentences in the list of sentences provided below.The goal is to extract independent and self-contained claims from each sentence. Each extracted claim should be understandable on its own without additional context.This means all entities must be referred to by name rather than definite noun phrases (e.g., 'the teacher') whenever possible. If a definite noun phrase is used, be sure to add modifiers (e.g., an embedded clause, a prepositional phrase, etc.) to appropriately narrow down the scope of the noun phrase.Each extracted claim must be a complete sentence including a verb that is either explicitly stated in the sentence or strongly implied by the context.Pay particular attention to adjectives and relative clauses that modify a noun or noun phrase. Most often, adjectives and subordinate clauses make an implicit claim about the state or existence of the modified noun or noun phrase. Such an implicit claim must be extracted as a claim whenever the same information is not mentioned elsewhere in the context.When breaking down a sentence into claims observe the following guidelines:1. In general, if the sentence lists two or more qualities or attributes each of these should be extracted as individual claims. However, if the list of qualities or attributes, taken as a unit, is part of a chain of argumentation it is best to leave the list intact (claims or not split) within the chain of argumentation.2. If the sentence describes a change or effect and then lists two or more objects affected by the change or effect, then a claim should be extracted for each of the objects.3. If the sentence describes a multi-step process it will be necessary for the claims extracted to make reference to the entire process. If the multi-step process is not overly lengthy then perhaps a single claim will be a good fit for the extraction.However, if the multi-step process is sufficiently lengthy it could be better to extract multiple claims. The downside to extracting multiple claims is the overhead of making each claim self-contained by summarizing the prior steps in the multi-stepprocess. In all cases it is crucial to never extract claims without the context in which they originally occurred.5. Each extracted claim must include relevant time and location information as implied by the summary and the context.6. Quotations must be extracted verbatim with the name of the source whenever possible.7. Avoid redundant or trivial claims or claims about existence of universally known entities.EXAMPLES STARTEXAMPLE 1Sentence: 'The team's greatest accomplishment was to participate in the World Cup Final in 1998.'Claims:1. The team participated in the World Cup Final in 1998; and,2. The participation in the World Cup Final in 1998 is the team's greatest accomplishment.Note that the second claim is extracted because the adjective 'greatest' (and nothing else) implies it.EXAMPLE 2Sentence: 'The Lexus LS400 / LS430, a rear-drive V-8 powered luxury sedan, costs tens of thousands less than an equivalent Mercedes-Benz, BMW or Jaguar.' Claims:1. The Lexus LS400 / LS430 is a rear-drive luxury sedan;2. The Lexus LS400 / LS430 is a V-8 powered luxury sedan; and,3. The Lexus LS400 / LS430 costs tens of thousands less than an equivalent Mercedes-Benz, BMW or Jaguar.Note that the first two claims are extracted from a relative clause that provides nontrivial information about the subject of the sentence. Also note that each of the first two claims speak to one attribute or feature of the subject.EXAMPLES END
[0054] In some embodiments, the prompt further includes instructions regarding the formatting of the output, such as instructing the language model to output as an object in JavaScript Object Notation (JSON). The prompt further includes fields for the input to be supplied (e.g., the input text that claims will be extracted from).
[0055] Some aspects of embodiment of the present disclosure further relate to training or fine-tuning a low-parameter language model 430 (e.g., Llama 3.1 7B from Meta®, Phi-3B from Microsoft®, or Gemma-3B from Google® although not limited thereto) to train or fine-tune a claims extraction language model 440 that is trained for the specific task of extracting claims from input text (e.g., update, based on backpropagation, the parameters of a pre-trained neural network implementing the language model based on the additional training data associated with the task of claims extraction to improve the quality of the generated outputs on this task).
[0056] In some embodiments, the training data for training or fine tuning (at 445) the claims extraction language model 440 includes synthetic data (e.g., the claims extracted from source data by the large language model prompted as described above), human-generated claims (e.g., claims manually extracted by humans from source data provided through a user interface such as label studio 450, and ensembles thereof.
[0057] The claims extraction language model 440 (a language model that is finetuned, as discussed above, to extract claims from input text) may then be used in place of the large language model (e.g., the general-purpose language model such as Llama 3.1 405B) to extract claims from source data. In some embodiments, human feedback may be used to revise the claims extracted by the claims extraction language model 440 to generate additional training data for re-training or further fine-tuning the claims extraction language model at 445 to generate claims that more accurately represent the content of the source data.
[0058] At 307, the clams extracted from the source data are organized into a claims graph. Entities (concepts, in general) are organized into a knowledge graph based on relationships between those entities, and statements made by and about those entities are organized and indexed in claims graphs. Claims are represented within the claims graph as corresponding nodes, and edges of the graph represent relations from claims to other claims in the claims graph or to entities and / or events in the knowledge graph. These types of relations include subsumption (broader, narrower), entailment, and the like.
[0059] FIG. 5 is a schematic depiction of a portion of a claims graph and a knowledge graph, according to one embodiment of the present disclosure.
[0060] As shown in FIG. 4, a knowledge graph (KG) I names, entities and relations (NER) extraction and indexing module 460 splits extracts named entities, events, and relations from the source data and indexes these entities into a knowledge graph 510, where relationships 511 between entities 513 and events 515 (e.g., within and across sentences) are also extracted from the documents. In some embodiments, a trained knowledge graph extraction language model is used toextract the knowledge graph from the source documents. In more detail, like the claims extraction language model, a large language model (e.g., Llama 3.1 405B) may initially be prompted to extract entities and events and relationships between entities and events from some given text, and training data may further be used to train or fine-tune a low-parameter language model (e.g., Llama 3.1 7B) to train a knowledge graph extraction language model to extract a knowledge graph 510.
[0061] In some embodiments, the entities extracted from some given text are disambiguated within the document (e.g., entities referred to by different names such as “Biden” and “the President”) and also disambiguated against a knowledge base (e.g., disambiguated to “Joseph Robinette Biden”).
[0062] As noted above, the claims extraction and indexing module 410 extracts fully-qualified claims 520 from sentences in the document and links these claims to associated entities 513 and events 515, where a fully-qualified or de-contextualized claim represents a complete and comprehensive statement. In some embodiments, the claims further include attribution to a source, where the source (author, publisher, etc.) are also entities in the knowledge graph.
[0063] Claims are categorized into different types (e.g., fact, value, policy, definition, etc.) and linked to other claims based on a claims taxonomy, where the claims are organized based on their relationships to one another. In the example of FIG. 5, the taxonomy indicates that Claim 1.1 is supported by Claim 2.3, and that Claim 2.1 provides details of claim 1.2. In some embodiments, the claims graph is computed automatically by a language model that is configured to (e.g., prompted to and / or fine-tuned to) identify any relationships between given pairs of claims (e.g., “supported by”, “provides details about”, and the like), where the language model is provided with a current claim to be added to the claims graph and one or more existing claims within the claims graph.
[0064] In some embodiments, the claims are indexed (e.g., at 305) using appropriate embeddings (e.g., vector embeddings) to capture their semantic representation, and the claims are assigned weights based on their importance, uniqueness, and pedigree (e.g., reputation of the source or publisher) in a collection of source documents or source data.
[0065] Referring again to FIG. 3, at 309, an LLM is prompted (e.g., with input user prompt 101 or 201) to generate a response 311 (e.g., response 203 of FIG. 2).
[0066] Continuing the example shown in FIG. 1 , suppose that the LLM response 103 to the user prompt 101 “Why is education important?” continued with another paragraph shown in Table 2:Table 2In addition to its economic benefits, education also has a profound impact on an individual's quality of life. Studies have shown that educated individuals tend to live longer, healthier lifestyles. For example, a report from the CDC's National Center for Health Statistics found that people with a bachelor's degree or higher live about nine years longer than people who don't graduate high school [5], Moreover, education empowers individuals by providing self-determination, self-confidence, resilience, and motivation, leading to greater happiness and life satisfaction [3], This is supported by research from various universities, including Oxford University, University of Arizona, and University of California Davis, which suggests that learning is a key ingredient of happiness and thriving.
[0067] The response of the LLM may then be stored at 313 and then claims are extracted from the response at 315 using, for example, a claims extraction language model (e.g., the claims extraction language model 440 described above with respect to FIG. 4).
[0068] From this second paragraph of the LLM response, the following claims shown in Table 3 are automatically extracted by the claims extraction language model (e.g., at 315):Table 3Sentence (2, 1): In addition to its economic benefits, education also has a profound impact on an individual's quality of life.• Claim (2, 1, 1): Education has a profound impact on an individual's quality of life, in addition to its economic benefits.Sentence (2, 2): Studies have shown that educated individuals tend to live longer, healthier lifestyles.• Claim (2, 2, 1): Educated individuals tend to live longer lifestyles.• Claim (2, 2, 2): Educated individuals tend to live healthier lifestyles.Sentence (2, 3): For example, a report from the CDC's National Center for Health Statistics found that people with a bachelor's degree or higher live about nine years longer than people who don't graduate high school [5],• Claim (2, 3, 1 ): A report from the CDC's National Center for Health Statistics found that people with a bachelor's degree or higher live about nine years longer than people who don't graduate high school. Sentence (2, 4): Moreover, education empowers individuals by providing self-determination, self-confidence, resilience, and motivation, leading to greater happiness and life satisfaction [3],• Claim (2, 4, 1): Education empowers individuals by providing self- determination.• Claim (2, 4, 2): Education empowers individuals by providing selfconfidence.• Claim (2, 4, 3): Education empowers individuals by providing resilience.• Claim (2, 4, 4): Education empowers individuals by providing motivation.• Claim (2, 4, 5): The empowerment provided by education leads to greater happiness and life satisfaction.Sentence (2, 5): This is supported by research from various universities, including Oxford University, University of Arizona, and University of California Davis, which suggests that learning is a key ingredient of happiness and thriving.• Claim (2, 5, 1): Research from Oxford University, University of Arizona, and University of California Davis suggests that learning is a key ingredient of happiness and thriving.
[0069] Optionally, at 317, a response claims graph is generated from the claims extracted from the response.
[0070] At 319, the claims extracted from the response (and / or the response claim graph) are used by the contribution engine 250 to retrieve semantically similar source claims stored in the claims graph generated at 307. Table 4, below, shows examples of the retrieved relevant source claims for Claim (2, 2, 2):Table 4Claim (2, 2, 2): Educated individuals tend to live healthier lifestyles.• Source D01 : “Learning Is A Sure Path To Happiness: Science Proves It” https: / / www.forbes.com / sites / tracybrower / 2021 / 10 / 17 / learning-is-a-sure-path- to-happiness-science-proves-it / • Claim C12 (7-4-2): “More education tends to result in longer lifespans.”[0.1233]• Claim C15 (7-10-5): “When countries support greater educational attainment, their citizens are healthier.” [0.2546]• Source D04: “9 Secrets To Living Longer”https: / / time.com / 81573 / how-to-live-longer / • Claim C3 (4-2-1 ): “Education is correlated with a longer life.” [0.1434] • Claim C4 (4-3-1 ): “A 2012 report from the CDC’s National Center for Health Statistics found a correlation between education level and life expectancy.” [0.335]• Claim C5 (4-3-2): “People with a bachelor’s degree or higher live about nine years longer than people who don’t graduate high school.” [0.10] • Claim C6 (4-6-3): “Educated people are more likely to make healthier lifestyle choices.” [0.0134]
[0071] The contribution engine 250 then determines, at 321 , the contribution score of each source claim to the associated response claim, and computes, at 323, a total contribution score for each source.
[0072] FIG. 6 is a flowchart depicting a workflow 600 for determining contributions of source data to a response generated by a large language model, according to one embodiment of the present disclosure, where workflow 600 provides additional details on the operations at 321 and 323 shown in FIG. 3.
[0073] At 610, the contribution engine 250 retrieves semantically similar source claims stored in the claims graph (as shown at 319 of FIG. 3). In some embodiments, these semantically similar source claims are stored in a vector database in association with embeddings (e.g., vectors representing the semantic meanings of the claims). Accordingly, the contribution engine 250 may compute an embedding for each of the response claims, search the vector database for source claims that have embeddings that are similar to the embedding for a given response claim (e.g., based on cosine similarities between the embeddings, such as computed with a dot product). At 620, the contribution engine 250 evaluates the response and scores the claim features 630. The overall contribution score of a given source claim that is associated with a response claim depends on a plurality of factors or features 630. In the example shown in FIG. 6, these factors or features 630 include a source pedigree 631 (e.g., giving a higher contribution feature score to more reputable sources), a response claim importance 632 (e.g., giving higher contribution feature score when the claim is more important to the overall response generated by the LLM versus being an aside or digression in the response), a source claim uniqueness 633 (e.g., how frequently does the claim appear within the sources, giving a higher contribution feature score to unique claims), and a source claim publishing date (e.g., giving a higher contribution feature score to claims that came from earlier sources rather than later sources that may have taken the concept from the earlier source). Embodiments of the present disclosure are not limited to the four factors or features 630 described above, and other factors or features may be included in determining a contribution score of a given source claim.
[0074] FIG. 7 provides an example of a workflow 700 for extracting and verifying claims and determining percentage contributions of source documents, according to one embodiment of the present disclosure. As used herein, verifying a response claim and a source claim refers to confirming that the meaning of an associated source claim is logically consistent with the meaning of the response claim (e.g., the claim appearing in the LLM output), as described in more detail below.
[0075] FIG. 7 depicts similar steps of generating 710 a response 720 using a large language model and extracting 730 claims from the response, as describedabove. FIG. 7 further depicts generating pairings 740 of a response claim and an associated source claim (a (response claim, source claim) pair), and verifying each pair at 750. As noted above, any given response claim may be associated with one or more source claims (e.g., k source claims), and therefore one or more corresponding pairings would be generated (e.g., k pairings, where each pairing has the same response claim and one of the k different source claims).
[0076] Continuing the above example, Table 5 shows a process of verifying that one of the source claims (source claim C15) supports Claim (2, 2, 2), where the verification is performed by the contribution engine 250 by supplying the (response claim, source claim) pair as part of an input to a language model that is prompted and / or fine-tuned to compute determinations as to whether a claim is supported: Table 5LLM Response Claim (2, 2, 2): Educated individuals tend to live healthier lifestyles.• Source D01 : “Learning Is A Sure Path To Happiness: Science Proves It” https: / / www.forbes.com / sites / tracybrower / 2021 / 10 / 17 / learning-is-a-sure-path- to-happiness-science-proves-it / Source claim C15: “When countries support greater educational attainment, their citizens are healthier.” [0.2546]Verification:Step 1 :“A relationship between education and health at a country level implies a similar relationship at an individual level.”Thus“Educated individuals tend to be healthier.”Step 2:“Being healthier implies living healthier lifestyles.”Thus:“Educated individuals tend to live healthier lifestyles.”Therefore, the response claim is supported by the source claim.
[0077] In some embodiments, the contribution engine 250 performs claim verification to filter out source claims that do not support the response claim, such that the associated source claims that do not support the response claim do not receive contribution credit.
[0078] At 640, the contribution engine 250 calculates a contribution score for each source claim based on a combination of the contribution feature scores. In some embodiments, the contribution score is computed as a linear combination (or weighted sum) of the contribution feature scores. In some embodiments, thecontribution score is computed as a product of the contribution feature scores. At 650, the contribution engine ranks the sources based on the contribution scores of the source claims, and at 660 calculates a percentage contribution by each source.
[0079] In some embodiments, the percentage contribution by each source is supplied to a fractional compensation engine 280, which computes (at 325 of FIG. 3 and at 670 of FIG. 6) fractional compensation or fractional credit to each source based on the percentage contributions computed by the contribution engine 250.
[0080] As a concrete example, a prompt to an LLM may be: “What is the best time of year to visit Seoul, Korea?”
[0081] The LLM response to the prompt may be: “The best time to visit Seoul is spring or autumn. Spring offers pleasant temperatures and blooming cherry blossoms, while autumn brings cool weather and stunning fall foliage. These periods are less crowded than summer, providing an ease of exploration.”
[0082] This LLM response may be broken down into nine response claims (RC1 through RC9):RC1 : The best time to visit Seoul is Spring or Autumn.RC2: Spring in Seoul offers pleasant temperatures; and,RC3: Spring in Seoul offers blooming cherry blossoms.RC4: Autumn in Seoul brings cool weather; and,RC5: Autumn in Seoul brings stunning fall foliage.RC6: Spring in Seoul is less crowded than summer;RC7: Autumn in Seoul is less crowded than summer;RC8: It is easier to explore Seoul in the Spring because it is less crowded than the summer; and,RC9: It is easier to explore Seoul in the Spring because it is less crowded than the summer.
[0083] These response claims can then be matched to source claims SC in the claims graph generated from the source documents, such as:SC1 : The best times of year to visit Seoul are Spring and Autumn. (RC1)SC2: Travelers enjoy visiting in the spring for the pleasant temperatures and beautiful blooming cherry blossoms. (RC2, RC3)SC3: However, some prefer Autumn for the cooler temperatures. (RC4)SC4: Autumn in Seoul is known for its stunning fall foliage. (RC5)SC5: Avoid the summer because it is very hot and crowded, making it difficult to explore. (RC6, RC7, RC8, RC9)
[0084] The identified source claims can then be connected to underlying source documents S:S1: Best Times to Visit Seoul - U.S. News Travel (SC1-RC1, SC4-RC5)S2: What are the best months to visit Korea? - Reddit r / koreatravel (SC5 - RC6, RC7, RC8, RC9)S3: The seasons of Seoul - Lonely Planet (SC2-RC2,RC3, SC3- RC4)
[0085] The contribution engine 250 then calculates contribution scores for each source claim and source:S1: 40% (high pedigree and most important claim)S2: 25% (many claims but low pedigree)S3: 35% (good pedigree and second most claims)
[0086] FIG. 7 further provides an example of a collection of verified response claims RC 780 in the response, each having an associated list of verified source claims that support those response claims. In the example of FIG. 7, there is a larger number of possible sources (e.g., S1 through S99) and the assignment of fractional contributions to six sources 790: S5, S10, S24, S37, S55, and S99.
[0087] While various aspects of the present disclosure relate to calculating contribution scores for source claims including source claims that were extracted from publisher content 241 provided by publisher sources 240, embodiments of the present disclosure are not limited thereto. For example, some embodiments of the present disclosure relate to determining contributions of source documents to a response generated by a large language model, where the source documents are provided together with the LLM response. In some embodiments, this functionality is exposed to users through an application programming interface, thereby providing attribution-as-a-service.
[0088] For example, some language models are prompted to perform research, such as by supplying queries to internet search engines to obtain documents (e.g., web pages, books, published papers, and the like), ingest the documents, and generate a response from a prompt or context that includes the content of those documents (or summaries thereof). These responses may further include hyperlinks and / or citations to the source documents. In such a case, some aspects of embodiments of the present disclosure relate to receiving the LLM response at 311 of FIG. 3, retrieving copies of the documents (by retrieving the files located at the URLs of the hyperlinks or by retrieving a document based on its citation, such as by providing the citation to an internet search engine), and providing the retrieved documents as the source content 301 shown in FIG. 3. Accordingly, the documents identified in the LLM response are processed by extracting source claims from those documents and by generating a claims graph. The remaining operations of FIG. 3 proceed as described above, such that these embodiments of the present disclosure compute contribution scores for source documents identified in an LLM response,even if those source documents were not part of an original collection of publisher content provided to the system 200.
[0089] According to one embodiment of the present disclosure, a system includes: a neural network trained to extract claims from an output response generated by a large language model, a graph structure stored in a database, the graph structure including a plurality claims, entities, events, concepts, or beliefs and relationship links between the plurality of claims, the claims being extracted from a plurality of source data, and a contribution engine configured to determine contribution scores of the source data used to train the large language model to generate the output response.
[0090] The graph structure may be represented in a machine-readable format.
[0091] The relationship link may be representative of a subsumption or an entailment.
[0092] Each of the claims may be classified as a claims type.
[0093] The claims type may be a claim of fact, claim of value, claim of policy claim of definition, or claim of cause or effect.
[0094] A claim may be a statement made by or about one or more of the entities, events, concepts or beliefs and wherein each of the claims is connected by a relationship link to one or more other claims, entities, events, concepts or beliefs.
[0095] According to one embodiment of the present disclosure, a method for training a claims extraction language model, wherein the claims extraction language model includes fully qualified claims includes: providing a snippet portion of content from a plurality of content sources, chunking the snippet portion into one or more sentences, extracting one or more claims from the one or more sentences, decontextualizing the extracted claims into fully qualified claims using a claims taxonomy, and storing, in a database, the fully qualified claims as nodes in a claims graph.
[0096] The decontextualizing step may further include the step of using previously fully qualified claims as training data to fine tune the extracted claims into fully qualified claims.
[0097] The plurality of content sources may include text, films, music, games, images or video.
[0098] The step of decontextualizing fully qualified using a claims taxonomy may further include the step of identifying one or more claim types.
[0099] The one or more claims types may include a claim of fact, claim of value, claim of policy, claim of definition, or claim of cause or effect.
[0100] Each fully qualified claim may be a statement made by or about one or more entities, events, concepts or beliefs.
[0101] Each of the claims may be connected by at least one relationship link to one or more other claims, entities, events, concepts or beliefs.
[0102] The method may further include the step of storing the claims graph further including the step of storing the fully qualified claims linked together by at least one relationship link.
[0103] The method may further include converting each fully qualified claim into an embedding for semantically comparing the fully qualified claim with other claims.
[0104] The method may further include the step of using the claims extraction language model in the step of decontextualizing the extracted claims into fully qualified claims.
[0105] The content sources may be multimodal including text, images, video or audio.
[0106] According to one embodiment of the present disclosure, a method for determining a plurality content sources’ relative contribution to a large language model generated response includes: prompting the large language model to generate the response, using a claims extraction language model to extract response claims from the output response, using the claims extraction language model to associate semantically similar fully qualified source claims for each response claim of the response, associating the semantically similar source claims from the content sources with the associated response claims, and determining the importance of each of the source claims to the associated response claims.
[0107] The method may further include the step of ranking the content sources that are more important with relative contribution scores.
[0108] The content source providing semantically similar source claims for the most response claims may be ranked with a higher contribution score.
[0109] The content source providing one or more source claims for the most important response claim may be ranked with a higher contribution score.
[0110] The content source providing the most unique source claim may be ranked with a higher contribution score.
[0111] The content source providing the earliest publication date for the source claims may be ranked with a higher contribution score.
[0112] The content source having brand reputation for providing source claims with journalist integrity may be ranked with a higher contribution score.
[0113] The content source with stronger brand pedigree may be ranked with a higher contribution score.
[0114] The method may further include the step of sorting the content sources from highest to lowest contribution score ranking.
[0115] The method may further include assigning a credit score for each content source based on the highest to the lowest contribution score ranking.
[0116] The contribution score may be a percentage.
[0117] The content sources having higher contribution score rankings may be assigned higher fractional compensation.
[0118] The method may further include the step of assigning highest fractional compensation to the content source having the overall highest contribution for the output response.
[0119] According to one embodiment of the present disclosure, a method for determining a content source’s relative contribution to an output response generated by a large language model includes: prompting the large language model to generate the output response, extracting fully qualified response claims from the response, using a claims extraction language model to associate semantically similar fully qualified content source claims to the response claims, ranking the importance of each of the content source claims to each of the response claims, and assigning a relative contribution score for each of the content source claims to the response based on the ranking of each of the content source claims.
[0120] The method may further include the step of training the claims extraction language model with fully qualified source claims from a plurality of content sources and storing extracted claims in a claims graph database.
[0121] The method may further include the step of training the claims extraction language model with fully qualified source claims from a plurality of content sources including text, films, music, games, images or video.
[0122] The method may further include the step of representing the relative contribution score as a percentage.
[0123] The method may further include the step of representing higher percentages to relative contribution scores with the highest ranking source claims.
[0124] The method may further include the step of compensating publishers of the content sources in accordance with the percentages.
[0125] The method may further include the steps of determining the percentages and the compensation to publishers in real time for each of the generated responses.
[0126] The method may further include the step of aggregating the compensation due to a publisher over a period of time.
[0127] The step of aggregating the compensation due to a publisher may be over a period of time of a day, or a week, or a month or a quarter.
[0128] The content source claims providing semantically similar claims to the most response claims may be ranked with a higher contribution scores.
[0129] The content source claims providing one or more claims for the most important response claims may be ranked with a higher contribution scores.
[0130] The content source claims providing the most unique source claims may be ranked with higher contribution scores.
[0131] The content source claims providing the earliest publication date for the claims may be ranked with higher contribution scores.
[0132] The content source claims having brand reputation for providing claims with journalist integrity may be ranked with higher contribution scores.
[0133] The content source claims with stronger brand pedigree may be ranked with higher contribution scores.
[0134] The method may further include the step of sorting the content sources from highest to lowest rankings.
[0135] The method may further include assigning a contribution score for each content source claims based on the highest to the lowest rankings.
[0136] The publishers providing the content source claims having higher rankings or higher contribution scores may be assigned higher fractional compensation.
[0137] The method may further include the step of assigning highest fractional compensation to the publisher providing the content source claim having the highest contribution total score.
[0138] According to one embodiment of the present disclosure, a method for assessing the confidence of an output response produced by a large language model includes: training a claims extraction language model to extract fully qualified source claims from a plurality of content sources, prompting the large language model to obtain the output response, using the claims extraction language model to extract response claims from the response, using the claims extraction language model to associate semantically similar fully qualified source claims for each response claim, determining the relative importance of each of the response claims in the response, and for each of the content sources supporting the most important response claims, ranking the brand reputation of each content source, wherein the higher the ranking for the content sources, the higher the confidence in the response.
[0139] FIG. 8 is a block diagram illustrating components of a machine 800, according to some example embodiments, able to read instructions from a non-transitory machine-readable medium (e.g., a computer-readable storage medium) and perform any one or more of the methodologies discussed herein. Specifically, FIG. 8 shows a diagrammatic representation of the machine 800 in the example form of a computer system, within which instructions 810 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 800to perform any one or more of the methodologies discussed herein may be executed. As such, the instructions 810 may be used to implement modules or components described herein. The instructions 810 transform the general, non-programmed machine 800 into a particular machine 800 programmed to carry out the described and illustrated functions in the manner described. In alternative embodiments, the machine 800 operates as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 800 may operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 800 may include, but not be limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular telephone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing the instructions 810, sequentially or in parallel or concurrently, that specify actions to be taken by the machine 800. Further, while only a single machine 800 is illustrated, the term “machine” or “processing circuit” shall also be taken to include a collection of machines that individually or jointly execute the instructions 810 to perform any one or more of the methodologies discussed herein.
[0140] The machine 800 may include processors 804 (including processors 808 and 812), memory / storage 806, and I / O components 818, which may be configured to communicate with each other such as via a bus 802. The memory / storage 806 may include a memory 814, such as a main memory, or other memory storage, and a storage unit 816, both accessible to the processors 804 such as via the bus 802. The storage unit 816 and memory 814 store the instructions 810 embodying any one or more of the methodologies or functions described herein. The instructions 810 may also reside, completely or partially, within the memory 814, within the storage unit 816, within at least one of the processors 804 (e.g., within the processor’s cache memory), or any suitable combination thereof, during execution thereof by the machine 800. Accordingly, the memory 814, the storage unit 816, and the memory of the processors 804 are examples of machine-readable media.
[0141] The I / O components 818 may include a wide variety of components to receive input, provide output, produce output, transmit information, exchange information, capture measurements, and so on. The specific I / O components 818 that are included in a particular machine will depend on the type of machine. For example, portable machines such as mobile phones may include a touch input device or other such input mechanisms, while a headless server machine will likely not include sucha touch input device. It will be appreciated that the I / O components 818 may include many other components that are not shown in FIG. 8. The I / O components 818 are grouped according to functionality merely for simplifying the following discussion, and the grouping is in no way limiting. In various example embodiments, the I / O components 818 may include output components 826 and input components 828. The output components 826 may include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibratory motor, resistance mechanisms), other signal generators, and so forth. The input components 828 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), tactile input components (e.g., a physical button, a touch screen that provides location and / or force of touches or touch gestures, or other tactile input components), audio input components (e.g., a microphone), and the like.
[0142] In further example embodiments, the I / O components 818 may include biometric components 830, motion components 834, environment components 836, or position components 838, among a wide array of other components. For example, the biometric components 830 may include components to detect expressions (e.g., hand expressions, facial expressions, vocal expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, perspiration, or brain waves), identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or electroencephalogrambased identification), and the like. The motion components 834 may include acceleration sensor components (e.g., accelerometer), gravitation sensor components, rotation sensor components (e.g., gyroscope), and so forth. The environment components 836 may include, for example, illumination sensor components (e.g., photometer), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., barometer), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors (e.g., gas sensors to detect concentrations of hazardous gases for safety or to measure pollutants in the atmosphere), or other components that may provide indications, measurements, or signals corresponding to a surrounding physical environment. The position components 438 may include location sensor components (e.g., a GlobalPositioning System (GPS) receiver component), altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude may be derived), orientation sensor components (e.g., magnetometers), and the like.
[0143] Communication may be implemented using a wide variety of technologies. The I / O components 818 may include communication components 840 operable to couple the machine 800 to a network 832 or devices 820 via a coupling 824 and a coupling 822, respectively. For example, the communication components 840 may include a network interface component or other suitable device to interface with the network 832. In further examples, the communication components 840 may include wired communication components, wireless communication components, cellular communication components, Near Field Communication (NFC) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components to provide communication via other modalities. The devices 820 may be another machine or any of a wide variety of peripheral devices (e.g., a peripheral device coupled via a USB).
[0144] Moreover, the communication components 840 may detect identifiers or include components operable to detect identifiers. For example, the communication components 840 may include Radio Frequency Identification (RFID) tag reader components, NFC smart tag detection components, optical reader components (e.g., an optical sensor to detect one-dimensional bar codes such as Universal Product Code (UPC) bar code, multi-dimensional bar codes such as Quick Response (QR) code, Aztec code, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar code, and other optical codes), or acoustic detection components (e.g., microphones to identify tagged audio signals). In addition, a variety of information may be derived via the communication components 840, such as location via Internet Protocol (IP) geo-location, location via Wi-Fi® signal triangulation, location via detecting an NFC beacon signal that may indicate a particular location, and so forth. T
[0145] The term non-transitory computer-readable medium is to be understood herein to refer to one or more non-transitory computer-readable media, such as a single solid-state drive, multiple solid-state drives connected in a redundant array of independent drives, one or more hard disk drives (e.g., magnetic data storage media), one or more optical (e.g., CD-ROM or DVD-ROM) media, one or more pools of data storage devices connected to one or more computer servers, and the like.
[0146] It should be understood that the sequence of steps of the processes described herein in regard to various methods and with respect various flowcharts is not fixed, but can be modified, changed in order, performed differently, performed sequentially, concurrently, or simultaneously, or altered into any desired order consistent with dependencies between steps of the processes, as recognized by aperson of skill in the art. Further, as used herein and in the claims, the phrase “at least one of element A, element B, or element C” is intended to convey any of: element A, element B, element C, elements A and B, elements A and C, elements B and C, and elements A, B, and C.
[0147] A person of ordinary skill in the art would appreciate, in view of the present disclosure in its entirety, that each suitable feature of the various embodiments of the present disclosure may be combined or combined with each other, partially or entirely, and may be technically interlocked and operated in various suitable ways, and each embodiment may be implemented independently of each other or in conjunction with each other in any suitable manner.
[0148] While the present invention has been described in connection with certain exemplary embodiments, it is to be understood that the invention is not limited to the disclosed embodiments, but, on the contrary, is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims, and equivalents thereof.
Claims
WHAT IS CLAIMED IS:
1. A system comprising:a neural network trained to extract claims from an output response generated by a large language model,a graph structure stored in a database, the graph structure comprising a plurality claims, entities, events, concepts, or beliefs and relationship links between the plurality of claims, the claims being extracted from a plurality of source data, and a contribution engine configured to determine contribution scores of the source data used to train the large language model to generate the output response.
2. The system in claim 1 wherein the graph structure is represented in a machine-readable format.
3. The system in claim 1 wherein the relationship link is representative of a subsumption or an entailment.
4. The system in claim 1 wherein each of the claims is classified as a claims type.
5. The system in claim 4 where in the claims type is a claim of fact, claim of value, claim of policy claim of definition, or claim of cause or effect.
6. The system of claim 1 wherein a claim is a statement made by or about one or more of the entities, events, concepts or beliefs and wherein each of the claims is connected by a relationship link to one or more other claims, entities, events, concepts or beliefs.
7. A method for training a claims extraction language model, wherein the claims extraction language model comprises fully qualified claims, the method comprising:providing a snippet portion of content from a plurality of content sources, chunking the snippet portion into one or more sentences,extracting one or more claims from the one or more sentences, decontextualizing the extracted claims into fully qualified claims using a claims taxonomy, andstoring, in a database, the fully qualified claims as nodes in a claims graph.
8. The method of claim 7 wherein the decontextualizing step further comprises the step of using previously fully qualified claims as training data to fine tune the extracted claims into fully qualified claims.
9. The method of claim 7 wherein the plurality of content sources comprises text, films, music, games, images or video.
10. The method of claim 7 wherein the step of decontextualizing fully qualified using a claims taxonomy further comprises the step of identifying one or more claim types.
11. The method of claim 10 wherein the one or more claims types comprise a claim of fact, claim of value, claim of policy, claim of definition, or claim of cause or effect.
12. The method of claim 7 wherein each fully qualified claim is a statement made by or about one or more entities, events, concepts or beliefs.
13. The method of claim 12 wherein each of the claims is connected by at least one relationship link to one or more other claims, entities, events, concepts or beliefs.
14. The method of claim 12 further comprises the step of storing the claims graph further comprising the step of storing the fully qualified claims linked together by at least one relationship link.
15. The method of claim 12 further comprises converting each fully qualified claim into an embedding for semantically comparing the fully qualified claim with other claims.
16. The method of claim 7 further comprises the step of using the claims extraction language model in the step of decontextualizing the extracted claims into fully qualified claims.
17. The method of claim 7 wherein the content sources are multimodal comprising text, images, video or audio.
18. A method for determining a plurality content sources’ relative contribution to a large language model generated response comprising:prompting the large language model to generate the response,using a claims extraction language model to extract response claims from the output response,using the claims extraction language model to associate semantically similar fully qualified source claims for each response claim of the response, associating the semantically similar source claims from the content sources with the associated response claims, anddetermining the importance of each of the source claims to the associated response claims.
19. The method of claim 18 further comprises the step of ranking the content sources that are more important with relative contribution scores.
20. The method of claim 19 wherein the content source providing semantically similar source claims for the most response claims is ranked with a higher contribution score.
21. The method of claim 18 wherein the content source providing one or more source claims for the most important response claim is ranked with a higher contribution score.
22. The method of claim 18 wherein the content source providing the most unique source claim is ranked with a higher contribution score.
23. The method of claim 18 wherein the content source providing the earliest publication date for the source claims is ranked with a higher contribution score.
24. The method of claim 18 wherein the content source having brand reputation for providing source claims with journalist integrity is ranked with a higher contribution score.
25. The method of claim 18 wherein the content source with stronger brand pedigree is ranked with a higher contribution score.
26. The method of claim 18 further comprises the step of sorting the content sources from highest to lowest contribution score ranking.
27. The method of claim 18 further comprises assigning a credit score for each content source based on the highest to the lowest contribution score ranking.
28. The method of claim 27 wherein the contribution score is a percentage.
29. The method of claim 28 wherein the content sources having higher contribution score rankings are assigned higher fractional compensation.
30. The method of claim 29 further comprises the step of assigning highest fractional compensation to the content source having the overall highest contribution for the output response.
31. A method for determining a content source’s relative contribution to an output response generated by a large language model comprising:prompting the large language model to generate the output response, extracting fully qualified response claims from the response,using a claims extraction language model to associate semantically similar fully qualified content source claims to the response claims,ranking the importance of each of the content source claims to each of the response claims, andassigning a relative contribution score for each of the content source claims to the response based on the ranking of each of the content source claims.
32. The method of claim 31 further comprises the step of training the claims extraction language model with fully qualified source claims from a plurality of content sources and storing extracted claims in a claims graph database.
33. The method of claim 32 further comprising the step of training the claims extraction language model with fully qualified source claims from a plurality of content sources comprising text, films, music, games, images or video.
34. The method of claim 31 further comprises the step of representing the relative contribution score as a percentage.
35. The method of claim 34 further comprises the step of representing higher percentages to relative contribution scores with the highest ranking source claims.
36. The method of claim 35 further comprises the step of compensating publishers of the content sources in accordance with the percentages.
37. The method of claim 36 further comprises the steps of determining the percentages and the compensation to publishers in real time for each of the generated responses.
38. The method of claim 37 further comprises the step of aggregating the compensation due to a publisher over a period of time.
39. The method of claim 38 wherein the step of aggregating the compensation due to a publisher is over a period of time of a day, or a week, or a month or a quarter.
40. The method of claim 31 wherein the content source claims providing semantically similar claims to the most response claims are ranked with higher contribution scores.
41. The method of claim 31 wherein the content source claims providing one or more claims for the most important response claims are ranked with higher contribution scores.
42. The method of claim 31 wherein the content source claims providing the most unique source claims are ranked with higher contribution scores.
43. The method of claim 31 wherein the content source claims providing the earliest publication date for the claims is ranked with higher contribution scores.
44. The method of claim 31 wherein the content source claims having brand reputation for providing claims with journalist integrity are ranked with higher contribution scores.
45. The method of claim 31 wherein the content source claims with stronger brand pedigree are ranked with higher contribution scores.
46. The method of claim 31 further comprises the step of sorting the content sources from highest to lowest rankings.
47. The method of claim 31 further comprises assigning a contribution score for each content source claims based on the highest to the lowest rankings.
48. The method of claim 47 wherein the publishers providing the content source claims having higher rankings or higher contribution scores are assigned higher fractional compensation.
49. The method of claim 48 further comprises the step of assigning highest fractional compensation to the publisher providing the content source claim having the highest contribution total score.
50. A method for assessing the confidence of an output response produced by a large language model comprising:training a claims extraction language model to extract fully qualified source claims from a plurality of content sources,prompting the large language model to obtain the output response, using the claims extraction language model to extract response claims from the response,using the claims extraction language model to associate semantically similar fully qualified source claims for each response claim,determining the relative importance of each of the response claims in the response, andfor each of the content sources supporting the most important response claims, ranking the brand reputation of each content source, wherein the higher the ranking for the content sources, the higher the confidence in the response.