Systems and methods using predictive question and answer generation
The system addresses the limitations of LLMs and RAG systems by generating pre-defined question-answer pairs from trusted sources, ensuring fast, accurate, and transparent responses in legal and medical fields.
Patent Information
- Application Number
- PCT/US2025/037864
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-17
- Filing Date
- 2025-07-16
- Publication Date
- 2026-01-22
AI Technical Summary
Current Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems in legal and medical fields lack precision, reliability, and integration of up-to-date domain-specific knowledge, leading to inaccurate, verbose, and opaque responses that are slow and difficult to verify.
A system and method that leverages pre-generated question-answer pairs from trusted sources to provide accurate, concise, and transparent answers, using computational preprocessing to generate millions of potential question-answer pairs from a corpus of information, enabling rapid and reliable responses.
The system provides fast, accurate, and interpretable answers grounded in trusted sources, reducing response latency, minimizing hallucinations, and enhancing reliability through automated verification processes.
Smart Images

Figure US2025037864_22012026_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS USING PREDICTIVE QUESTION AND ANSWERGENERATIONCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This Application claims the benefit of U.S. Provisional Application No. 63 / 672,704, filed on 17-JUL-2024, titled “SYSTEMS AND METHODS USING PREDICTIVE QUESTION AND ANSWER GENERATION” which is incorporated in its entirety by this reference.TECHNICAL FIELD
[0002] This invention relates generally to the field of digital document search tools, and more specifically to a new and useful system and method for using predictive question and answer generation.BACKGROUND OF THE INVENTION
[0003] Large Language Models (LLMs), Retrieval -Augmented Generation (RAG), and vector databases represent significant advancements in natural language processing and information retrieval. However, their application in critical industries like legal and medical fields is limited by several factors. LLMs, while capable of generating human-like text, often lack the precision and reliability required for legal and medical documentation. RAG systems, which combine LLMs with retrieval mechanisms to improve factual accuracy, still struggle with accuracy, context relevance, and integration of up-to-date, domain-specific knowledge. Vector databases, essential for managing the vast unstructured data required for these technologies, face challenges in ensuring the accuracy and relevance of retrieved information.
[0004] Current LLM-enabled search and research tools are plagued by numerous computational and technical problems. They may be inaccurate and therefore unreliable, especially for applications requiring high precision. For example, legal or academic research depends on having highly accurate information. Additionally, they maybe slow, requiring minutes-long waits for responses due to real-time processing overhead. LLM-enabled responsescan also be overly verbose, sometimes containing multiple-page answers that can be burdensome and make it harder to obtain direct and quick answers. Furthermore, responses from such LLM-enabled research tools are opaque for how their responses are generated. They often will lack traceability to source materials. Verifying information generated by such research tools can be difficult and time consuming, often removing time advantages of such tools. These make such LLM-enabled search and research tools limited in their use in many industries especially those where efficiency and accuracy are critical.
[0005] Thus, there is a need in the digital document search field to create a new and useful system and method for using predictive question and answer generation. This invention provides such a new and useful system and method.BRIEF DESCRIPTION OF DRAWINGS
[0006] FIGURE 1 is a schematic of a system variation.
[0007] FIGURE 2 is a schematic representation of a system variation for content validation.
[0008] FIGURE 3 is a schematic representation of a system variation for validation and annotation of a document.
[0009] FIGURE 4 is a flowchart of a method variation.
[0010] FIGURE 5 is a detailed flowchart of a variation for processing a corpus of source information.
[0011] FIGURE 6 is a diagram representation of a method variation.
[0012] FIGURE 7 is a diagram representation of different potential considerations in some method variations.
[0013] FIGURE 8 is a flowchart representation of a method variation for validating content in a query.
[0014] FIGURE 9 is a detailed flowchart representation of a method variation for content validation.
[0015] FIGURES 10A and 10B are diagrams of an example processing of a segment.
[0016] FIGURE 11 is an exemplary system architecture that may be used in implementing the system and / or method.DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] The following description of the embodiments of the invention is not intended to limit the invention to these embodiments but rather to enable a person skilled in the art to make and use this invention.1. Overview
[0018] The systems and methods for LLM-enabled query responses can leverage large language models (LLMs) for processing one or more reference sources of information to generate accurate, concise, and interpretable (i.e., transparent in their source) research answers for given queries. The reference sources may be trusted sources used as a ground truth, assumed trusted sources, sources to which comparison maybe desired, and / or any suitable information or collection of information of interest. The systems and methods may use computational preprocessing of a corpus of sources of information to algorithmically pre-generate potential question-answer pairs that may be addressed by the information. In this way when a query is made, these pre-generated question-answer pairs may be leveraged to provide fast and highly accurate responses that are grounded in the source documents. The pre-generated question pairs from a corpus of reference source information can be used for improved answering of questions, citing appropriate sources, and / or comparing alignment of submitted information to the referenced source information.
[0019] In particular, the systems and methods may be used for providing high quality answers that are grounded in trusted sources of information for received questions or prompts. This pre-generation of question-answer pairs may generate millions, billions, or even trillions of potential question-answer pairs so as to computationally predict the possible answers for many of the potential questions that can be addressed by the source documents.
[0020] Additionally, or alternatively, the systems and methods may be used as a computational mechanism for content validation and verification. In this variation, a set of reference source materials maybe processed to generate reference question-answer pairs. The answers may be trusted to correctly answer the questions based on the source material or to at least reflect the answers to questions based on such reference source materials. Then another material or collection of materials may be processed to generate comparison question-answer pairs. The questions generated from the reference source materials may then be used as an index for matching questions generated from the other materials within a data system. Then a computational comparison of corresponding answers maybe used to generate an automatedassessment of corresponding statements. This validation approach maybe applied to evaluate whole documents, research papers, or articles, or may be used for validating individual statements or assertions within content.
[0021] Herein, the reference source material may be characterized as trusted source information. In the primary variation, the reference source materials serve as a source of truth or present an assumed ground truth for validation purposes. As trusted source information, the reference source materials provide information with high confidence of accuracy, enabling reliable validation and fact-checking operations. However, depending on the application and implementation, the reference source materials may alternatively represent any collection of source information to which comparison is desired, regardless of absolute truth value.
[0022] The systems and methods are not limited to traditionally "trusted" sources and maybe applied with various reference material configurations. In one example, when using the systems and methods for validation, different news articles from multiple publications may be included as reference source materials. Content could then be validated against this corpus of reference news sources to determine alignment with different publications or editorial perspectives. This approach may be used to show which publications or viewpoints the validated content aligns with most closely. As a further extension, content across the internet may be used as a comprehensive corpus of reference source information for comparative analysis.
[0023] While the systems and methods are primarily described herein with respect to trusted source materials for clarity of example, the systems and methods are not limited to trusted sources and may be applied to any suitable collection of reference source information, whether characterized as trusted, authoritative, representative, or comparative baseline material.
[0024] The systems and methods may be used as one stand-alone search / research / verification / validation tool. However, the systems and methods maybe used in combination with other search / research / verification / validation tools. For example, the systems and methods may be used in combination with other LLM or RAG-based document search solutions. The systems and methods may leverage LLMs in a manner that significantly mitigates the risk of hallucination. LLM and / or other algorithmic processing may be used in isolated, specific processes which are orchestrated to reduce opportunities for hallucination by constraining the scope of language model operations. The systems and methods may also include computational redundancies and automated verification processes to further enhance reliability and accuracy.
[0025] The systems and methods preferably use LLMs to predictively generate questionanswer pairs from a corpus of source information. In some variations, these question-answer pairs maybe used to rapidly retrieve answers to a general query that are backed by sourced resources. When a query is made to the computing platform, the query may be used to find potential question-answer pairs that correspond to the query (e.g., vector embedding search, keyword search, grammar-tree based search, etc.). The potential question-answer pairs may then be algorithmically evaluated to determine a confirmed answer and generate a response.
[0026] The systems and methods may be implemented across various digital platforms and interfaces. In one exemplary implementation, the systems and methods may be deployed in a query-based digital interface, such as a web application or native application, that can receive questions as queries and provide answers along with citations to relevant source information that supports the answers. In another exemplary implementation, the systems and methods may be integrated as a feature or service within existing Al chat interfaces or search products. For example, the systems and methods may be used to augment LLM-based chatbots by providing verified answers and citations that can be incorporated into chat conversation responses.
[0027] In another exemplary implementation, the systems and methods may be used to provide Al-facilitated search results within a search interface. For example, an Al-generated answer along with citations or links to search results (which were processed as source information) may be provided in a search result page for a search engine.
[0028] The systems and methods may also be implemented for information verification applications. This may include scoring or annotating alignment of individual statements with reference source information and generating overall validity scores for documents which may contain numerous individual statements. For example, the systems and methods maybe integrated with browser plug-ins or tools that can automatically check validity of one or more statements on websites against reference source information, providing real-time verification indicators to users.
[0029] The systems and methods may additionally be applied to various types of source information including textual documents, video and audio media (e.g., through transcript generation or multimodal model processing), structured database content, and computer code repositories. While the systems and methods are primarily described herein with respect to textbased information processing as the primary example, the systems and methods are not limited to text-based sources and maybe applied to any suitable type of source information. This versatility allows the systems and methods to be deployed across diverse industries and applications where reliable, source-grounded information retrieval and validation are critical.
[0030] The system and method may provide a number of potential benefits. The system and method are not limited to always providing such benefits and are presented only as exemplary representations for how the system and method may be put to use. The list of benefits is not intended to be exhaustive, and other benefits may additionally or alternatively exist.
[0031] As one potential benefit, the systems and methods may be a more accurate digital search and research tool. The systems and methods can be tightly coupled and restricted to responses based on trusted content. In particular, the operation of the systems and methods can be grounded in the assertions made in source information which may reduce likelihood of generating unsupported responses.
[0032] As another potential benefit, the systems and methods may be considerably faster than traditional LLM and RAG based approaches. Responses may be generated in a few seconds or less, compared to tens of seconds or even minutes using existing LLM approaches. This computational efficiency is achieved through pre-processing that converts intensive real-time language model operations into rapid database retrieval operations, significantly reducing response latency.
[0033] As another potential benefit, the systems and methods may provide more control over the form of responses. In particular, answers can be concise. In some cases, answers may be one to a few word answers. As the systems and methods pre-generates answers, they can be configured to keep the answers in a form expected for the application. This maybe used to optimize response formatting for specific user interface requirements and / or downstream processing systems. A related additional or alternative benefit of response control offered by the systems and methods is that it may also reduce harmful or disruptive discourse from an LLM (e.g., adding bias, inappropriateness, rambling, and the like).
[0034] As another potential benefit, the systems and methods may be more transparent in how the answers are derived. The answers will have direct correlation to portions of referenced information. A user could be directed to the exact text that supports a given answer and the broader context in which that answer fits. This transparency can also make query tools using the systems and methods more reliable. Users can quickly see sources supporting an answer, which may make use of such a query tool more feasible in industries where high accuracy is required.
[0035] As another potential benefit, the systems and methods may further provide enhanced reliability through computational redundancies and automated verification processes. The systems and methods can implement multiple verification layers, including pre-verification of question-answer pairs during generation and / or on-demand verification during queryprocessing. This multi-stage verification approach reduces the likelihood of errors and enhances system dependability for critical applications where accuracy is critical.
[0036] As another potential benefit, the systems and methods may provide automated content validation and verification capabilities. The question-answer generation mechanism serves as a computational indexing system for comparing statements across different sources. When used with trusted sources and untrusted sources, individual statements in untrusted sources may be algorithmically scored and evaluated using the generated question-answer pairs. This may enable systematic assessment of content reliability, identification of potential misinformation, and automated fact-checking workflows that can be integrated into content management systems or publishing platforms.
[0037] As another potential benefit, the systems and methods may provide superior scalability and responsiveness for processing large information corpora. The pre-processing approach enables efficient handling of extensive document collections by distributing computational load across time, with intensive processing performed offline during corpus preparation rather than during real-time query serving. This architecture allows the systems and methods to scale to handle millions, billions, or even trillions of question-answer pairs and support high query volumes without degrading response performance.
[0038] As another potential benefit, the systems and methods may provide domainspecific customization capabilities. The systems and methods can be adapted for various industries and applications by tailoring question generation patterns, answer formats, and verification criteria based on specific domain requirements. For example, legal applications may emphasize citation accuracy and regulatory compliance, while medical applications may prioritize precision and risk assessment, allowing the same underlying technology to serve diverse specialized use cases. The formats of answers may be customized for different industries. Similarly, the corpus of information maybe curated and customized for different applications or instances of use.
[0039] As another potential benefit, the systems and methods may be readily integrated with existing digital tools and platforms to augment and improve their capabilities. The systems and methods can be implemented as modular components that enhance search engines, chatbots, content management systems, web browsers, and research platforms without requiring complete system replacement. This integration capability allows organizations to leverage existing technology investments while adding advanced question-answering and content validation functionality through APIs, plugins, or embedded services.2. System
[0040] As shown in FIGURE 1, a system may include a question-answer pair generator 110, a question-answer pair data system 120, a query interface 130, and a question-answer pair matching engine 140. Depending on configuration and intended application of the system, the system may additionally include a verification engine 150. In some variations, the verification engine 150 may be configured to function as or otherwise include a content validation engine as shown in FIGURE 2 to facilitate validation of content or statements received as a query input.
[0041] In some variations, the system may be configured for validating multiple statements from a set of content (e.g., a whole document or collection of documents) as shown in FIGURE 3. The question-answer pair generator 110 or a similar system for content processing module maybe used to segment the content into generated question-answer pairs where the answers correspond to statements or assertions from the content. The generated questionanswer pairs may then be processed by a question-answer pair matching engine 140 to find similar question-answer pairs in the question-answer pair data system 120. The similar question-answer pairs may then be processed by the verification engine 150 (or other suitable system) to validate or otherwise assess how the answers from the found question-answer pairs align with the corresponding answer (or statement) from the content. This system variation may be used to batch validate information in a set of content.
[0042] The question-answer pair generator 110 functions to generate a set of questionanswer pairs from different segments of a corpus of source information. The source information could be a text, a collection of text, or any suitable source of information. The source information may be static but could alternatively be a dynamic changing source of information. The question-answer pair generator 110 may include and / or interface with a large language model or other natural language services to perform the computational processing of source information into question-answer pairs. The question-answer pair generator 110 may include specialized processing modules for different source types including text processing modules, media transcription modules, database analysis modules, and code repository analysis modules. The question-answer pair generator 110 may also include segmentation algorithms for dividing source information into processable segments, answer identification algorithms for extracting potential answers from segments, and question generation algorithms for creating questions corresponding to identified answers.
[0043] The question-answer pair generator 110 may include a source processing module that handles ingestion and initial processing of various source information types. The sourceprocessing module may include format conversion capabilities, metadata extraction functions, and preprocessing algorithms that prepare source information for question-answer pair generation. The module may interface with external systems to retrieve source information and maintain source attribution throughout the processing pipeline.
[0044] The question-answer pair generator 110 may include a segmentation engine that divides source information into processable segments of varying sizes. The segmentation engine may implement algorithms for creating overlapping segments, composing segments from multiple sources, and adding contextual information to segments. The engine may utilize natural language processing techniques to identify optimal segment boundaries and maintain semantic coherence within segments.
[0045] The question-answer pair data system 120 functions to store and host the question-answer pairs. The question-answer pairs can be stored with reference to the source information and associated metadata. The data system 120 may comprise various storage architectures including relational databases, key-value stores, document databases, or vector databases optimized for embedding-based retrieval. The question-answer pair data system 120 may implement multiple indexing strategies including B-tree indexes for metadata, full-text search indexes for content, and vector indexes for semantic similarity computation. The question-answer pair data system 120 may also include caching mechanisms, replication systems, and backup processes to ensure data reliability and retrieval performance.
[0046] The question-answer pair data system 120 may include a storage engine that manages the physical storage and organization of question-answer pairs. The storage engine may implement compression algorithms, data partitioning strategies, and optimization techniques for efficient storage utilization and retrieval performance.
[0047] The question-answer pair data system 120 may include an indexing system that creates and maintains various indexes for efficient searching and retrieval. The indexing system may generate vector embeddings for semantic search, maintain keyword indexes for exact matching, and create metadata indexes for filtering operations.
[0048] The query interface 130 functions to provide a mechanism through which a request may be made. The query interface 130 may include a user interface, which may be accessible through a web application, a native application, a chat interface, an audio interface, or any suitable user-facing computer interface. The query interface 130 may additionally or alternatively include a programmatic interface such as an application programming interface. In another variation, the programmatic interface could be an internal interface used within a computing system or application. The query interface 130 may include input validationmechanisms, query parsing algorithms, and formatting systems that prepare received queries for processing by downstream components.
[0049] The question-answer pair matching engine 140 functions to process and facilitate generating a response for a given query. The question-answer pair matching engine 140 may facilitate finding relevant question-answer pairs for a given query and then evaluating those candidate question-answer pairs to determine a response. In the variation for generating answers, the matching engine 140 may identify candidate question-answer pairs that are used to form an answer response. The matching engine 140 may include vector similarity computation algorithms, keyword matching systems, semantic analysis capabilities, and ranking algorithms that prioritize candidate question-answer pairs based on relevance scores. The engine 140 may also include filtering mechanisms that apply metadata constraints and similarity thresholds to refine candidate selection.
[0050] The question-answer pair matching engine 140 may include a search algorithm module that implements various search strategies including vector database search, embeddingbased similarity computation, keyword search, and hybrid search approaches. The module may utilize multiple similarity metrics and ranking algorithms to identify and prioritize candidate question-answer pairs.
[0051] The question-answer pair matching engine 140 may include a filtering and ranking system that applies query-specific filters and ranks candidate question-answer pairs based on relevance, source credibility, recency, and metadata criteria. The system may implement machine learning algorithms for continuous improvement of ranking performance.
[0052] The question-answer pair verification engine 150 functions to analyze and assess question-answer pairs for accuracy and relevance. As shown in FIGURE 1, verification of question-answer pairs may be performed as part of servicing a query. However, in some other variations, the verification engine 150 maybe used in connection with the question-answer pair generator 110. The verification engine 150 may include large language model interfaces for performing verification tasks, multi-choice verification algorithms, and caching systems for storing verification results. The engine 150 may implement both pre-verification processes during question-answer pair generation and on-demand verification during query processing.
[0053] The verification engine 150 can be used to perform various checks and assessments of the question-answer pairs. In one variation, the verification engine 150 may be configured to verify various quality conditions for question and answer pairs. For example, the verification engine 150 may be used to make sure an answer suitably addresses a question. This may be used to make sure a generated answer adequately addresses a question from a queryinput. This may also be used to make sure the question-answer pairs stored in the questionanswer pair data system 120 include an answer that addresses the question.
[0054] As another variation, the verification may be configured to process input received through the query interface and to evaluate how the information relates to answers stored in the question-answer pair data system 120. This validation may alternatively be performed by a separate subsystem (e.g., a content validation engine).
[0055] As in a variation for generating an answer response for a question, a query input may be received by the system, but in this variation, the system may receive queries containing individual statements or multiple statements requiring validation against reference source information. The content received may be similarly processed to generate one or more associated question-answer pairs from the query input and then match these to pregenerated question-answers in the question-answer pair data system 120. The verification engine 150 may then review the matched question-answer pairs to assess how the answers from the source question-answer pairs compare to the answers from the query input content. This may be used to generate a validation score for one or more statements based on comparison of statements in the content and the matched answers as shown in FIGURE 2.
[0056] The verification engine 150 may process single factual claims, collections of statements from documents, full documents, a collection of documents or resources, and / or continuous streams of content requiring real-time validation assessment. The input received for a validation query input may be a form submission or any suitable data input that may be used to indicate one or more pieces of content for validation. For example, a document maybe uploaded and submitted through a query interface for validation. In another example, access maybe granted to a document storage system (e.g., such as a cloud storage folder or document) so that one or more documents may be processed for validation.
[0057] When the input includes content with multiple statements for validation, the system may process the query input to segment the content into multiple statements for validation. This “segmentation” may be facilitated by the question-answer pair generator 110 used in populating the data system 120 where each statement may be used to create associated question-answer pairs. During segmentation, content may be separated into distinct statements or assertions and then question-answer pairs may be generated to correspond with the statement. As shown in FIGURE 3, the query input may be passed through question-answer pair generator 110 and then each statement may then be validated. Alternatively, a separate or different content processing module may be used for processing query input.
[0058] In a validation variation, the verification engine 150 may include a validation scoring system that generates numerical or categorical validity scores for statements and documents. The scoring system may implement weighted scoring algorithms, confidence measures, and threshold-based classification systems for categorizing content reliability based on alignment with trusted source answers. The validation scoring system may similarly use the verification techniques related to testing alignment of questions and answers to test how answers align.
[0059] Additionally, the verification engine may include an annotation engine that generates digital annotations, citations, footnotes, and visual indicators for validated content. For example, individual statements maybe marked or annotated with indication of validation scores for different statements. As shown in FIGURE 3, statements may have an approval indicator (e.g., a “star”), have a warning when there maybe issues (e.g., a warning triangle symbol), or an alert indicator when there is a conflict (e.g., an exclamation alert symbol). The annotation engine may interface with various content management systems and provide APIs for browser plugins, content editors, and publishing platforms to display validation results.3. Method
[0060] As shown in FIGURE 4, a method may include: processing a corpus of source information to generate a set of question-answer pairs S110; storing the question-answer pairs in a question-answer pair data system with reference to segments of the source information S120; receiving a query through a query interface S130; identifying candidate question-answer pairs by searching the stored question-answer pairs using the query S140; and generating a response to the query by processing the candidate question-answer pairs S150. The corpus of source information maybe processed using a large language model.
[0061] The method in one variation may be used for answer retrieval / generation for a received question query. In this variation, as shown in FIGURE 5, the method includes: processing a corpus of source information to generate a set of question-answer pairs S110, wherein processing the corpus comprises processing multiple information segments of the source information S112, identifying potential answers for a given information segment S114, and for each potential answer, generating a set of questions that can be answered with the potential answer based on the segment S116; storing the question-answer pairs in a questionanswer pair data system with reference to segments of the source information S120; receiving a query through a query interface S130; identifying candidate question-answer pairs by searchingthe stored question-answer pairs using the query S140; and generating a response to the query by processing the candidate question-answer pairs S150, wherein generating a response comprises providing an answer from at least one of the candidate question-answer pairs S152. The method may involve orchestration of one or more language models within the process, where an LLM may be used for identifying the potential answers from segments in S114, generating questions for potential answers S116, and / or other processes. Additionally or alternatively, the method may use grammar tree analysis and other natural language processing techniques for answer identification, question generation, and / or other processes.
[0062] As shown in FIGURE 6, the method preferably uses source information to pregenerate question-answer pairs, which can then be used at query-time to retrieve and generate accurate answers. As shown in FIGURE 7, various considerations maybe made at the different stages of the process.
[0063] As shown in FIGURE 8, an alternative variation may apply the method for validation of one or more statements or facts that were submitted as a query. This validation approach may be used for validating individual statements, facts, or multiple statements from one or more sources or documents by treating validation targets as query inputs. The query to be validated could be, for example, a single statement, LLM-generated text, a large document or collection of documents, or any amount of content requiring validation assessment. In such a validation variation, the method may include: processing a corpus of source information to generate a set of question-answer pairs S110; storing the question-answer pairs in a questionanswer pair data system with reference to segments of the source information S120; receiving a query with at least one statement for validation through a query interface S131; for each statement, generating one or more questions based on the statement S137; using the generated question to find corresponding pre-generated question-answer pairs S142; and generating a validation score for each statement based on computational comparison between the answer and the statement S152. In the case of receiving a query input with a large amount of content for validation, the method may additionally include processing the query to identify and extract statements for validation S135, in which case block S137 may include, for each identified statement, generating one or more questions based on the statement using language model processing (S137). S135 and S137 may use similar processing as used when processing the corpus of source information. Accordingly, the process for generating a set of question-answer pairs in S110 may be similarly used for generating question-answer pairs from statements in the query input. The generated question-answer pairs associated with the validation input may thenbe used for matching and comparison to found answers in order to assess validity of the statements.
[0064] More generally, the validation approach enables any statement or collection of statements, possibly derived from one or more documents or resources, to be validated by generating question-answer pairs from the statements and matching them against a reference set of pre-generated question-answer pairs from select sources. This approach allows systematic validation of content regardless of source format, length, or complexity by leveraging the same question-answer generation and matching processing used for answer-query processing but used to assess content accuracy and reliability against established reference materials.
[0065] A method for information validation, as shown in FIGURE 9, may alternatively be characterized as: processing source information using a large language model to generate reference question-answer pairs S210; storing the reference question-answer pairs in a database with reference to segments of the source information S220; processing untrusted content using the large language model to generate comparison question-answer pairs S230; identifying questions from the comparison pairs that match questions from the reference pairs S240; evaluating correspondence between the reference answers from the matching reference question-answer pairs and statements in the untrusted content S250; and generating a content validity score for the statements based on the evaluated correspondence S260. This validation method functions to compare statements from submitted content to answers extracted from a corpus of source information to determine content validity scores for the untrusted content through systematic computational analysis. In this exemplary variation, processes S210 and S230 may function similarly to S110. Untrusted content may be received through a process such as S130. S220, S240, S260 may be substantially similar to S120, S140 and S150 respectively.
[0066] Block S110, which includes processing a corpus of source information to generate a set of question-answer pairs, functions to predictively generate possible questions and their answers that may be answered by a source of information. Processing the corpus of source information may include using a large language model to generate the set of question-answer pairs, though additional or alternative processing techniques may be used.
[0067] The processing may generate millions, billions, or even trillions of potential question-answer pairs to computationally predict possible answers for all potential questions that can be addressed by the source information. In some implementations, processing the corpus may generate more than one trillion predicted question-answer pairs, enabling comprehensive coverage of the information space contained within the source materials. In some applications, the method may be used for processing any arbitrarily large source ofreference information. For example, the source material may include a large portion of resources on the internet. In such variations, the source material may be an ever-growing body of content. Accordingly, the set of pre-generated question-answer pairs may be scaled to capture the body of content used for search and / or validation processes.
[0068] Processing of a corpus may be substantially performed in advance of any received query. In some cases, the processing of the corpus maybe a one-time process. For example, a collection of legal documents may be processed to generate a static set of pre-generated question-answer pairs.
[0069] In other variations, processing of a corpus of source information may be a periodic or continuous process. In some cases, source information may change. Accordingly, the method may include updating question-answer pairs in response to changing content in previously processed source information. Additionally, new source information may be added such as when the method is used with an ever-evolving source of information.
[0070] As discussed herein, in some variations, the generated question-answer pairs may be verified after generation, which functions to ensure the questions and answers make sense. This may alternatively be deferred to when the question-answer pairs are used to answer a query.
[0071] The source information may comprise various types of data, media files, or digital resources that contain text-based information or from which text-based information maybe extracted. As one main variation, source information includes data that has text-based information or content from which text descriptions or information may be computationally extracted. This can include text-based documents but can also include any digital content from which textual information may be derived through automated processing. Additionally or alternatively, multimodal language models may be used to interpret media in non-text formats, enabling broader content analysis capabilities.
[0072] In one variation, the source information comprises textual data which can include any suitable type of text-based content. The textual data may be from a set of text documents including legal documents, research papers, regulatory materials, academic publications, technical manuals, news articles, encyclopedias, paper or article databases, or other text-based content. Text-based documents may also include for example, articles, webpage content, social media, or any suitable type of document or media containing text.
[0073] In another variation, the source information comprises audio-based media content. In one variation, audio content may be converted to text for processing in a manner similar to the text-based information. Accordingly, when source information includes audio-based media content, the method may further include generating transcripts from the media content, wherein processing the corpus using the large language model processes at least the transcripts of the media content. Audio-based media content may include video files, audio recordings, podcasts, lectures, interviews, interactive media, multimedia presentations, and / or any suitable media format with audio. Transcript generation may be performed prior to corpus processing using automated speech recognition systems.
[0074] In some variations, the method may also process visual-based media content, where video or image content analysis may include visual interpretation of video frames, image content analysis, on-screen text extraction, and multimodal content understanding to generate comprehensive question-answer pairs from both audio and visual information streams. Accordingly, the source information may include video-based media content, and processing the corpus of source information may include converting the video content into textual representations using language model processing prior to generating question-answer pairs. In some variations, converting the video content into textual representations may include at least one of: generating transcripts from audio portions of the video content, generating descriptions of visual content using multimodal language models, and extracting textual information displayed within the video content depending on the type of model used and application.
[0075] In another variation, the source information comprises database content, further comprising generating questions and answers from data contained within the database. This may include structured data such as numerical data, tabular data, relational database content, or data warehouse information. Potential answers may be derived from database schema analysis, computational data analysis including statistical calculations and comparative analytics, and generated database queries that form answers for potential questions. For example, questions about data relationships, statistical summaries, trend analysis, or specific data retrievals may be generated with corresponding answers computed from the underlying database content.
[0076] In another variation, the source information comprises computer code repository content including software projects, documentation, and related technical materials. Processing comprises generating question-answer pairs from function signatures, API documentation, code comments, program logic analysis, software architecture descriptions, and implementation details. Questions may address code functionality, parameter requirements, return values, usage examples, debugging information, and software design patterns, with answers derived from automated code analysis and documentation parsing.
[0077] Various techniques may be used to generate question-answer pairs. In one variation, LLMs or other NLP processes may first determine potential answers in the segment and then figure out questions that could be answered with that answer based on the segment. In one variation, as shown in FIGURE 5, processing a corpus of source information using a large language model to generate a set of question-answer pairs S110 may include: processing multiple information segments of the corpus of source information S112; identifying potential answers for a given information segment S114; and for each potential answer, generate a set of questions S116. The generated question-answer pairs may then be stored in a question-answer pair data system as part of block S120 discussed below. In some variations, the question-answer pairs may be verified or checked to satisfy different requirements.
[0078] Block S112, which includes processing multiple information segments of the corpus of source information, functions to process discrete portions of a source of information. The source information herein is generally characterized as textual information but could additionally include other forms of information such as numerical data, tabular data, visual information, structured data, multi-media, and the like.
[0079] Processing multiple information segments can include algorithmically segmenting the source information. Segmenting source information could include segmenting based on sentences, phrases, paragraphs, text windows, sections or any suitable way of segmenting the source information. The segments may be consistent, but the method may also include segmenting with varying segment sizes. The segments may be distinct windows but could additionally overlap wherein a subset of segments overlap with at least one other segment. For example, one segment could be a paragraph, and another segment could be a sentence from that paragraph. Having varying segment sizes may afford different answers of facts to be determined from the content from which a question-answer pair can be generated.
[0080] Segmenting, in some variations, may additionally include composing segments from the source information. Different segments (e.g., small segments in the source) may be combined, merged, or used in combination as one segment. For example, composing segments may include combining or merging related information data to form segments, such as placing related text snippets together or injecting definitions where defined terms are used.
[0081] Segmenting may additionally include adding supplemental contextual information to segments. For example, segments may have accompanying information like chapter titles, section headers, preceding / proceeding sentences or paragraphs, term definitions, and / or other organizational metadata or contextual information. This supplemental contextualinformation maybe provided as context when generating, evaluating and / or verifying questionanswer pairs.
[0082] Block S114, which includes identifying potential answers for a given information segment, functions to algorithmically generate answers that could be derived from the segment. This may include identifying nouns and / or verbs or parsing grammatical structures for factual assertions or phrases that can form answers or facts. Potential answers could also include positive or negative affirmations. In some variations, the potential answers maybe directly taken from content in the segment. In other variations, a potential answer may be some assessment or logical conclusion of the segment. For example, if a sentence segment states that events happen in the order of A, B, and C, then one potential answer could be "first", "second", and / or "last". In this example, corresponding generated questions could be "when does A occur in the order of events?", "when does B happen in the order of events?", and "when does C happen in the order of events?" Potential answers could also be negations of any of the other potential answers. Various other types of answers may be identified or more generally generated from the segment. Identifying potential answers from a segment may use a large language model, a grammar tree, or other NLP techniques. When processing a given segment, one, multiple, or no potential answers may be identified.
[0083] Block S116, which includes generating a set of questions for each potential answer, functions to algorithmically determine one or more questions that can be answered with the potential answer and supported by the segment. An LLM or other NLP process may be used. For an LLM variation, an LLM may be prompted for each potential answer to generate one or more questions that are answered by the potential answer and supported by the segment. There may be multiple different questions that are generated for each potential answer, wherein generating the set of questions comprises generating multiple different questions that can be answered by the same potential answer. In some cases, the LLM may determine that the potential answer does not have a corresponding question supported by the segment in which case no generated question-answer pair may be generated for that potential answer. This may help mitigate errors caused by hallucinations and other LLM errors.
[0084] Herein the generation of answers and questions are described as being generated with answers first and then questions using those answers and the segment. However, questionanswer pairs may alternatively or additionally be generated in other ways as well.
[0085] In one example of processing a corpus of source information to generate a set of question-answer pairs, one segment in a legal text may read: "A general conservatorship terminates only on the conservatee's death or by court order." This segment maybe processedusing a large language model to identify multiple potential answers including "A general conservatorship," "conservatee's death," "court order," and "terminates." For the potential answer "A general conservatorship," an LLM may be prompted to generate questions. For example, an LLM prompt may be formulated as: "Generate one question from the following sentence, where the answer to the question is provided by the following phrase. Sentence: 'A general conservatorship terminates only on the conservatee's death or by court order.' Phrase: 'A general conservatorship'". The LLM might generate questions such as "What type of legal arrangement terminates only on the conservatee's death or by court order?" or "What terminates only on the conservatee's death or by court order?" Each generated question can then be stored with the corresponding potential answer as a question-answer pair, with reference to the source segment.
[0086] FIGURE 10A and FIGURE 10B show examples of a segment and the generated question-answer pairs that may be generated for a set of different potential answers derived from the same segment, demonstrating how multiple question-answer pairs can be algorithmically extracted from a single source segment.
[0087] Block S120, which includes storing question-answer pairs in a question-answer pair data system, functions to make each determined question-answer pair accessible for later use. Each question-answer pair stored may store the generated question and corresponding answer. A reference or other suitable citation record maybe used to map the question-answer pair to the segment of the source information.
[0088] Different texts or sources may be processed and imported with corresponding question-answer pairs that maintain metadata associations. The question-answer pairs may be stored in a manner where searched pairs can be algorithmically filtered based on metadata of the source material such as author, publication date, citation count, peer review status, publication name, publication type, publication venue, document type, and / or other information metadata. This filtering capability enables domain-specific queiying and source credibility assessment when responding to a query. For example, source metadata maybe used to find answers from documents from a particular time period, from a particular source, or any suitable subset of resources.
[0089] The question-answer pair data system may utilize various storage architectures optimized for retrieval performance. In some variations, the data system may comprise a relational database with indexed tables storing question-answer pairs, source references, and metadata. Alternatively, the system may use key- value stores for rapid retrieval, document databases for flexible schema management, or specialized vector databases for embedding-based similarity search. The storage system may implement multiple indexing strategies including B-tree indexes on metadata fields, full-text search indexes on question and answer content, hash indexes for exact matching, and vector indexes for semantic similarity computation. In some implementations, hybrid storage approaches may be used, combining relational databases for structured metadata with vector databases for semantic search capabilities. The indexing processes may include pre-computing embeddings for questions and answers, creating inverted indexes for keyword search, and maintaining metadata indexes for filtering operations. This multi-layered storage and indexing architecture may enable efficient retrieval across different query types and similarity computation methods.
[0090] As discussed in block S110, processing of source information may be performed initially, and then kept static. While in some variations, where source information is added, removed, or edited, the question-answer data system can be algorithmically updated with repeated processing. As such, storing of potential question-answer pairs or updating potential question-answer pairs (e.g., editing or removing) maybe performed in response to updated processing of the source information. If some segment from the source information is removed, the corresponding question-answer pairs may be automatically removed. If some segment from the source information is edited, the corresponding question-answer pairs may be automatically updated.
[0091] In some variations, the method may involve computational processes to consolidate, condense, or address issues where question-answer pairs maybe the same or similar to other question-answer pairs for different segments. In some cases, different segments may yield similar question-answer pairs. The method may include determining at least two question-answer pairs satisfying a similarity condition using algorithmic comparison techniques and consolidating the similar question-answer pairs by storing a consolidated question-answer pair. In some variations, the consolidated question-answer pair maybe stored with multiple references to source information segments. This may enable comprehensive source attribution while reducing data redundancy. The consolidation process may also normalize or otherwise adjust questions and / or answers to consistent phrasing to improve matching and retrieval accuracy. Individual records may be maintained for each source segment, while the consolidated approach optimizes storage efficiency and query performance.
[0092] This consolidation may additionally include identifying and resolving conflicting question-answer pairs. Conflicting pairs may arise when similar questions from different segments have contradictory or inconsistent answers. The method may use computational analysis to detect such conflicts by comparing answer semantics, identifying logicalcontradictions, and / or flagging pairs where the same question yields substantially different answers from different source segments. When conflicts are detected, the method may trigger a review process that examines the source material of segments that generated the conflicting questions to reevaluate whether the question-answer pairs accurately reflect the source content. This review process may include algorithmically widening the segment boundaries to incorporate greater contextual information, reprocessing segments with enhanced contextual data, or adjusting segmentation parameters to improve accuracy. The review may also involve re-analyzing the segment with different language model parameters or alternative questionanswer generation approaches. Conflict resolution may involve several approaches: maintaining separate question-answer pairs with source attribution to preserve conflicting viewpoints, prioritizing answers based on source credibility metrics, generating composite answers that acknowledge the conflict, or flagging conflicted pairs for manual review. The resolution approach may be configurable based on the application domain and reliability requirements of the system.
[0093] Block S130, which includes receiving a query through a query interface, functions to receive some input that triggers an analysis process leveraging the potential question-answer pairs. The query maybe received through various possible technical interfaces including user interfaces such as web applications, mobile applications, chat interfaces, voice interfaces, or through programmatic interfaces such as application programming interfaces (APIs), webhooks, or embedded service calls. The query may be subsequently used to determine similar pregenerated question-answer pairs and then use those to generate a response in blocks S140 and S150.
[0094] In an answer-generation variation, the query comprises a submitted question or other prompt with an objective of receiving an answer to the query. The query will generally be a question seeking specific information, but may also include complex prompts, natural language requests, or structured queries.
[0095] In one exemplary variation, a query prompt may be an explicit question which may be used to find a corresponding question in the pre-generated question-answer pairs.
[0096] In another exemplary variation, a query prompt may be a message from which one or more implied questions maybe generated using language model analysis. For example, the query prompt may be a message from a user detailing some request or objective. There may be one or more implied questions that relate to this prompt. In some cases, the question associated with a query may not even be from the received query but could be from some related content such as an Al chatbot-generated response to the query. The method may algorithmicallyparse complex prompts to identify constituent questions or information requests embedded within the prompt or from content derived from the query prompt. In this way, the method may be used to provide answers to specific questions even when those questions are not explicitly expressed in the prompt, but answering the questions would work towards fulfilling the overall prompt objective.
[0097] In another variation where the method is used for validating statements, the query could comprise a statement or collection of statements to be validated through the method. For example, a query could be long-form LLM-generated text or a document comprising many statements. In this validation variation, the query includes at least one assertion, claim, or factual statement that requires validation against source information. The method may process such validation queries by generating questions based on the statements and using those questions to find corresponding pre-generated question-answer pairs for comparison and validation scoring.
[0098] The query interface may also enable selection and customization of the source information used for providing answers. For example, a query may be configured to consider source information across a comprehensive collection of legal texts or may be restricted to search within a few select legal texts based on user preferences, domain requirements, or credibility filters. This source selection capability allows for targeted querying and domainspecific result filtering.
[0099] The query may be received through various exemplary implementations depending on the implementation context. In one exemplary implementation, the query may be received through a dedicated question-answer search interface, such as a web-based or native application specifically designed for querying the pre-generated question-answer pairs. In another exemplary implementation, the query may be received through a search engine interface where the method augments traditional search results with Al-generated answers derived from the question-answer pairs. The query in some variations may include an uploaded or referenced document or digital resource(s). For example, the query may include an uploaded document, a link to a website (e.g., a URL), a reference to a data storage location (e.g., a cloud storage link to a document or collection of resources for processing), or any suitable indicator of content for processing. The query may also be received programmatically through API calls from external systems or applications seeking to integrate the question-answering capabilities. Additionally, the query maybe received in response to a message in an Al chatbot interface, where the method provides verified answers to enhance chatbot responses with source- grounded information. Other exemplary implementations may include voice-activated systems,browser plugins for content validation, or embedded widgets within content management systems.
[0100] As described above, the source information may be in a variety of media formats including, but not limited to, text-based resources, video content, audio-based content, and other forms of multi-media, as well as other resource formats like database content or code repository content. The query, in a similar way, maybe specified in similar alternative forms. For example, audio-based media content or video content could similarly be supplied as inputs.
[0101] Accordingly, in some variations, the query may include query input in at least one format selected from text format, audio format, and video format. In some variations, the method may further include audio and / or video formats processed to extract textual information prior to query processing. However, multi-modal models may alternatively be used for direct interpretation of the audio and / or video media content.
[0102] Block S140, which includes identifying candidate question-answer pairs, functions to find corresponding question-answer pairs from the stored question-answer pair data system. In particular, the question-answer pairs may be searched using the stored questions as the mechanism for matching to a query. Search may use vector embedding search (e.g., using cosine similarity between the query and questions), keyword search, grammar-based search, semantic similarity algorithms, and / or other search processes that can match a query to a set of corresponding questions of the question-answer pairs. In one variation, the method may include performing vector database search using embedding -based similarity computation between the query and stored questions of the question-answer pairs in the question-answer pair data system. In general, a ranked set of question-answer pairs may be identified based on relevance scores computed through the various search mechanisms.
[0103] In some variations, the method may include computationally evaluating candidate questions for similarity to the query. This can include generating similarity scores or other types of computational characterization of similarity between the query and candidate questions. The similarity scoring may utilize various metrics including cosine similarity of embedding vectors, semantic distance measures, keyword overlap scores, and weighted combinations of multiple similarity factors. For a question-answer pair to be included in the candidate set of question-answer pairs, the question maybe required to satisfy a minimum similarity score threshold to be used for answering the query. In other words, if the most similar generated question isn't very similar to the query, then that question-answer pair would not be used for an answer response.
[0104] The identification process may use filtering parameters associated with the query to filter or otherwise select subsets of question-answer pairs in which candidates may be found. These filtering parameters maybe embedded within the query itself, specified as separate query parameters, configured in user settings, determined by system defaults based on the application context, or otherwise determined. For example, filtering parameters may restrict searches to question-answer pairs from specific sources, publication date ranges, author credentials, document types, subject matter domains, or source credibility metrics. The filtering maybe applied before similarity computation to reduce the search space and improve performance or applied after similarity scoring to refine results based on additional criteria.
[0105] As another variation, the ranking of candidate question-answer pairs may consider both similarity scores and metadata relevance to provide optimized results for the specific query context, enabling targeted searches within large question-answer pair databases while maintaining computational efficiency.
[0106] Block S150, which includes generating a response to the query by processing the candidate question-answer pairs, functions to algorithmically evaluate and determine answer(s) from the candidate question-answer pairs.
[0107] In a variation for generating an answer response, generating a response to the query by processing the candidate question-answer pairs comprises providing an answer from at least one of the candidate question-answer pairs. Once the candidate question-answer pairs are determined, then the answer and the corresponding segment of source information from each candidate question-answer pair may be used to computationally determine an answer response. In some variations, answers from each candidate question-answer pair may be supplied, possibly with corresponding citation or content segment for context. In some variations, a composite answer may be formed from the set of candidate question-answer pairs. This answergeneration approach leverages the pre-generated answers while maintaining traceability to the source segments that support each answer.
[0108] In one variation this may include: for each question-answer pair from the list of candidate question-answer pairs, evaluate or determine an answer by prompting an LLM to consider the source segment related to the question-answer pair and select the best answer to the received query from a list of options, where the options include at least the answer of the given question-answer pair. The options may additionally include false response options as well. The false response options may not necessarily be false but could be answers that are not the best answer or at least expected answer. The false response options could include a negation of the answer and / or an option that none of the options are accurate or all the options areaccurate. The false and negated answers reduce the odds that the LLM will select any particular answer, making the selection of an expected answer more likely to be based on the correctness of that answer than the probability of a random selection.
[0109] This step can ensure that the answer accurately answers the question of the received query and is supported by the segment of the source information. The configuration of the method can promote high reliability that the answer accurately responds to the query and is supported by verified information. Furthermore, because the forms of the answers are based on pre-generated formats, they may be phrased to match a desired format. In one exemplary implementation, the answers may be formatted to be concise and direct.
[0110] In one variation, if there are multiple different answers determined from different question-answer pairs, then these multiple answers may be returned in a response. Each answer may include or reference the associated segment of source information. In some variations, one or more answers from the candidate question-answer pairs may be the same. This may result in higher confidence scoring or ranking of this answer. This answer may be returned with the answer response including, or referencing, the different segments of source information that support the same conclusion.
[0111] In some variations, an automated clean-up process may be used to refine the generated output. This maybe done for generated question -answer pairs in S110 and / or for answers in S150. The question may be cleaned up so that the question has more natural phrasing, given the segment and / or the answer. Clean up processing may also generate a clean answer that is a more accurate and natural response to the cleaned-up question, given the segment. In some variations, the cleaned-up question may be generated to be short.Alternatively, the cleaned-up question may be generated to be long, for example by including relevant subject matter context. The cleaned-up question may also be generated to be similar to or in a format of a likely question one would ask within the query interface. The cleaned-up question maybe generated to use the language of the segment. The cleaned-up answer may also be very short and not repeat language from the cleaned-up question.
[0112] Additionally or alternatively, an automated clean-up process may similarly be used for refining the answer response generated for a query. Cleaning up an answer response may use an LLM or other NLP process to make the answer be in a form matching that of the query optimized for the intended application context.
[0113] In some variations, the question-answer pairs may be verified. Verification may be used for assessing alignment between questions and / or answers from the question-answer pair associated with the input and / or a stored reference question-answer pair.
[0114] In some variations, the question-answer pairs may be verified using a large language model to verify question-answer pair accuracy by evaluating whether an answer of a question-answer pair correctly answers a corresponding question based on the source segment. Verifying question-answer pair accuracy evaluates and determines if the generated questionanswer pairs make logical sense. This could be performed during the processing of questionanswer pairs. However, this may use processing / compute resources before they are needed. Accordingly, in some variations, verifying of question-answer pairs may be performed on- demand in response to a received query. For example, in some variations, using a large language model to verify question-answer pairs maybe performed on-demand in response to a received query to verify that a candidate answer correctly answers the received query before providing the answer in the response. In a similar manner, the verification may be performed just to confirm that an answer determined in S140 accurately answers the received query.
[0115] Different approaches may be used for verifying question-answer pairs. In some variations, generating a response to the query includes: for each question-answer pair from the candidate question-answer pairs, prompting a large language model to consider the source segment related to the question-answer pair and select the best answer to the received query from a list of options. The options include at least the answer of the given question-answer pair and false response options. This multi-choice verification approach may include presenting the question-answer pair to the large language model with multiple choice options including the candidate answer, false response options possibly including a negation of the answer, and options indicating none are correct. The verification process determines validity based on whether the large language model correctly selects the candidate answer, wherein the false response options include a negation of the answer to ensure the large language model correctly identifies the accurate answer.
[0116] In either case, question-answer pairs may be positively or negatively verified through automated assessment processes. If accuracy is confirmed (e.g., positively verified), then that question-answer pair may be updated such that verification is cached and not needed in subsequent processing. If the question-answer pair is found inaccurate (e.g., negatively verified), then that question-answer pair may be automatically removed from the data system or otherwise updated to mark as inaccurate. Such verification results can similarly be cached so that the question-answer pair need not be verified again in the future.
[0117] In some variations, using a large language model to verify question-answer pairs may be performed during initial processing of predicted question-answer pairs, wherein verification results are cached such that verification is not needed in subsequent processing. Inanother variation, a subset of question-answer pairs may be pre-verified while others may be verified on-demand. For example, a database of previous user queries maybe algorithmically evaluated to pre-verify related question-answer pairs of the top queries based on usage patterns and frequency analysis.
[0118] The processes and variations described herein may be implemented in varying ways. FIGURE 9 shows an exemplary architecture that maybe used in processing.
[0119] As mentioned, in some variations, verification may be used for validation which can function to assess information validity of a statement from the query input. Such a validation approach may use the same question-answer pair processing methodology but with comparison of a statement (or “answer”) from some input to the answers from reference question-answer pairs to assess content accuracy and reliability. This validation capability may enable applications such as automated fact-checking, misinformation detection, content credibility scoring, statement-to-source tracing, and augmentation of digital content based on its relationship to a reference set of sources (e.g., trusted sources or sources used as a point of comparison).
[0120] In one validation variation, the query comprises at least one statement for validation, and the method further comprises: generating a question based on the statement using language model processing; using the generated question to find corresponding pregenerated question-answer pairs; and generating a validation score for the statement based on comparison between the answer and the statement. In this approach, the statement for verification is processed to generate questions that can be used to find corresponding pregenerated question-answer pairs, then a validation score is generated for the statement based on computational comparison between the answer and the statement.
[0121] For example, a statement like "The capital of France is Paris" could generate the question "What is the capital of France?" which would be matched against pre-generated question-answer pairs from trusted geographical sources. The validation score would reflect the degree of correspondence between the statement and the trusted answer. This approach enables real-time fact-checking of individual claims or assertions against established knowledge bases.
[0122] The method may be particularly useful when using the statement validation process across a whole document or a collection of resources where multiple statements may be systematically validated and scored for how they align to one or more different reference information sources. As such, the method as shown in FIGURE 8 may include processing the query to identify and extract statements for validation and then for each statement: generating one or more questions based on the statement, using the generated question to findcorresponding pre-generated question-answer pairs and generating a validation score for the statement based on computational comparison between the answer and the statement.
[0123] A method variation for validation may alternatively be characterized as including: processing trusted source information using a large language model to generate reference question-answer pairs; storing the reference question-answer pairs in a database with reference to segments of the trusted source information; processing untrusted content using the large language model to generate comparison question-answer pairs; identifying questions from the comparison pairs that match questions from the reference pairs; evaluating correspondence between the reference answers from the matching reference question-answer pairs and statements in the untrusted content; and generating a content validity score for the statements based on the evaluated correspondence as shown in FIGURE 9. This validation method may enable systematic validation of entire documents, articles, LLM-generated text, or content collections against curated trusted sources such as peer-reviewed academic papers, government publications, and / or established reference materials. This validation method may also be used to enable tracing statements to sources in assumed, but not necessarily trusted, content like webpages on the internet. In general, the validation method described herein maybe used in finding appropriate materials for citation of information, for scoring (or otherwise characterizing) alignment to different reference materials.
[0124] The validation processing may include identifying statements within the untrusted content, and for each statement in the untrusted content, using the large language model to generate a set of questions that would be answered by the statement. For instance, processing a news article might identify statements about economic data, political events, or scientific claims, generating questions for each that can be cross-referenced with trusted sources. The identification of matching questions from comparison pairs and reference pairs comprises using the generated questions to search the database and identify matching reference question-answer pairs, enabling automated detection of potential misinformation or unsubstantiated claims.
[0125] In some variations, evaluating correspondence comprises directly comparing the reference answers to the statements in the untrusted content. This direct comparison approach is particularly effective for factual statements where exact matches or clear contradictions can be identified. In some variations, a large language model maybe used to evaluate comparison, but other approaches may be used.
[0126] Alternatively, evaluating correspondence may comprise: generating an answer from each statement in the untrusted content using the large language model; and comparingthe generated answers to the reference answers from the matching reference question-answer pairs.
[0127] The validation method may enable each statement in the untrusted content to receive a validation score based on correspondence between comparison answers and reference answers. These scores can range from high confidence (indicating strong alignment with trusted sources) to low confidence or conflict indicators (suggesting potential misinformation or disputed claims). The method may further include digitally annotating statements in the untrusted content with the content validity scores, wherein annotations include citations, footnotes, or indicators for confirmed, questionable, or inaccurate statements. This validation approach enables systematic assessment of content reliability and automated fact-checking capabilities that can be integrated into content management systems, publishing platforms, or browser-based verification tools. For example, a browser plugin could highlight validated statements in green, questionable content in yellow, and potential misinformation in red, while providing citations to the trusted sources that support or contradict each claim.4. Examples
[0128] Hereafter are described different examples of system and method variations. These examples are not intended to limit the systems and methods and their variations and do not include every variation and combination of variations of the systems and methods described herein.
[0129] Example 1.1: A method comprising: processing a corpus of source information using a large language model to generate a set of question-answer pairs; storing the questionanswer pairs in a predicted question-answer pair data system with reference to segments of the source information; receiving a query through a query interface; identifying candidate questionanswer pairs by searching the stored question-answer pairs using the query; generating a response to the query by processing the candidate question-answer pairs.
[0130] Example 2.1: A method comprising: processing trusted source information using a large language model to generate reference question-answer pairs; storing the reference question-answer pairs in a database with reference to segments of the trusted source information; processing untrusted content using the large language model to generate comparison question-answer pairs; identifying questions from the comparison pairs that match questions from the reference pairs; evaluating correspondence between the reference answers from the matching reference question-answer pairs and statements in the untrusted content;and generating a content validity score for the statements based on the evaluated correspondence.
[0131] Example 3.1: A method comprising: processing trusted source information using a large language model to generate reference question-answer pairs; storing the reference question-answer pairs in a database with reference to segments of the trusted source information; processing untrusted content to identify statements within the untrusted content; for each statement in the untrusted content, using the large language model to generate a set of questions that would be answered by the statement; using the generated questions to search the database and identify matching reference question-answer pairs; evaluating correspondence between the reference answers from the matching reference question-answer pairs and the statement; and generating a content validity score for the statement based on the evaluated correspondence.
[0132] Example 4.1: A non-transitory computer-readable medium storing instructions that, when executed by one or more computer processors of a computing platform, cause the computing platform to perform operations comprising: processing a corpus of source information using a large language model to generate a set of question-answer pairs; storing the question-answer pairs in a predicted question-answer pair data system with reference to segments of the source information; receiving a query through a query interface; identifying candidate question-answer pairs by searching the stored question-answer pairs using the query; generating a response to the query by processing the candidate question-answer pairs.
[0133] Example 5.1: A system comprising of: one or more computer-readable mediums storing instructions that, when executed by the one or more computer processors, cause a computing platform to perform operations comprising: processing a corpus of source information using a large language model to generate a set of question-answer pairs; storing the question-answer pairs in a predicted question-answer pair data system with reference to segments of the source information; receiving a query through a query interface; identifying candidate question-answer pairs by searching the stored question-answer pairs using the query; generating a response to the query by processing the candidate question-answer pairs.
[0134] Example 1.2: A variation of example 1.1, 4.1, 5.1 and / or any of the other system and method examples and variations herein, wherein processing the corpus of source information to generate the set of question-answer pairs comprises: processing multiple information segments of the source information; identifying potential answers for a given information segment; and for each potential answer, generating a set of questions that can beanswered with the potential answer based on the segment. Example 1.2 may be used in combination with example 1.1, 4.1, 5.1 and / or their variations or combinations.
[0135] Example 1.3: A variation of example 1.2 and / or any of the other system and method examples and variations herein, wherein generating a response to the query by processing the candidate question-answer pairs comprises providing an answer from at least one of the candidate question-answer pairs. Example 1.3 may be used in combination with example 1.1, 1.2, 4.1, 5.1 and / or their variations or combinations.
[0136] Example 1.4: A variation of example 1.2 and / or any of the other system and method examples and variations herein, wherein identifying potential answers from a segment uses a large language model. Example 1.4 maybe used in combination with example 1.1, 1.2, 1.3, 4.1, 5.1 and / or their variations or combinations.
[0137] Example 1.5: A variation of example 1.1, 4.1, 5.1 and / or any of the other system and method examples and variations herein, wherein identifying candidate question-answer pairs comprises performing a vector database search using embedding-based similarity computation between the query and stored questions of the predicted question-answer pairs in the predicted question-answer pair data system. Example 1.5 may be used in combination with example 1.1, 1.2, 1.3, 1.4, 4.1, 5.1 and / or their variations or combinations.
[0138] Example 1.6: A variation of example 1.2 and / or any of the other system and method examples and variations herein, wherein processing multiple information segments comprises composing segments from the source information. Example 1.6 may be used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 4.1, 5.1 and / or their variations or combinations.
[0139] Example 1.7: A variation of example 1.2 and / or any of the other system and method examples and variations herein, wherein processing multiple information segments comprises: segmenting the source information with varying segment sizes, wherein a subset of segments overlaps with at least one other segment. Example 1.7 may be used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 4.1, 5.1 and / or their variations or combinations.
[0140] Example 1.8: A variation of example 1.2 and / or any of the other system and method examples and variations herein, wherein generating the set of questions for each potential answer comprises generating multiple different questions that can be answered by the same potential answer. Example 1.8 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 4.1, 5.1 and / or their variations or combinations.
[0141] Example 1.9: A variation of example 1.1, 4.1, 5.1 and / or any of the other system and method examples and variations herein, further comprising: determining at least two question-answer pairs satisfying a similarity condition; and consolidating the at least twoquestion-answer pairs by storing a consolidated question-answer pair with multiple references to source information segments. Example 1.9 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 4.1, 5.1 and / or their variations or combinations.
[0142] Example 1.10: A variation of example 1.1, 4.1, 5.1 and / or any of the other system and method examples and variations herein, further comprising: using a large language model to verify question-answer pair accuracy by evaluating whether an answer of a question-answer pair correctly answers a corresponding question based on the source segment. Example 1.10 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 4.1, 5.1 and / or their variations or combinations.
[0143] Example 1.11: A variation of example 1.10 and / or any of the other system and method examples and variations herein, wherein generating a response to the query comprises: for each question-answer pair from the candidate question-answer pairs, prompting a large language model to consider the source segment related to the question-answer pair and select the best answer to the received query from a list of options, where the options include at least the answer of the given question-answer pair and false response options. Example 1.11 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 1.10, 4.1, 5.1 and / or their variations or combinations.
[0144] Example 1.12: A variation of example 1.10 and / or any of the other system and method examples and variations herein, wherein using a large language model to verify question-answer pairs is performed on-demand to verify that a candidate answer correctly answers the received query before providing the answer in the response. Example 1.12 may be used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 1.10, 1.11, 4.1, 5.1 and / or their variations or combinations.
[0145] Example 1.13: A variation of example 1.10 and / or any of the other system and method examples and variations herein, wherein using a large language model to verify question-answer pairs is performed during initial processing of predicted question-answer pairs. Example 1.13 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 1.10, 1.11, 1.12, 4.1, 5.1 and / or their variations or combinations.
[0146] Example 1.14: A variation of example 1.1, 4.1, 5.1 and / or any of the other system and method examples and variations herein, wherein the source information comprises textual data from a set of text documents. Example 1.14 may be used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 1.10, 1.11, 1.12, 1.13, 4.1, 5.1 and / or their variations or combinations.
[0147] Example 1.15: A variation of example 1.1, 4.1, 5.1 and / or any of the other system and method examples and variations herein, wherein the source information comprises audiobased media content, further comprising generating transcripts from the media content; and wherein processing a corpus of source information data using a large language model to generate a set of question-answer pairs processes at least the transcripts of the media content. Example 1.15 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 1.10, 1.11, 1.12, 1.13, 1.14, 4.1, 5.1 and / or their variations or combinations.
[0148] Example 1.16: A variation of example 1.1, 4.1, 5.1 and / or any of the other system and method examples and variations herein, wherein the source information comprises database content, further comprising generating questions and answers from data in the database content. Example 1.16 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5,1.6. 1.7. 1.8. 1.9. 1.10. 1.11. 1.12. 1.13. 1.14. 1.15. 4.1, 5.1 and / or their variations or combinations.
[0149] Example 1.17: A variation of example 1.1, 4.1, 5.1 and / or any of the other system and method examples and variations herein, wherein the source information comprises computer code repository content. Example 1.17 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 1.10, 1.11, 1.12, 1.13, 1.14, 1.15, 1.16, 4.1, 5.1 and / or their variations or combinations.
[0150] Example 1.18: Avariation of example 1.1, 4.1, 5.1 and / or any of the other system and method examples and variations herein, wherein the query with a statement for verification, and the method further comprises: generating a question based on the statement; using the generated question to find corresponding pre-generated question-answer pairs; and generating a validation score for the statement based on comparison between the answer and the statement. Example 1.18 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7,1.8. 1.9. 1.10. 1.11. 1.12. 1.13. 1.14. 1.15. 1.16. 1.17. 4.1, 5.1 and / or their variations or combinations.
[0151] Example 1.19: A variation of example, 1.1, 4.1 and 5.1 and / or any of the other system and method examples and variations herein, wherein the source information comprises video-based media content, further comprising converting the video content into textual representations using language model processing prior to generating question-answer pairs. In some variations, converting the video content into textual representations may include at least one of: generating transcripts from audio portions of the video content, generating descriptions of visual content using multimodal language models, and extracting textual information displayed within the video content. Example 1.19 maybe used in combination with example 1.1,1.2. 1.3. 1.4. 1.5. 1.6. 1.7. 1.8. 1.9. 1.10. 1.11. 1.12. 1.13. 1.14. 1.15. 1.16. 1.17. 1.18. 4.1, 5.1 and / or their variations or combinations.
[0152] Example 1.20: A variation of example, 1.1, 4.1 and 5.1 and / or any of the other system and method examples and variations herein, wherein the query comprises query input in at least one format selected from text format, audio format, and video format. In some variations, the method may further include wherein audio and / or video formats are processed to extract textual information prior to query processing. However, multi-modal models may alternatively be used for direct interpretation of the audio and / or video media content. Example 1.20 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 1.10, 1.11, 1.12, 1.13, 1.14, 1.15, 1.16, 1.17, 1.18, 1.19, 4.1, 5.1 and / or their variations or combinations.
[0153] Example 1.21: A variation of example 1.1, 4.1, 5.1 and / or any of the other system and method examples and variations herein, wherein processing the corpus generates more than one million predicted question-answer pairs. Example 1.21 maybe used in combination with example 1.1, 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 1.10, 1.11, 1.12, 1.13, 1.14, 1.15, 1.16, 1.17, 1.18, 1.19, 1.20, 4.1, 5.1 and / or their variations or combinations.
[0154] Example 2.2: A variation of example 2.1 and / or any of the other system and method examples and variations herein, further comprising: digitally annotating statements in the untrusted content with the content validity scores, wherein annotations include citations, footnotes, or indicators for confirmed, questionable, or inaccurate statements. Example 2.2 may be used in combination with example 2.1, 3.1 and / or their variations or combinations.
[0155] Example 2.3: A variation of example 2.1 and / or any of the other system and method examples and variations herein, wherein each statement in the untrusted content receives a validation score based on correspondence between comparison answers and reference answers. Example 2.3 may be used in combination with example 2.1, 2.2, 3.1 and / or their variations or combinations.
[0156] Example 2.4: A variation of example 2.1 and / or any of the other system and method examples and variations herein, wherein evaluating correspondence comprises directly comparing the reference answer to the statements in the untrusted content. Example 2.4 maybe used in combination with example 2.1, 2.2, 2.3, 3.1 and / or their variations or combinations.
[0157] Example 2.5: A variation of example 2.1 and / or any of the other system and method examples and variations herein, wherein evaluating correspondence comprises: generating an answer from each statement in the untrusted content using the large language model; and comparing the generated answers to the reference answers from the matching reference question-answer pairs. Example 2.5 maybe used in combination with example 2.1, 2.2, 2.3, 2.4, 3.1 and / or their variations or combinations.
[0158] Example 2.6: A variation of example 2.1 and / or any of the other system and method examples and variations herein, wherein: processing untrusted content using the large language model to generate comparison question-answer pairs comprises: identifying statements within the untrusted content, for each statement in the untrusted content, using the large language model to generate a set of questions that would be answered by the statement. Example 2.6 maybe used in combination with example 2.1, 2.2, 2.3, 2.4, 2.5, 3.1 and / or their variations or combinations.4. System Architecture
[0159] The systems and methods of the embodiments can be embodied and / or implemented at least in part as a machine configured to receive a computer-readable medium storing computer-readable instructions. The instructions can be executed by computerexecutable components integrated with the application, applet, host, server, network, website, communication service, communication interface, hardware / firmware / software elements of a user computer or mobile device, wristband, smartphone, or any suitable combination thereof. Other systems and methods of the embodiment can be embodied and / or implemented at least in part as a machine configured to receive a computer-readable medium storing computer- readable instructions. The instructions can be executed by computer-executable components integrated with apparatuses and networks of the type described above. The computer-readable medium can be stored on any suitable computer readable media such as RAMs, ROMs, flash memory, EEPROMs, optical devices (CD or DVD), hard drives, floppy drives, or any suitable device. The computer-executable component can be a processor, but any suitable dedicated hardware device can (alternatively or additionally) execute the instructions.
[0160] Tn one variation, a system comprising of one or more computer- readable mediums (e.g., non-transitory computer-readable mediums) storing instructions that, when executed by the one or more computer processors, cause a computing platform to perform operations comprising those of the system or method described herein such as: processing a corpus of source information to generate a set of question-answer pairs; storing the questionanswer pairs in a question-answer pair data system with reference to segments of the source information; receiving a query through a query interface; identifying candidate question-answer pairs by searching the stored question-answer pairs using the query; and generating a response to the query by processing the candidate question-answer pairs.
[0161] FIGURE 11 is an exemplary computer architecture diagram of one implementation of the system. In some implementations, the system is implemented in a plurality of devices in communication over a communication channel and / or network. In some implementations, the elements of the system are implemented in separate computing devices. In some implementations, two or more of the system elements are implemented in the same device. The system and portions of the system maybe integrated into a computing device or system that can serve as or within the system.
[0162] The communication channel 1001 interfaces with the processors 1002A-1002N, the memory (e.g., a random-access memory (RAM)) 1003, a read only memory (ROM) 1004, a processor-readable storage medium 1005, a display device 1006, a user input device 1007, and a network device 1008. As shown, the computer infrastructure may be used in connecting question-answer pair generator 1101, predicted question-answer pair data system 1102, query interface 1103, question-answer pair matching engine 1104, question-answer verification engine 1105, content validation engine 1106, and / or other suitable computing devices.
[0163] The processors 1002A-1002N may take many forms, such CPUs (CentralProcessing Units), GPUs (Graphical Processing Units), microprocessors, ML / DL (Machine Learning / Deep Learning) processing units such as a Tensor Processing Unit, FPGA (Field Programmable Gate Arrays, custom processors, and / or any suitable type of processor.
[0164] The processors 1002A-1002N and the main memory 1003 (or some subcombination) can form a processing unit 1010. In some embodiments, the processing unit includes one or more processors communicatively coupled to one or more of a RAM, ROM, and machine-readable storage medium; the one or more processors of the processing unit receive instructions stored by the one or more of a RAM, ROM, and machine-readable storage medium via a bus; and the one or more processors execute the received instructions. In some embodiments, the processing unit is an ASIC (Application-Specific Integrated Circuit). In some embodiments, the processing unit is a SoC (System-on-Chip). In some embodiments, the processing unit includes one or more of the elements of the system.
[0165] A network device 1008 may provide one or more wired or wireless interfaces for exchanging data and commands between the system and / or other devices, such as devices of external systems. Such wired and wireless interfaces include, for example, a universal serial bus (USB) interface, Bluetooth interface, Wi-Fi interface, Ethernet interface, near field communication (NFC) interface, and the like.
[0166] Computer and / or Machine-readable executable instructions comprising of configuration for software programs (such as an operating system, application programs, anddevice drivers) can be stored in the memory 1003 from the processor-readable storage medium 1005, the ROM 1004 or any other data storage system.
[0167] When executed by one or more computer processors, the respective machineexecutable instructions may be accessed by at least one of processors 1002A-1002N (of a processing unit 1010) via the communication channel 1001, and then executed by at least one of processors 1001A-1001N. Data, databases, data records or other stored forms data created or used by the software programs can also be stored in the memory 1003, and such data is accessed by at least one of processors 1002A-1002N during execution of the machine-executable instructions of the software programs.
[0168] The processor-readable storage medium 1005 is one of (or a combination of two or more of) a hard drive, a flash drive, a DVD, a CD, an optical disk, a floppy disk, a flash storage, a solid-state drive, a ROM, an EEPROM, an electronic circuit, a semiconductor memory device, and the like. The processor-readable storage medium 1005 can include an operating system, software programs, device drivers, and / or other suitable sub-systems or software.
[0169] As used herein, first, second, third, etc. are used to characterize and distinguish various elements, components, regions, layers and / or sections. These elements, components, regions, layers and / or sections should not be limited by these terms. Use of numerical terms maybe used to distinguish one element, component, region, layer and / or section from another element, component, region, layer and / or section. Use of such numerical terms does not imply a sequence or order unless clearly indicated by the context. Such numerical references may be used interchangeable without departing from the teaching of the embodiments and variations herein.
[0170] As a person skilled in the art will recognize from the previous detailed description and from the figures and claims, modifications and changes can be made to the embodiments of the invention without departing from the scope of this invention as defined in the following claims.
Claims
CLAIMSWe Claim:
1. A method comprising: processing a corpus of source information using a large language model to generate a set of question-answer pairs; storing the question-answer pairs in a predicted question-answer pair data system with reference to segments of the source information; receiving a query through a query interface; identifying candidate question-answer pairs by searching the stored questionanswer pairs using the query; generating a response to the query by processing the candidate question-answer pairs.
2. The method of claim 1, wherein processing the corpus of source information to generate the set of question-answer pairs comprises: processing multiple information segments of the source information; identifying potential answers for a given information segment; and for each potential answer, generating a set of questions that can be answered with the potential answer based on the segment.
3. The method of claim 2, wherein generating a response to the query by processing the candidate question-answer pairs comprises providing an answer from at least one of the candidate question-answer pairs.
4. The method of claim 2, wherein identifying potential answers from a segment uses a large language model.
5. The method of claim 1, wherein identifying candidate question-answer pairs comprises performing a vector database search using embedding-based similarity computation between the query and stored questions of the predicted questionanswer pairs in the predicted question-answer pair data system.
6. The method of claim 2, wherein processing multiple information segments comprises composing segments from the source information.
7. The method of claim 2, wherein processing multiple information segments comprises: segmenting the source information with varying segment sizes, wherein a subset of segments overlaps with at least one other segment.
8. The method of claim 2, wherein generating the set of questions for each potential answer comprises generating multiple different questions that can be answered by the same potential answer.
9. The method of claim 1, further comprising: determining at least two question-answer pairs satisfying a similarity condition; and consolidating the at least two questionanswer pairs by storing a consolidated question-answer pair with multiple references to source information segments.
10. The method of claim 1, further comprising: using a large language model to verify question-answer pair accuracy by evaluating whether an answer of a questionanswer pair correctly answers a corresponding question based on the source segment.
11. The method of claim 10, wherein generating a response to the query comprises: for each question-answer pair from the candidate question-answer pairs, prompting a large language model to consider the source segment related to the question-answer pair and select the best answer to the received query from a list of options, where the options include at least the answer of the given question-answer pair and false response options.
12. The method of claim 10, wherein using a large language model to verify questionanswer pairs is performed on-demand to verify that a candidate answer correctly answers the received query before providing the answer in the response.
13. The method of claim 10, wherein using a large language model to verify questionanswer pairs is performed during initial processing of predicted question-answer pairs.
14. The method of claim 1, wherein the source information comprises textual data from a set of text documents.
15. The method of claim 1, wherein the source information comprises audio-based media content, further comprising generating transcripts from the media content; and wherein processing a corpus of source information using a large language modelto generate a set of question-answer pairs processes at least the transcripts of the media content.
16. The method of claim i, wherein the source information comprises database content, further comprising generating questions and answers from data in the database content.
17. The method of claim 1, wherein the source information comprises computer code repository content.
18. The method of claim 1, wherein the query comprises a statement for verification, and the method further comprises: generating a question based on the statement; using the generated question to find corresponding pre-generated question-answer pairs; and generating a validation score for the statement based on comparison between the answer and the statement.
19. The method of claim 1, wherein the source information comprises video-based media content, further comprising converting the video content into textual representations using language model processing prior to generating question-answer pairs.
20. The method of claim 1, wherein the query comprises query input in at least one format selected from text format, audio format, and video format.
21. The method of claim 1, wherein processing the corpus generates more than one million predicted question-answer pairs.
22. A method comprising: processing a corpus of source information using a large language model to generate reference question-answer pairs; storing the reference question-answer pairs in a database with reference to segments of the source information; processing untrusted content using the large language model to generate comparison question-answer pairs; identifying questions from the comparison pairs that match questions from the reference pairs; evaluating correspondence between the reference answers from the matching reference question-answer pairs and statements in the untrusted content; andgenerating a content validity score for the statements based on the evaluated correspondence.
23. The method of claim 22, further comprising: digitally annotating statements in the untrusted content with the content validity scores, wherein annotations include citations, footnotes, or indicators for confirmed, questionable, or inaccurate statements.
24. The method of claim 22, wherein each statement in the untrusted content receives a validation score based on correspondence between comparison answers and reference answers.
25. The method of claim 22, wherein evaluating correspondence comprises directly comparing the reference answer to the statements in the untrusted content.
26. The method of claim 22, wherein evaluating correspondence comprises: generating an answer from each statement in the untrusted content using the large language model; and comparing the generated answers to the reference answers from the matching reference question-answer pairs.
27. The method of claim 22, wherein: processing untrusted content using the large language model to generate comparison question-answer pairs comprises: identifying statements within the untrusted content, for each statement in the untrusted content, using the large language model to generate a set of questions that would be answered by the statement.
28. A method comprising: processing a corpus of source information using a large language model to generate reference question-answer pairs; storing the reference question-answer pairs in a database with reference to segments of the source information; processing untrusted content to identify statements within the untrusted content; for each statement in the untrusted content, using the large language model to generate a set of questions that would be answered by the statement; using the generated questions to search the database and identify matching reference question-answer pairs;evaluating correspondence between the reference answers from the matching reference question-answer pairs and the statement; and generating a content validity score for the statement based on the evaluated correspondence.
29. A non-transitory computer-readable medium storing instructions that, when executed by one or more computer processors of a computing platform, cause the computing platform to perform operations comprising: processing a corpus of source information using a large language model to generate a set of question-answer pairs; storing the question-answer pairs in a predicted question-answer pair data system with reference to segments of the source information; receiving a query through a query interface; identifying candidate question-answer pairs by searching the stored questionanswer pairs using the query; generating a response to the query by processing the candidate question-answer pairs.
30. A system comprising of: one or more computer-readable mediums storing instructions that, when executed by the one or more computer processors, cause a computing platform to perform operations comprising: processing a corpus of source information using a large language model to generate a set of question-answer pairs; storing the question-answer pairs in a predicted question-answer pair data system with reference to segments of the source information; receiving a query through a query interface; identifying candidate question-answer pairs by searching the stored questionanswer pairs using the query; generating a response to the query by processing the candidate questionanswer pairs.
Citation Information
Patent Citations
Method and device for extracting structured knowledge from document, equipment and medium
CN117194632A
Question and answer pair extraction method and device, electronic equipment and storage medium
CN117407502A
A method and device for constructing question-answering data based on a large language model
CN117591661B
Dialogue method and system based on document retrieval enhanced machine language model
CN117807199A
Code retrieval processing method and device
CN117933395A
Cited By
Synthetic knowledge ingestion for enhancing large language model performance
US20260119537A1