LLM Continuation Filtering Against Training Text for Original Output
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer applications struggle with processing heterogeneous and broad data collections, leading to impractical schema building and coding challenges, and current large language models (LLMs) generate inaccurate and unreliable outputs due to their statistical nature and lack of true understanding.
Innovation Solution
Utilize a structured, machine-readable representation of data, such as a universal language (UL), to interact with LLMs, improve output generation, and validate factual accuracy, while avoiding hallucinations and adding citations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If statistical machine learning and deep learning are used to process natural language, then significant progress can be made with many problems, but the results cannot be explained in a way that makes sense to human users and the system lacks real understanding of the data
Solution Approach 1:
The patent introduces an intermediary layer between the statistical learning model and the user interface. This intermediary translates the model's internal representations into human-understandable explanations, mediating between the black-box computational process and human comprehension needs.
Solution Approach 2:
The patent replaces purely statistical correlation-based systems with a hybrid approach that incorporates symbolic reasoning and explanation-generation mechanisms, substituting mechanical statistical processing with more interpretable computational processes.
2Adaptability or versatility
If statistical machine learning models are trained with randomly initiated weights improved through training, then the models can learn from data, but the system is inherently unreliable and unable to reliably know when the result it produces is accurate
Solution Approach 1:
The patent implements feedback mechanisms where the system continuously evaluates its own confidence in generated results and adjusts its output accordingly. The model receives feedback about the reliability of its predictions and uses this to determine when to express uncertainty or seek additional information.
Solution Approach 2:
The patent performs preliminary confidence assessment before final output generation. The system evaluates the reliability of its predictions in advance and prepares appropriate responses, including indicating uncertainty when confidence is low, before presenting results to users.
3Ease of operation
If LLMs generate text output based on statistical patterns, then they can produce continuous text, but the output is frequently incorrect and often describes things that are not true
Solution Approach 1:
The patent merges statistical language generation capabilities with factual verification mechanisms. The system combines the fluency and coherence of LLMs with fact-checking processes that verify the accuracy of generated content before presentation to users.
Solution Approach 2:
The patent applies preliminary anti-action by implementing fact-checking and verification processes that counteract potential factual errors before they are presented to users. The system proactively identifies and corrects inaccurate information in generated text.
Data Source
AI summary
There is provided a computer-implemented method for ensuring that a large language model (LLM) generates original text, including (i) providing or accessing a database of previous text that the LLM should not generate, wherein the database includes text used to train the LLM; (ii) checking potential continuations generated by the LLM against the database; (iii) when a potential continuation generated by the LLM matches text in the database, adjusting the potential continuation generated by the LLM to no longer match that text in the database, to produce an adjusted potential continuation, and (iv) storing the adjusted potential continuation.


