Language Model Prompting with Embedding Vectors for Citation Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Language models (LMs) face challenges such as hallucinations, inconsistencies across versions, and unreliable information, which affect their accuracy and reliability in generating responses to natural language questions.
Innovation Solution
A computer-implemented method that involves receiving a natural language question, determining its relevance to a reference document, computing an embedding vector for the question, selecting relevant question-and-answer pairs, creating a prompt for a language model that includes the question, selected pairs, and an expected output format requesting a quotation or citation from the reference document, submitting the prompt to the language model, computing response scores to evaluate the accuracy of the responses, and selecting the most accurate response to determine the answer to the question.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If language models are used to generate responses to natural language questions, then the ability to process and comprehend natural language is improved, but hallucinations and unreliable information occur
Solution Approach 1:
The patent introduces an intermediary verification system that acts as a mediator between the language model and the final output. This system includes: (1) extracting entities and claims from the language model's response, (2) searching for evidence in knowledge bases and web sources, (3) verifying claims against found evidence, and (4) generating corrected responses when hallucinations are detected. This intermediary layer maintains the natural language processing capability while significantly improving reliability by filtering and verifying the model's outputs.
Solution Approach 2:
The patent implements a feedback mechanism where the system continuously monitors language model outputs for hallucinations and unreliable information. When issues are detected, the system provides feedback by: (1) identifying specific problematic claims, (2) searching for correct information, (3) generating corrected responses, and (4) using these corrections to improve future model outputs. This closed-loop feedback system progressively improves reliability while maintaining ease of operation.
2Productivity
If language models generate responses without verification, then response speed is improved, but inconsistencies across versions occur
Solution Approach 1:
The patent applies preliminary action by pre-processing the language model's response before it is fully generated or immediately upon generation. The system extracts entities and claims, searches for evidence, and verifies accuracy in advance of final output delivery. This preliminary verification ensures consistency across different model versions without significantly impacting response speed, as the verification process is efficiently integrated into the generation workflow.
Solution Approach 2:
The patent replaces manual verification mechanisms with automated computational systems. Instead of relying on human reviewers or complex manual consistency checks, the system uses automated entity extraction, claim verification algorithms, and evidence-matching processes. This substitution maintains high response speed while ensuring consistent verification across different model versions, as the automated system applies the same verification criteria uniformly.
3Ease of operation
If language models are used without reference documents, then ease of use is improved, but factual accuracy deteriorates
Solution Approach 1:
The patent implements self-service by enabling the language model system to automatically retrieve and verify information from reference documents, knowledge bases, and web sources without requiring user intervention. The system autonomously: (1) identifies when verification is needed, (2) searches appropriate sources, (3) validates claims, and (4) generates accurate responses. This maintains ease of use for users while dramatically improving factual accuracy through automated reference checking.
Data Source
AI summary
Applications of a language model are improved by creating a prompt that includes representations that improve the accuracy and reliability of the language model. The representations may be added to a prompt generated by a user. The representations are selected using embedding vectors of the representations and the user-generated prompt. Responses of the language model are selected by computing scores of the responses.


