Grounded Document Generation via LLM Framework and Source Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large language models (LLMs) face challenges in generating accurate and reliable outputs, often producing false or incorrect information (hallucinations), which affects their reliability and trustworthiness, and existing solutions are limited by linear chat capabilities, limited context, and reliance on public web data without support for private documents or constrained responses.

Innovation Solution

The method involves using one or more LLMs to automatically generate documents with sections and subsections, providing references to data sources used, allowing users to customize the framework and data sources, and generating reports or surveys with focused content based on user prompts, thereby ensuring accuracy and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If LLMs generate text fluently and coherently, then the quality and readability of output improve, but accuracy and reliability deteriorate due to hallucinations

Engineering Contradiction:
Improvetext generation qualityVSAvoidoutput accuracy
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent introduces an intermediary verification mechanism where retrieved data sources and references act as mediators between the LLM generation process and the final output. The system retrieves evidence from external data sources and uses these as intermediaries to verify and ground the LLM's generated content, reducing hallucinations while preserving fluency

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback loops where the generated output is continuously verified against retrieved data sources. The system provides feedback by checking whether generated claims are supported by evidence from external sources, and iteratively refines the output to improve both accuracy and reliability

Inventive Principle:
Principle #23Feedback

2Reliability

If LLMs are constrained to provide grounded responses with references, then reliability improves, but flexibility and ease of operation worsen

Engineering Contradiction:
Improveresponse groundednessVSAvoidresponse flexibility
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent implements dynamic response generation where the level of grounding and referencing is adaptable based on user preferences and query context. The system can dynamically adjust between highly grounded responses with full references and more flexible responses, allowing users to choose their preferred interaction mode

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If LLMs rely on public web data, then data accessibility improves, but support for private documents and controlled responses deteriorates

Engineering Contradiction:
Improvedata source accessibilityVSAvoiddata source control
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal data retrieval system that can handle multiple types of data sources including public web data, private documents, and proprietary databases through a unified interface. The system uses multi-functionality to support both open-access and controlled-environment data sources, allowing users to specify which data sources to use based on their needs

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240428005A1Generating grounded documents using large language models
Publication Date: 2024.12.26 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240428005A1 patent drawing
  • US20240428005A1 patent drawing
  • US20240428005A1 patent drawing

AI summary

The present disclosure relates to methods and systems for automatically generating documents for a specific topic using large language models. The methods and systems receive an input query that identifies a topic for the document. The methods and systems automatically generate, using the large language models, a framework for the document with sections and subsections for the document. The methods and systems write the document, using the large language models, and provide references for the data sources used to obtain the data that the large language model used to write the document.