LLM Context Structuring for Reliable Heterogeneous Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer systems struggle to deeply understand and process natural language, leading to limitations in handling heterogeneous and broad data sets required for complex applications, such as health data management and accounting, where existing methods rely on structured data or statistical language models that are unreliable and lack explainability.

Innovation Solution

A method involving a structured, machine-readable representation of data, such as a universal language, is used to improve the output of large language models by providing new context and enabling improved continuation text generation, validation of natural language for factual accuracy, and avoidance of hallucination, while ensuring original text generation and adding citations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If structured data is used to store information for processing, then data processing reliability is improved, but the ability to handle heterogeneous and broad data sets deteriorates

Engineering Contradiction:
Improvedata processing reliabilityVSAvoidability to handle heterogeneous data
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary layer (structured data representation with semantic meaning) between natural language input and computer processing. This intermediary maintains structured organization for reliable processing while being able to represent heterogeneous information through flexible schema design and natural language integration.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a composite data structure that combines elements of structured data (for reliability) and natural language (for versatility). This composite approach allows the system to handle both organized information and unstructured heterogeneous data within a unified framework.

Inventive Principle:
Principle #40Composite materials

2Adaptability or versatility

If statistical language models are used to process natural language, then language processing capability is improved, but reliability and explainability deteriorate

Engineering Contradiction:
Improvelanguage processing capabilityVSAvoidmodel reliability and explainability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent replaces the statistical/machine learning-based language processing mechanism with a rule-based, logic-driven approach. This substitution maintains language processing capability while improving reliability and explainability by using deterministic rules and structured reasoning instead of probabilistic models.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system incorporates self-validation and self-correction mechanisms where the structured data representation and logical processing rules enable the system to verify its own output reliability and explain its reasoning processes without external intervention.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If a broad schema is created to cover all possible data types, then data coverage is improved, but system complexity and difficulty of maintenance deteriorate

Engineering Contradiction:
Improvedata coverageVSAvoidschema complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the broad schema into modular, hierarchical components that can be independently managed and maintained. This segmentation allows comprehensive data coverage through composition of smaller units while reducing overall system complexity through organized modularity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates universal data structure templates and patterns that can represent multiple types of heterogeneous data through parameterization and inheritance. This universality achieves broad data coverage without requiring separate complex schemas for each data type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12067362B2Computer implemented methods for the automated analysis or use of data, including use of a large language model
Publication Date: 2024.08.20 UNLIKELY ARTIFICIAL INTELLIGENCE LTD
  • US12067362B2 patent drawing
  • US12067362B2 patent drawing
  • US12067362B2 patent drawing

AI summary

Methods are provided, such as a method of interacting with a large language model (LLM), including the step of a processing system using a structured, machine-readable representation of data that conforms to a machine-readable language, such as a universal language, to provide new context data for the LLM, in order to improve the output, such as continuation text output, generated by the LLM in response to a prompt; and such as a method of interacting with a LLM, including the step of providing continuation data generated by the LLM to a processing system that uses a structured, machine-readable representation of data that conforms to a machine-readable language, such as a universal language, in which the processing system is configured to analyse the continuation output generated by the LLM in response to a prompt to enable an improved version of that continuation output to be provided to a user. Related computer systems are provided.