Structured Data Tokenization for Better LLM Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large language models struggle with input data formats that are not in natural language, leading to poor quality outputs and unobservable inner workings, making them difficult to use effectively.

Innovation Solution

Processing event and tabular data into sequential data tokens, embeddings, and other forms suitable for input to large language models using techniques like alphanumeric transformation, template, embedding projection, clustering, and decision trees, to improve data suitability and output quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If event data and tabular data are provided directly to large language models without processing, then the input process is simple, but the output quality is poor and the model cannot effectively process non-natural language formats

Engineering Contradiction:
Improveoutput qualityVSAvoiddata processing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by transforming event data and tabular data into sequential token formats before they are input to the large language model. This pre-processing step converts structured data into a format the model can effectively process, thereby improving output quality without requiring changes to the model itself. The transformation includes mapping data values to token sequences that preserve the original data's semantic meaning while adhering to the model's input requirements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary data transformation layer between the structured data source and the large language model. This intermediary component acts as a mediator that converts event data and tabular data into sequential token representations, enabling the model to process non-natural language inputs effectively. The intermediary transformation process includes techniques such as value mapping, sequence generation, and format conversion, which bridge the gap between structured data and the model's natural language processing capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If large language models are used for processing structured data, then linguistic tasks can be performed, but the inner workings of the model remain unobservable and difficult to interpret

Engineering Contradiction:
Improvetask performance capabilityVSAvoidmodel interpretability
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies feedback by implementing observation and interpretation mechanisms that provide insights into the large language model's processing of transformed data. The system observes how the model processes sequential tokens derived from structured data and provides feedback on the transformation effectiveness. This feedback loop enables users to understand and measure the model's inner workings, improving interpretability while maintaining the model's versatile task performance capability.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If data transformation techniques are applied to improve data suitability, then output accuracy improves, but the processing time and computational resources increase

Engineering Contradiction:
Improveclassification accuracyVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies parameter changes by systematically transforming data parameters from structured formats into sequential token representations suitable for large language models. The transformation process modifies data parameters such as value encoding, sequence length, and tokenization strategies to optimize both accuracy and processing efficiency. By carefully managing these parameter transformations, the system achieves improved classification accuracy while minimizing the time and computational resources required for data preparation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12608546B2Processing event data and/or tabular data for input to one or more machine learning models
Publication Date: 2026.04.21 CAPITAL ONE SERVICES LLC
  • US12608546B2 patent drawing
  • US12608546B2 patent drawing
  • US12608546B2 patent drawing

AI summary

Aspects described herein may relate to techniques and/or methods that process certain forms of event data and/or tabular data for input to one or more machine learning models, such as a large language model. Additional aspects may relate to using the output of a large language model as part of a process for detecting fraud based on the event data and/or tabular data. In some variations, the event data and/or tabular data may be processed into data tokens, embeddings, or other forms of data suitable for use as input to a large language model.