Structured Data Tokenization for Better LLM Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large language models struggle with input data formats that are not in natural language, leading to poor quality outputs and unobservable inner workings, making them difficult to use effectively.
Innovation Solution
Processing event and tabular data into sequential data tokens, embeddings, and other forms suitable for input to large language models using techniques like alphanumeric transformation, template, embedding projection, clustering, and decision trees, to improve data suitability and output quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If event data and tabular data are provided directly to large language models without processing, then the input process is simple, but the output quality is poor and the model cannot effectively process non-natural language formats
Solution Approach 1:
The patent applies preliminary action by transforming event data and tabular data into sequential token formats before they are input to the large language model. This pre-processing step converts structured data into a format the model can effectively process, thereby improving output quality without requiring changes to the model itself. The transformation includes mapping data values to token sequences that preserve the original data's semantic meaning while adhering to the model's input requirements.
Solution Approach 2:
The patent introduces an intermediary data transformation layer between the structured data source and the large language model. This intermediary component acts as a mediator that converts event data and tabular data into sequential token representations, enabling the model to process non-natural language inputs effectively. The intermediary transformation process includes techniques such as value mapping, sequence generation, and format conversion, which bridge the gap between structured data and the model's natural language processing capabilities.
2Adaptability or versatility
If large language models are used for processing structured data, then linguistic tasks can be performed, but the inner workings of the model remain unobservable and difficult to interpret
Solution Approach 1:
The patent applies feedback by implementing observation and interpretation mechanisms that provide insights into the large language model's processing of transformed data. The system observes how the model processes sequential tokens derived from structured data and provides feedback on the transformation effectiveness. This feedback loop enables users to understand and measure the model's inner workings, improving interpretability while maintaining the model's versatile task performance capability.
3Measurement precision
If data transformation techniques are applied to improve data suitability, then output accuracy improves, but the processing time and computational resources increase
Solution Approach 1:
The patent applies parameter changes by systematically transforming data parameters from structured formats into sequential token representations suitable for large language models. The transformation process modifies data parameters such as value encoding, sequence length, and tokenization strategies to optimize both accuracy and processing efficiency. By carefully managing these parameter transformations, the system achieves improved classification accuracy while minimizing the time and computational resources required for data preparation.
Data Source
AI summary
Aspects described herein may relate to techniques and/or methods that process certain forms of event data and/or tabular data for input to one or more machine learning models, such as a large language model. Additional aspects may relate to using the output of a large language model as part of a process for detecting fraud based on the event data and/or tabular data. In some variations, the event data and/or tabular data may be processed into data tokens, embeddings, or other forms of data suitable for use as input to a large language model.


