Metadata-Enriched Text Chunk Routing for Accurate Data Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data analysis platforms are inefficient and inaccurate in ingesting, analyzing, and processing large datasets, particularly those with unlabeled data, leading to inaccurate or spurious data generation, which can result in ill-informed decisions.

Innovation Solution

The system processes raw data into text chunks, classifies them, augments with metadata, and distributes them to specialized data stores using machine learning models, embedding and sequencing to improve data management and retrieval.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional systems process large datasets, then data processing capacity is maintained, but data analysis accuracy deteriorates leading to hallucinations and spurious data

Engineering Contradiction:
Improvedata analysis accuracyVSAvoiddata processing capacity
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent divides large datasets into smaller text chunks and processes them through specialized machine learning models. Each chunk is handled individually through embedding, classification, and routing to appropriate data stores, which improves processing accuracy while maintaining overall system capacity through parallel processing of multiple chunks.

Inventive Principle:
Principle #1Segmentation

2Productivity

If conventional systems ingest diverse data formats, then data coverage is maintained, but data comprehension efficiency deteriorates

Engineering Contradiction:
Improvedata comprehension efficiencyVSAvoiddata format compatibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary processing layer that converts diverse data formats into a standardized text chunk format. This intermediary step includes text extraction, cleaning, and normalization processes that transform various source formats into a uniform structure suitable for efficient machine learning model processing, thereby maintaining format versatility while improving comprehension efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If conventional systems retrieve data from large databases, then data availability is maintained, but retrieval accuracy deteriorates due to information overload

Engineering Contradiction:
Improveretrieval accuracyVSAvoiddata volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments large databases into multiple specialized data stores, each containing specific types of processed text chunks. When a query is received, the system routes it to the appropriate specialized data store based on the query type, thereby reducing the search space from the entire database to a relevant subset and significantly improving retrieval accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the traditional single-database structure into a multi-dimensional organized system with specialized data stores arranged by data type and classification. This dimensional reorganization allows the system to navigate and retrieve information more efficiently by moving from a flat, monolithic structure to a hierarchical, categorized structure.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250245259A1Systems and methods for text data processing and chunk distribution
Publication Date: 2025.07.31 LLAMALAB INC
  • US20250245259A1 patent drawing
  • US20250245259A1 patent drawing
  • US20250245259A1 patent drawing

AI summary

Disclosed herein are systems, methods, and media for processing and distributing text data from a dataset. The disclosed embodiments include receiving raw data and converting the raw data into a set of text chunks. The disclosed embodiments include determining a set of classifications for the raw data. The disclosed embodiments include. augmenting a text chunk in the set of text chunks with metadata. The disclosed embodiments include generating a windowed chunk by appending context to the augmented chunk. The disclosed embodiments include embedding the windowed chunk and the extracted retrieval metadata into a chunk embedding. The disclosed embodiments include distributing the chunk embedding by determining a data store in the set of data stores corresponding to the assigned classification and assigning the chunk embedding to the determined data store.