Automated Data Type Interpretation for Structured Files

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual identification of data types in large business data files is error-prone, time-consuming, and costly due to the lack of standardized formats and reliance on human interpretation, which becomes inefficient as data volumes grow.

Innovation Solution

An automated method using a rich contextual framework with multiple subsystems to identify and validate data types across various perspectives, leveraging oracle subsystems to interpret data patterns and relationships, reducing reliance on provided layouts and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manual identification of data types is performed by human reviewers, then flexibility in interpreting data context is improved, but accuracy and consistency deteriorate due to human error and fatigue

Engineering Contradiction:
Improveflexibility in interpreting data contextVSAvoidaccuracy of data type identification
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent replaces the manual mechanical process of human reviewers examining and labeling data fields with an automated computational system. The system uses machine learning models and algorithms to automatically identify data types, patterns, and relationships across millions of records, eliminating human error while maintaining adaptability through programmable logic and contextual analysis capabilities.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables data files to be self-analyzed through automated pattern recognition and validation. The computational system independently examines data patterns, infers data types, and validates consistency across records without requiring human intervention, allowing the data itself to reveal its structure through systematic analysis.

Inventive Principle:
Principle #25Self-service

2Reliability

If manual data type identification is performed, then contextual understanding can be applied, but processing time increases significantly for large data volumes

Engineering Contradiction:
Improvecontextual understanding of dataVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the analysis process into distinct computational stages: pattern recognition phase, data type inference phase, validation phase, and consistency checking phase. Each stage processes specific aspects of the data independently and in parallel where possible, enabling comprehensive contextual understanding to be achieved through systematic breakdown rather than sequential manual review.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis by examining data patterns, formats, and relationships across sample records before finalizing data type assignments. Pre-processing steps include identifying delimiters, detecting fixed-width boundaries, and analyzing value distributions to establish contextual frameworks that guide subsequent automated classification, enabling fast processing without sacrificing understanding.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated methods are used for data type identification, then processing speed is improved, but accuracy deteriorates without rich contextual analysis

Engineering Contradiction:
Improveprocessing speedVSAvoidaccuracy of data type identification
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent implements feedback loops where the system continuously validates its data type assignments by checking consistency across multiple records and against inferred patterns. The automated system adjusts its classifications based on validation results, cross-references data against known patterns and relationships, and iteratively refines its accuracy while maintaining high processing speed through efficient algorithmic operations.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system moves beyond simple single-field analysis to multi-dimensional contextual analysis, examining relationships between adjacent fields, patterns across record sequences, and hierarchical structures within data. By analyzing data from multiple dimensions simultaneously (field-level, record-level, file-level patterns), the automated system achieves high accuracy that rivals or exceeds manual review while maintaining computational speed.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Loss of time

If provided layout information is used, then processing time is reduced, but accuracy deteriorates when layout is inaccurate or incomplete

Engineering Contradiction:
Improveprocessing timeVSAvoidaccuracy of layout interpretation
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary validation layer between the provided layout information and the final data type assignments. The system uses the provided layout as a starting hypothesis but automatically validates and corrects it by analyzing actual data patterns, field relationships, and consistency across records. This intermediary analysis step ensures that even when provided layouts are inaccurate or incomplete, the final interpretation achieves high accuracy while maintaining efficient processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10838919B2Automated interpretation for the layout of structured multi-field files
Publication Date: 2020.11.17 ACXIOM LLC
  • US10838919B2 patent drawing
  • US10838919B2 patent drawing
  • US10838919B2 patent drawing

AI summary

An entirely automated system for the interpretation of the field layout for multi-field files uses a rich contextual framework constructed by the interaction of three subsystems to provide a holistic view of the contexts of a structured data file as defined by the location and data type of each field. The roles of each of the subsystems are (1) the determination of the file's metadata and positions of the different data fields; (2) the use of fallible oracles (i.e., no oracle must be capable of identifying the type for every record) to provide a set of interpretations of the fields at several levels; and (3) the accurate determination of the location and specific data type for each field without the necessity to interpret every record correctly, even in the presence of ambiguity of data. The system may operate on both delimited and fixed-width structure files.