Unstructured Data Extraction via Historical Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual entry of data into structured data collections is inflexible, inconvenient, and time-consuming, and existing methods struggle to efficiently extract relevant information from unstructured data due to ambiguity.

Innovation Solution

A system and method that receives unstructured data, identifies attributes by using historical data, additional user data, and external information to classify and resolve ambiguities, allowing for efficient conversion and storage of unstructured data into structured data collections, including natural language processing and optical character recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual entry of data is used, then data can be entered into structured data collection, but it is inflexible, inconvenient, and time consuming

Engineering Contradiction:
Improveconvenience of data inputVSAvoidtime consuming
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces the mechanical manual keyboard entry system with an automated speech recognition and text processing system. The computing system captures speech inputs, converts them to text, and automatically extracts and structures data without requiring manual typing, thereby improving convenience and reducing time consumption.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service data entry where users simply speak their information naturally without needing to know the structure of the data collection. The computing system automatically processes the unstructured speech input, identifies relevant data elements, and populates the structured format without user intervention beyond the initial speech input.

Inventive Principle:
Principle #25Self-service

2Productivity

If unstructured data is processed to extract terms, then data extraction efficiency improves, but ambiguity in term identification occurs

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoidaccuracy of term identification
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system uses feedback loops where historical data and previously identified terms are incorporated into the processing of new unstructured data. The computing system learns from past extractions and uses this feedback to improve the accuracy of term identification in current processing, reducing ambiguity while maintaining high efficiency.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary actions by pre-processing the unstructured data to identify potential terms and their contexts before final classification. The system prepares candidate terms and uses historical data to pre-evaluate their relevance, which streamlines the extraction process and improves accuracy by reducing ambiguous classifications.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If historical data and additional user data are used to identify terms, then term identification accuracy increases, but system complexity increases

Engineering Contradiction:
Improveaccuracy of term identificationVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The computing system performs multiple functions using a unified processing framework: it processes speech to text conversion, performs optical character recognition, extracts terms from unstructured data, compares with historical data, and structures the output. This multi-functional approach increases accuracy without proportionally increasing system complexity by consolidating functions into a single versatile system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Ease of operation

If predefined attributes are provided for selection, then flexibility of data collection changes, but ease of operation improves

Engineering Contradiction:
Improveease of attribute selectionVSAvoidflexibility of data collection
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system dynamically adjusts between predefined attributes and custom attribute creation based on the processing needs. When processing unstructured data, the system can select from predefined attributes for common data types, but also allows for dynamic creation of custom attributes when novel data elements are identified, maintaining both ease of operation and adaptability.

Inventive Principle:
Principle #15Dynamics

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach provides a more convenient and efficient user interface for data input, increases the accuracy of term identification, and allows for dynamic attribute selection and improvement of data extraction rules, enhancing the flexibility and accuracy of data collection.

Implementation Method 1

converting the natural language utterance into the text

Methodology Applied
Scientific EffectNatural language processing:

Implementation Method 2

performing optical character recognition on the image to identify the text

Methodology Applied
Scientific EffectOptical character recognition:

Data Source

PatentUS9299041B2Obtaining data from unstructured data for a structured data collection
Publication Date: 2016.03.29 BUSINESS OBJECTS SOFTWARE
  • US9299041B2 patent drawing
  • US9299041B2 patent drawing
  • US9299041B2 patent drawing

AI summary

Techniques for obtaining data from unstructured data for a structured data collection include receiving unstructured data that includes text; identifying an attribute associated with a structured data collection; obtaining at least one of historical data associated with the attribute or additional data associated with a user of the computing system; identifying one or more terms from the unstructured data as being associated with the attribute based on at least one of the historical data or the additional data; and storing the identified one or more terms in a data record of the unstructured data collection.