Unstructured Data Extraction via Historical Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual entry of data into structured data collections is inflexible, inconvenient, and time-consuming, and existing methods struggle to efficiently extract relevant information from unstructured data due to ambiguity.
Innovation Solution
A system and method that receives unstructured data, identifies attributes by using historical data, additional user data, and external information to classify and resolve ambiguities, allowing for efficient conversion and storage of unstructured data into structured data collections, including natural language processing and optical character recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual entry of data is used, then data can be entered into structured data collection, but it is inflexible, inconvenient, and time consuming
Solution Approach 1:
The patent replaces the mechanical manual keyboard entry system with an automated speech recognition and text processing system. The computing system captures speech inputs, converts them to text, and automatically extracts and structures data without requiring manual typing, thereby improving convenience and reducing time consumption.
Solution Approach 2:
The system enables self-service data entry where users simply speak their information naturally without needing to know the structure of the data collection. The computing system automatically processes the unstructured speech input, identifies relevant data elements, and populates the structured format without user intervention beyond the initial speech input.
2Productivity
If unstructured data is processed to extract terms, then data extraction efficiency improves, but ambiguity in term identification occurs
Solution Approach 1:
The system uses feedback loops where historical data and previously identified terms are incorporated into the processing of new unstructured data. The computing system learns from past extractions and uses this feedback to improve the accuracy of term identification in current processing, reducing ambiguity while maintaining high efficiency.
Solution Approach 2:
The patent performs preliminary actions by pre-processing the unstructured data to identify potential terms and their contexts before final classification. The system prepares candidate terms and uses historical data to pre-evaluate their relevance, which streamlines the extraction process and improves accuracy by reducing ambiguous classifications.
3Measurement precision
If historical data and additional user data are used to identify terms, then term identification accuracy increases, but system complexity increases
Solution Approach 1:
The computing system performs multiple functions using a unified processing framework: it processes speech to text conversion, performs optical character recognition, extracts terms from unstructured data, compares with historical data, and structures the output. This multi-functional approach increases accuracy without proportionally increasing system complexity by consolidating functions into a single versatile system.
4Ease of operation
If predefined attributes are provided for selection, then flexibility of data collection changes, but ease of operation improves
Solution Approach 1:
The system dynamically adjusts between predefined attributes and custom attribute creation based on the processing needs. When processing unstructured data, the system can select from predefined attributes for common data types, but also allows for dynamic creation of custom attributes when novel data elements are identified, maintaining both ease of operation and adaptability.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach provides a more convenient and efficient user interface for data input, increases the accuracy of term identification, and allows for dynamic attribute selection and improvement of data extraction rules, enhancing the flexibility and accuracy of data collection.
Implementation Method 1
converting the natural language utterance into the text
Implementation Method 2
performing optical character recognition on the image to identify the text
Data Source
AI summary
Techniques for obtaining data from unstructured data for a structured data collection include receiving unstructured data that includes text; identifying an attribute associated with a structured data collection; obtaining at least one of historical data associated with the attribute or additional data associated with a user of the computing system; identifying one or more terms from the unstructured data as being associated with the attribute based on at least one of the historical data or the additional data; and storing the identified one or more terms in a data record of the unstructured data collection.


