Machine Learning Data Extraction for Document Information Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users of documents, such as news articles, face limitations in accessing comprehensive data related to the subject matter, as relevant information is often scattered throughout the text and not easily retrievable.

Innovation Solution

A client-server-based network architecture that utilizes machine learning models to parse and extract relevant data from source documents, providing it to users in an automated and unified manner through a graphical user interface, integrating with data repositories to offer graphical representations and numerical data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If readers access information only from the source document text, then information completeness is maintained within the document boundaries, but information accessibility and comprehensiveness deteriorate because relevant data is scattered and hard to retrieve

Engineering Contradiction:
Improveinformation completenessVSAvoidinformation accessibility
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The patent introduces an intermediary system comprising a processing module and data repository that mediates between the source document and the reader. The processing module extracts entities and concepts from the document, queries the data repository for related information, and presents consolidated results to the reader, thereby improving accessibility without losing information completeness

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system segments the information retrieval process into distinct components: document parsing, entity extraction, data repository querying, and result presentation. This segmentation allows each component to specialize in specific tasks, improving overall information accessibility while maintaining completeness

Inventive Principle:
Principle #1Segmentation

2Reliability

If factual information is distributed throughout the article text, then information context is preserved, but information retrieval efficiency deteriorates because readers must manually search through the entire document

Engineering Contradiction:
Improveinformation contextVSAvoidinformation retrieval efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The processing module extracts key entities, concepts, and factual information from the distributed text into a structured format. By taking out this information from its scattered locations in the source document and organizing it systematically, the system enables efficient retrieval while preserving the contextual relationships through the data repository structure

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system transforms the one-dimensional linear text structure into a multi-dimensional information space by creating associations between entities, concepts, and related data in the repository. This dimensional transformation allows readers to access information from multiple entry points and perspectives, improving retrieval efficiency while maintaining contextual integrity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Loss of information

If the system integrates external data repositories to provide comprehensive related data, then information comprehensiveness is improved, but system complexity deteriorates due to integration requirements

Engineering Contradiction:
Improveinformation comprehensivenessVSAvoidsystem integration complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The data repository is designed as a universal multi-functional component that can store and provide various types of related information (statistical data, contextual information, supplementary facts) for different domains and topics. This universality reduces system complexity by using a single integrated repository rather than multiple specialized databases

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Measurement precision

If the system automatically extracts and presents related data from source documents, then information retrieval accuracy is improved, but processing time increases due to automated analysis requirements

Engineering Contradiction:
Improveinformation retrieval accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-processing and indexing the source document content before user queries are submitted. Entities, concepts, and key information are extracted and stored in an optimized format in advance, so that when users seek information, the system can quickly retrieve and present accurate results without performing extensive real-time analysis

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10108907B2Method and system to provide related data
Publication Date: 2018.10.23 KNOEMA CORP
  • US10108907B2 patent drawing
  • US10108907B2 patent drawing
  • US10108907B2 patent drawing

AI summary

Methods and systems of providing related information to a source document are described. The method may include accessing the source document displayed to a user in a graphical user interface (GUI) of a client device. The source document includes numerical data and text. Discovered data corresponding to the numerical data included in the source document is then identified. Further, a database trained with a machine-learning algorithm to identify time series data related data associated with the text is accessed. The discovered data with a discovered data identifier and the time series related data is then displayed in the GUI. In example embodiments, the methods and systems described herein interact with applications such as spreadsheets applications, email clients, word processing applications, webpages and the like.