Contextual Data Extraction System for Multimodal Richly Formatted Sources

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Richly formatted data poses challenges for machine learning as it often requires preserving context, which is difficult due to its multimodal nature, including images, audio, and video, and existing systems struggle to effectively process and extract information from such data while maintaining its inherent structure and formatting.

Innovation Solution

A system and method for automated scalable contextual data collection and extraction that accesses richly formatted data sources, processes information across different modalities, and transforms the data into graph and time series-based datasets for enriched knowledge base construction and data loss prevention monitoring.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If existing systems process richly formatted data, then data extraction speed improves, but context preservation deteriorates

Engineering Contradiction:
Improvedata extraction speedVSAvoidcontext preservation
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system segments richly formatted data into distinct modalities (text, images, audio, video) and processes each through specialized extraction engines that preserve modality-specific context. This segmentation allows parallel processing of different data types while maintaining their inherent structure and relationships, resolving the contradiction between extraction speed and context preservation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transforms extracted data from traditional tabular formats into graph-based representations that add dimensional context. By representing data as nodes and relationships as edges in a graph structure, the system preserves contextual relationships that would be lost in conventional processing, enabling both high-speed extraction and context retention through dimensional transformation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If the system processes multiple modalities, then data enrichment improves, but system complexity increases

Engineering Contradiction:
Improvemultimodal processing capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system employs a universal graph-based data structure that can represent multiple modalities (text, images, audio, video) within a single unified framework. This universal representation approach allows the same processing pipeline to handle diverse data types, reducing system complexity while maintaining multimodal versatility through a single multi-functional architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces graph-based data structures as an intermediary layer between raw multimodal inputs and processing engines. This intermediary representation standardizes diverse modalities into a common format that preserves contextual relationships, simplifying the architecture by providing a universal interface that mediates between different data types and processing functions.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If the system extracts detailed information, then data accuracy improves, but processing time increases

Engineering Contradiction:
Improvedata extraction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary organization of richly formatted data into graph structures before detailed extraction occurs. By pre-organizing data with contextual relationships already established in graph format, the system eliminates redundant processing steps during extraction, achieving both high accuracy and reduced processing time through advance preparation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates graph-based copies of the original richly formatted data that preserve contextual relationships. These graph copies serve as optimized representations that can be processed quickly while maintaining extraction accuracy, as the contextual structure is already encoded in the graph representation rather than requiring repeated analysis of original formats.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11314764B2Automated scalable contextual data collection and extraction system
Publication Date: 2022.04.26 QOMPLX INC
  • US11314764B2 patent drawing
  • US11314764B2 patent drawing
  • US11314764B2 patent drawing

AI summary

A system for contextual data collection and extraction is provided, comprising an extraction engine configured to receive context from a user for desired information to extract, connect to a data source providing a richly formatted dataset, retrieve the richly formatted dataset, process the richly formatted dataset and extract information from a plurality of linguistic modalities within the richly formatted, and transform the extracted data into a extracted dataset; and a knowledge base construction service configured to retrieve the extracted dataset, create a knowledge base for storing the extracted dataset, and store the knowledge base in a data store.