Unstructured Data Extraction and Correlation System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems are limited in processing and leveraging public data, particularly unstructured data, offering little more than typical search engine functionality, failing to provide users with meaningful information that ties together multiple data sources effectively.

Innovation Solution

A networked system with a data extractor and correlator that retrieves public data from various sources, extracts information from unstructured data, and correlates it with structured data to generate enriched knowledge, which can be stored and used in subsequent queries, along with a user profile module and feedback mechanism to adapt to user needs and preferences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If typical search engine functionality is used to process public data, then the system is simple and easy to operate, but the ability to provide meaningful information and leverage unstructured data is limited

Engineering Contradiction:
Improveability to process and leverage unstructured dataVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the data processing task into distinct modules: a data extractor module that retrieves unstructured data from public sources, a processor module that analyzes and structures the extracted information, and a integration module that combines it with existing structured data. This segmentation enables sophisticated unstructured data handling while maintaining manageable system complexity through modular design

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing layer between raw public data and the user. This intermediary includes natural language processing components, entity recognition systems, and data correlation engines that transform unstructured data into meaningful structured information, bridging the gap between simple search functionality and intelligent data leverage

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If more public data is retrieved and processed, then the quantity of information available to users increases, but the time and computational resources required increase

Engineering Contradiction:
Improvequantity of informationVSAvoidprocessing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system performs preliminary actions by pre-fetching and pre-processing public data in advance. The data extractor continuously retrieves relevant unstructured data from public sources and stores it in a processed state, so that when users query, the information is already prepared and ready for rapid delivery, reducing real-time processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by selectively processing only the most relevant portions of retrieved data based on query context and user preferences. Rather than processing all retrieved data equally, the system identifies and processes only the critical subsets needed to answer specific questions, reducing overall processing time while maintaining information quality

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9858332B1Extracting and leveraging knowledge from unstructured data
Publication Date: 2018.01.02 SRI INTERNATIONAL
  • US9858332B1 patent drawing
  • US9858332B1 patent drawing
  • US9858332B1 patent drawing

AI summary

A system may include a machine-implemented data extractor and correlator configured to retrieve data from at least one data source. The data extractor and correlator may extract information from unstructured data within the retrieved data and correlate the extracted information with previously stored structured data to generate additional structured data. The system may also include a storage device configured to store the previously stored structured data and the additional structured data.