Adaptive Data Mining System for Real-Time Unindexed Source Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data mining technologies are limited in analyzing real-time data across multiple sources, particularly failing to extract relevant information from unindexed web pages and proprietary data silos, which restricts their ability to provide comprehensive and timely insights.

Innovation Solution

A method and system for generating and expanding queries based on topics of interest, executing them across various data sources, including both open-source and proprietary data, and monitoring selected sources for matches, using processors and memory units to facilitate real-time data extraction and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If typical search engines are used for data analysis, then query execution is simple and fast, but the ability to analyze multiple data points in real time and access unindexed web pages is limited

Engineering Contradiction:
Improvereal-time data analysis capabilityVSAvoidaccess to unindexed web pages and proprietary data silos
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a system that performs multiple functions: it executes traditional search engine queries while simultaneously accessing unindexed web pages and proprietary data silos through specialized data extraction tools. This multi-functional approach allows the system to overcome the limitations of typical search engines by integrating diverse data access capabilities into a unified platform that can analyze data from multiple sources in real time.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Loss of information

If search engines are limited to indexed web sites and exact search terms, then system complexity is low, but the comprehensiveness of information obtained is limited

Engineering Contradiction:
Improveinformation completeness from unindexed sourcesVSAvoiddata extraction and monitoring system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent introduces intermediary components including specialized data extraction tools, query expansion modules, and monitoring agents that act as mediators between the search system and unindexed data sources. These intermediaries handle the complexity of accessing and extracting data from proprietary data silos and unindexed web pages, shielding the core system from direct complexity while enabling comprehensive information retrieval from previously inaccessible sources.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of time

If real-time monitoring of multiple data sources is implemented, then data freshness and timeliness improve, but system resource consumption and complexity increase

Engineering Contradiction:
Improvedata freshness and timelinessVSAvoidmonitoring system complexity
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by pre-configuring monitoring agents and data extraction tools to be ready before data needs to be retrieved. The system establishes continuous monitoring connections to multiple data sources in advance, pre-processes data where possible, and maintains ready-state extraction capabilities. This allows the system to provide real-time data freshness without the complexity of ad-hoc data retrieval, as the monitoring infrastructure is already in place and actively collecting data from multiple sources.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10108678B2Real-time and adaptive data mining
Publication Date: 2018.10.23 RULE 14 LLC
  • US10108678B2 patent drawing
  • US10108678B2 patent drawing
  • US10108678B2 patent drawing

AI summary

A method of analyzing data is presented. The method includes generating a query based on a topic of interest, expanding search terms of the query, executing the query on one or more data sources, monitoring a specific data source selected from the one or more data sources. The monitoring is performed to monitor for matches to the query.