Adaptive Data Mining Query Expansion for Real-Time Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data mining technologies are limited in analyzing real-time data across multiple sources, particularly due to their inability to search unindexed web pages and proprietary data silos, which restricts the availability of information for real-time analysis.

Innovation Solution

A method and apparatus for generating and expanding queries based on topics of interest, executing them across various data sources, including both open-source and proprietary data, and monitoring selected data sources for matches, enabling real-time data extraction and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If typical search engines are used to query data, then the search is limited to exact search terms and indexed web sites, but this restricts the ability to analyze multiple data points in real time and access unindexed data sources

Engineering Contradiction:
Improvedata source coverageVSAvoidinformation availability
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system implements a universal query execution capability that can access multiple data source types (indexed websites, unindexed web pages, proprietary data silos, social media, web feeds) through a single unified interface, eliminating the limitation of search engines being restricted to only indexed content

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces specialized data extraction and monitoring tools as intermediaries that bridge the gap between query requirements and diverse data sources, enabling access to unindexed pages and proprietary silos that traditional search engines cannot reach

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If traditional data mining tools are used, then the analysis is limited to structured data sources, but this excludes seventy percent of web pages that are not indexed

Engineering Contradiction:
Improvedata source accessibilityVSAvoiddata volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system employs multi-functional data extraction capabilities that can handle both structured and unstructured data from diverse sources including unindexed web pages, social media platforms, and proprietary databases, expanding access to the 70% of web content traditionally unreachable

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system dynamically adapts its data extraction and monitoring approach based on the type of data source being accessed, adjusting its methodology to effectively retrieve information from unstructured sources like unindexed web pages while maintaining efficiency with structured databases

Inventive Principle:
Principle #15Dynamics

3Productivity

If real-time data analysis is implemented across multiple data sources, then comprehensive information extraction is achieved, but the system complexity increases

Engineering Contradiction:
Improvereal-time analysis capabilityVSAvoidsystem architecture
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system segments the complex task of multi-source real-time data analysis into specialized modular components: query generation modules, expansion modules, execution modules for different data source types, and monitoring modules, allowing each to handle specific aspects independently while working together as a unified system

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10114872B2Real-time and adaptive data mining
Publication Date: 2018.10.30 RULE 14 LLC
  • US10114872B2 patent drawing
  • US10114872B2 patent drawing
  • US10114872B2 patent drawing

AI summary

A method of analyzing data is presented. The method includes generating a query based on a topic of interest, expanding search terms of the query, executing the query on one or more data sources, monitoring a specific data source selected from the one or more data sources. The monitoring is performed to monitor for matches to the query.