Adaptive Data Mining Query Expansion for Real-Time Multi-Source Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data mining technologies are limited in analyzing real-time data across multiple sources, particularly due to the inability of typical search engines to handle unstructured and proprietary data, leading to incomplete and non-real-time information extraction.
Innovation Solution
A method and apparatus for generating and expanding queries based on topics of interest, executing them across various data sources, including both open-source and proprietary data, and monitoring selected sources for matches, using a processor and memory unit to facilitate real-time data extraction and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If typical search engines are used to query data sources, then the search is limited to indexed structured data sources, but real-time analysis of multiple data sources including unstructured and proprietary data cannot be performed
Solution Approach 1:
The system employs a universal query execution framework that can operate across multiple data source types (structured, unstructured, proprietary, open-source) through a common interface. The query expansion module generates multiple query variations that can be adapted to different data source formats, enabling a single system to handle diverse data sources without losing information from any particular source type.
Solution Approach 2:
The patent introduces an intermediary query expansion and execution layer between the user and multiple data sources. This intermediary module expands the original query into multiple specialized queries tailored for different data source types, ensuring comprehensive information retrieval while maintaining a unified user interface. The intermediary handles the complexity of accessing proprietary and unstructured data sources that typical search engines cannot reach.
2Loss of information
If query execution is performed on multiple data sources, then comprehensive data coverage is achieved, but real-time monitoring and processing becomes computationally complex
Solution Approach 1:
The system segments the query execution process into distinct modules: query expansion, query execution, result aggregation, and monitoring. Each data source type is handled by specialized execution components that process queries independently. This segmentation allows comprehensive data coverage across multiple sources while managing system complexity through modular architecture, where each segment handles a specific aspect of the multi-source querying process.
3Loss of information
If search terms are expanded to cover more topics, then data extraction comprehensiveness improves, but query execution time and processing resources increase
Solution Approach 1:
The system performs preliminary query expansion before execution, generating all necessary query variations in advance. The expansion module pre-processes the original query to create multiple specialized queries for different data sources and contexts. This preliminary action ensures comprehensive data extraction coverage while optimizing execution time, as the expansion work is done beforehand rather than during the actual query execution and monitoring phases.
Data Source
AI summary
A method of analyzing data is presented. The method includes generating a query based on a topic of interest, expanding search terms of the query, executing the query on one or more data sources, monitoring a specific data source selected from the one or more data sources. The monitoring is performed to monitor for matches to the query.


