Adaptive Data Mining Query Expansion for Unindexed Sources
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current search engines are limited in analyzing multiple data points in real-time and are restricted to querying only indexed web sites, missing nearly seventy percent of web pages, including proprietary data silos and unindexed sources, which hampers real-time data analysis and automated data extraction.
Innovation Solution
A method and apparatus for generating and expanding queries based on topics of interest, executing them across various data sources, including both open-source and proprietary data, and monitoring selected data sources for matches, using a processor and memory unit to perform data mining and visualization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If typical search engines are used to query data, then the search process is simple and fast, but the coverage is limited to indexed web sites only, missing nearly seventy percent of web pages
Solution Approach 1:
The patent introduces web crawlers and data extraction tools as intermediary components that bridge the gap between search engines and unindexed data sources. These crawlers systematically navigate and extract data from proprietary data silos, web sites behind firewalls, and comment sections that traditional search engines cannot access, thereby expanding coverage without requiring direct modification of the search engine itself
Solution Approach 2:
The system segments the data collection process into multiple specialized components: web crawlers for navigating unindexed sites, data extraction modules for retrieving information from proprietary sources, and query expansion modules for generating comprehensive search terms. This segmentation allows each component to specialize in accessing specific types of data sources, collectively achieving broad coverage while maintaining manageable system complexity
2Loss of information
If real-time data analysis is performed across multiple data sources, then the completeness of information increases, but the processing time and computational resources increase significantly
Solution Approach 1:
The system performs preliminary actions by pre-crawling and indexing data from multiple sources before real-time analysis is needed. Data extraction tools continuously populate data repositories with information from proprietary sources and unindexed web pages in advance, so that when real-time analysis is required, the data is already available and structured for rapid processing
Solution Approach 2:
The patent implements continuous data extraction and monitoring processes that operate in the background, maintaining an up-to-date repository of data from multiple sources. This continuous action ensures that information is fresh and complete without requiring intensive batch processing at the time of analysis, thereby reducing processing time while maintaining information completeness
3Loss of information
If query expansion is performed to cover more search terms, then the comprehensiveness of data extraction improves, but the complexity of query processing increases
Solution Approach 1:
The query expansion module operates autonomously, automatically generating expanded search terms and queries based on the initial query without requiring manual intervention. The system self-manages the complexity of query processing by implementing automated algorithms that analyze the initial query, identify relevant concepts, and generate comprehensive search term variations, thereby improving data extraction completeness while containing processing complexity through automation
Data Source
AI summary
A method of analyzing data is presented. The method includes generating a query based on a topic of interest, expanding search terms of the query, executing the query on one or more data sources, monitoring a specific data source selected from the one or more data sources. The monitoring is performed to monitor for matches to the query.


