Self-learning crawler with rule-based data mining for information extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for automatic information extraction from multi-dimensional data sets are inefficient due to high time and processor complexity, and lack accuracy, especially when dealing with varying documentation standards and the need for periodic updates of large repositories.
Innovation Solution
A self-learning based crawling and rule-based data mining system that determines and adapts to documentation standards, performing assisted and recursive crawling to extract meaningful information, and applies classification rules and decision tree data mining for accurate information retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If independent crawling and data mining processes are used, then implementation simplicity is maintained, but time complexity and processor complexity increase significantly
Solution Approach 1:
The patent combines crawling and data mining into a single integrated process where data is crawled and mined simultaneously rather than through separate independent processes. This merging reduces the overall time complexity and processor complexity while maintaining implementation simplicity, directly addressing the technical contradiction presented in the background.
2Measurement precision
If manual updating of large repositories is performed, then data accuracy can be verified, but time consumption and error proneness increase
Solution Approach 1:
The system performs automated information extraction and repository updating through self-service mechanisms, eliminating the need for manual verification and updating. The automated process maintains data accuracy while significantly improving update efficiency and reducing errors, directly resolving the contradiction between accuracy verification and update productivity.
3Speed
If traditional data mining techniques are used, then processing speed is maintained, but accuracy in extracting meaningful information decreases due to varying documentation standards
Solution Approach 1:
The patent introduces dynamic adaptation mechanisms that allow the data mining process to adjust to varying documentation standards in real-time. The system learns and adapts to different standards while maintaining processing speed, resolving the contradiction between speed and extraction accuracy by making the mining process flexible rather than static.
4Loss of information
If comprehensive crawling is performed on all data sources, then information completeness is improved, but resource consumption and complexity increase
Solution Approach 1:
The system performs preliminary actions by pre-processing and filtering data during the crawling phase before applying data mining techniques. This preliminary filtering reduces the data volume that requires complex processing while ensuring information completeness is maintained, thereby reducing overall system complexity while preserving comprehensive information extraction.
Data Source
AI summary
Methods and Systems for automatic information extraction by performing self-learning crawling and rule-based data mining is provided. The method determines existence of crawl policy within input information and performs at least one of front-end crawling, assisted crawling and recursive crawling. Downloaded data set is pre-processed to remove noisy data and subjected to classification rules and decision tree based data mining to extract meaningful information. Performing crawling techniques leads to smaller relevant datasets pertaining to a specific domain from multi-dimensional datasets available in online and offline sources.


