Synchronizing Document Crawling Schedules via Text Mining Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In environments with multiple data sources, periodical document crawling often results in time lags and inconsistent analysis due to differing schedules, which can lead to discrepancies in metadata-driven document classification and security protection.
Innovation Solution
A computer-implemented method and system that uses text mining metadata to synchronize crawling schedules across data sources, ensuring that documents with similar metadata are crawled simultaneously, thereby maintaining consistent analysis and security protection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If different crawling schedules are used for documents in different data sources, then each data source can be crawled independently with its own timing, but time lags occur in crawling documents leading to inconsistent analysis
Solution Approach 1:
The patent merges the crawling schedules of multiple data sources by synchronizing them to a common schedule. The system determines that when added metadata in internal documents or metadata in original documents necessitates the same crawling schedule, it changes respective crawling schedules of at least two data sources to the same crawling schedule, thereby eliminating time lags and ensuring consistent analysis across all data sources.
2Reliability
If the same crawling schedule is enforced across all data sources, then consistent analysis is maintained, but flexibility in optimizing individual data source crawling is reduced
Solution Approach 1:
The patent implements a dynamic approach where the crawling schedule is not rigidly fixed but can be adjusted based on metadata analysis. The system dynamically determines whether to synchronize schedules by evaluating whether added metadata or original metadata necessitates the same crawling schedule, allowing flexibility to maintain different schedules when not required and synchronize when consistency is needed.
3Reliability
If manual coordination of crawling schedules is performed, then some consistency can be achieved, but the complexity of schedule management increases
Solution Approach 1:
The patent implements an automated self-service system that eliminates manual schedule coordination. The system automatically determines whether synchronization is needed by analyzing metadata, automatically changes crawling schedules to the same schedule when necessary, and manages the coordination without human intervention, thereby reducing management complexity while maintaining consistency.
4Device complexity
If crawling schedules are not synchronized, then system simplicity is maintained, but metadata-driven document classification and security protection become inconsistent
Solution Approach 1:
The patent implements a feedback mechanism where the system continuously monitors metadata from crawled documents and uses this information to determine whether schedule synchronization is needed. The system analyzes added metadata in internal documents or metadata in original documents, and based on this feedback, automatically adjusts crawling schedules to ensure consistent metadata-driven classification and security protection across all data sources.
Data Source
AI summary
A computer-implemented method, a computer program product, and a computer system for coordinating schedules of crawling documents based on metadata added to documents by text mining. A computer system determines whether added metadata in internal documents or metadata in original documents necessitates that the original documents in at least two of respective data sources be crawled by an application with a same crawling schedule. A computer system changes respective crawling schedules of at least two of the respective data sources to the same crawling schedule, in response to determining that the same crawling schedule is needed. A computer system crawls the original documents in at least two of the respective data sources, according to the same crawling schedule.


