Synchronizing Document Crawling Schedules via Text Mining Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In environments with multiple data sources, periodical document crawling often results in time lags and inconsistent analysis due to differing schedules, which can lead to discrepancies in metadata-driven document classification and security protection.

Innovation Solution

A computer-implemented method and system that uses text mining metadata to synchronize crawling schedules across data sources, ensuring that documents with similar metadata are crawled simultaneously, thereby maintaining consistent analysis and security protection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If different crawling schedules are used for documents in different data sources, then each data source can be crawled independently with its own timing, but time lags occur in crawling documents leading to inconsistent analysis

Engineering Contradiction:
ImproveIndependent crawling schedule for each data sourceVSAvoidConsistency of analysis across data sources
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent merges the crawling schedules of multiple data sources by synchronizing them to a common schedule. The system determines that when added metadata in internal documents or metadata in original documents necessitates the same crawling schedule, it changes respective crawling schedules of at least two data sources to the same crawling schedule, thereby eliminating time lags and ensuring consistent analysis across all data sources.

Inventive Principle:
Principle #5Merging (Combining)

2Reliability

If the same crawling schedule is enforced across all data sources, then consistent analysis is maintained, but flexibility in optimizing individual data source crawling is reduced

Engineering Contradiction:
ImproveConsistency of analysis across data sourcesVSAvoidFlexible crawling schedule for each data source
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic approach where the crawling schedule is not rigidly fixed but can be adjusted based on metadata analysis. The system dynamically determines whether to synchronize schedules by evaluating whether added metadata or original metadata necessitates the same crawling schedule, allowing flexibility to maintain different schedules when not required and synchronize when consistency is needed.

Inventive Principle:
Principle #15Dynamics

3Reliability

If manual coordination of crawling schedules is performed, then some consistency can be achieved, but the complexity of schedule management increases

Engineering Contradiction:
ImproveConsistency of crawling timingVSAvoidComplexity of schedule coordination
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements an automated self-service system that eliminates manual schedule coordination. The system automatically determines whether synchronization is needed by analyzing metadata, automatically changes crawling schedules to the same schedule when necessary, and manages the coordination without human intervention, thereby reducing management complexity while maintaining consistency.

Inventive Principle:
Principle #25Self-service

4Device complexity

If crawling schedules are not synchronized, then system simplicity is maintained, but metadata-driven document classification and security protection become inconsistent

Engineering Contradiction:
ImproveSimplicity of crawling schedule managementVSAvoidPrecision of metadata-driven classification
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent implements a feedback mechanism where the system continuously monitors metadata from crawled documents and uses this information to determine whether schedule synchronization is needed. The system analyzes added metadata in internal documents or metadata in original documents, and based on this feedback, automatically adjusts crawling schedules to ensure consistent metadata-driven classification and security protection across all data sources.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20230252065A1Coordinating schedules of crawling documents based on metadata added to the documents by text mining
Publication Date: 2023.08.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20230252065A1 patent drawing
  • US20230252065A1 patent drawing
  • US20230252065A1 patent drawing

AI summary

A computer-implemented method, a computer program product, and a computer system for coordinating schedules of crawling documents based on metadata added to documents by text mining. A computer system determines whether added metadata in internal documents or metadata in original documents necessitates that the original documents in at least two of respective data sources be crawled by an application with a same crawling schedule. A computer system changes respective crawling schedules of at least two of the respective data sources to the same crawling schedule, in response to determining that the same crawling schedule is needed. A computer system crawls the original documents in at least two of the respective data sources, according to the same crawling schedule.