Web Content Crawler Using Signature Files for Automated Updates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for updating web-content in data repositories are time-consuming, inefficient, prone to inaccuracies, and economically costly due to reliance on manual intervention, requiring continuous database upgrades to store new HTML documents.
Innovation Solution
A system that acquires a web-page signature file, analyzes it to identify modifications, compares the content, uses a machine learning algorithm to determine the importance of changes, and crawls the repository based on predefined thresholds to update the content automatically.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual intervention is used to update web-content in data repository, then accuracy of content update can be controlled, but time consumption increases and productivity decreases
Solution Approach 1:
The system performs automated crawling and comparison of web-content with the data repository without requiring manual intervention. The automated system independently detects changes, compares content, and updates the repository, eliminating the need for human operators to manually check and update content while maintaining accuracy through systematic automated comparison processes
Solution Approach 2:
The patent replaces manual mechanical operations with automated computational processes. Instead of human operators manually comparing HTML documents and updating databases, the system uses automated software that performs web crawling, content comparison, and database updates through computational algorithms, thereby increasing productivity while maintaining controlled accuracy
2Measurement precision
If manual comparison of HTML documents is performed, then update accuracy can be maintained, but computation time and labor requirements increase
Solution Approach 1:
The system continuously and automatically crawls web-content and compares it with the data repository in real-time or near-real-time. This continuous automated operation eliminates the intermittent manual comparison process, maintaining constant readiness for updates without the time loss associated with manual intervention cycles
Solution Approach 2:
The patent substitutes manual document comparison with automated computational comparison systems. The automated system processes HTML documents, detects changes, and updates the repository through computational algorithms, eliminating the time-consuming manual comparison process while maintaining systematic accuracy
3Device complexity
If conventional manual update technique is used, then system complexity remains low, but reliability and accuracy of updates deteriorate due to human error
Solution Approach 1:
The automated system performs self-service operations for web-content crawling, comparison, and updates without human intervention. This self-service mechanism eliminates human error and improves reliability while the systematic automated process manages complexity through standardized algorithms and procedures
Solution Approach 2:
The system incorporates feedback mechanisms where the automated crawler continuously monitors web-content changes, compares them with the data repository, and updates the repository based on detected modifications. This feedback loop ensures reliable and accurate updates by systematically detecting and responding to content changes without human error
4Loss of information
If continuous database upgrades are performed to store new HTML documents, then data freshness is maintained, but economic costs increase due to large database requirements
Solution Approach 1:
The system extracts only the necessary web-content information and updates the data repository selectively based on detected changes and predefined criteria. Instead of continuously storing all HTML documents, the system extracts and stores only relevant updated content, reducing storage requirements while maintaining data freshness
Solution Approach 2:
The patent implements parameter changes in the update process by using predefined criteria and thresholds for determining which content to update. The system changes the approach from storing all content to storing only content that meets specific criteria, thereby reducing database storage volume while maintaining freshness of relevant information
Data Source
AI summary
Disclosed is a system comprising: a data repository storing web-content; a data processing arrangement communicatively coupled to data repository, wherein data processing arrangement is configured to: acquire a web-page signature file associated to web-content, from a web-server hosting a website for displaying web-content, wherein web-page signature file includes a plurality of data related to web-content; analyse plurality of data included in web-page signature file to identify a modification in website; compare web-content stored in data repository with web-content displayed on website to determine additional web-content included in web-content displayed on website; use a machine learning algorithm to determine an importance value for additional web-content using a set of predefined parameters; crawl web-content stored in data repository based on additional web-content upon determining importance value to be greater than a predefined threshold value; and predict a time for crawling web-content using forecast module.

