Web Content Crawler Using Signature Files for Automated Updates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for updating web-content in data repositories are time-consuming, inefficient, prone to inaccuracies, and economically costly due to reliance on manual intervention, requiring continuous database upgrades to store new HTML documents.

Innovation Solution

A system that acquires a web-page signature file, analyzes it to identify modifications, compares the content, uses a machine learning algorithm to determine the importance of changes, and crawls the repository based on predefined thresholds to update the content automatically.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual intervention is used to update web-content in data repository, then accuracy of content update can be controlled, but time consumption increases and productivity decreases

Engineering Contradiction:
Improveaccuracy of content updateVSAvoidspeed of content update
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs automated crawling and comparison of web-content with the data repository without requiring manual intervention. The automated system independently detects changes, compares content, and updates the repository, eliminating the need for human operators to manually check and update content while maintaining accuracy through systematic automated comparison processes

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical operations with automated computational processes. Instead of human operators manually comparing HTML documents and updating databases, the system uses automated software that performs web crawling, content comparison, and database updates through computational algorithms, thereby increasing productivity while maintaining controlled accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual comparison of HTML documents is performed, then update accuracy can be maintained, but computation time and labor requirements increase

Engineering Contradiction:
Improveupdate accuracyVSAvoidtime for manual comparison
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system continuously and automatically crawls web-content and compares it with the data repository in real-time or near-real-time. This continuous automated operation eliminates the intermittent manual comparison process, maintaining constant readiness for updates without the time loss associated with manual intervention cycles

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The patent substitutes manual document comparison with automated computational comparison systems. The automated system processes HTML documents, detects changes, and updates the repository through computational algorithms, eliminating the time-consuming manual comparison process while maintaining systematic accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Device complexity

If conventional manual update technique is used, then system complexity remains low, but reliability and accuracy of updates deteriorate due to human error

Engineering Contradiction:
Improvesystem simplicityVSAvoidupdate reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The automated system performs self-service operations for web-content crawling, comparison, and updates without human intervention. This self-service mechanism eliminates human error and improves reliability while the systematic automated process manages complexity through standardized algorithms and procedures

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback mechanisms where the automated crawler continuously monitors web-content changes, compares them with the data repository, and updates the repository based on detected modifications. This feedback loop ensures reliable and accurate updates by systematically detecting and responding to content changes without human error

Inventive Principle:
Principle #23Feedback

4Loss of information

If continuous database upgrades are performed to store new HTML documents, then data freshness is maintained, but economic costs increase due to large database requirements

Engineering Contradiction:
Improvedata freshnessVSAvoiddatabase storage volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The system extracts only the necessary web-content information and updates the data repository selectively based on detected changes and predefined criteria. Instead of continuously storing all HTML documents, the system extracts and stores only relevant updated content, reducing storage requirements while maintaining data freshness

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements parameter changes in the update process by using predefined criteria and thresholds for determining which content to update. The system changes the approach from storing all content to storing only content that meets specific criteria, thereby reducing database storage volume while maintaining freshness of relevant information

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11321400B2System and method for crawling web-content
Publication Date: 2022.05.03 INNOPLEXUS AG
  • US11321400B2 patent drawing
  • US11321400B2 patent drawing

AI summary

Disclosed is a system comprising: a data repository storing web-content; a data processing arrangement communicatively coupled to data repository, wherein data processing arrangement is configured to: acquire a web-page signature file associated to web-content, from a web-server hosting a website for displaying web-content, wherein web-page signature file includes a plurality of data related to web-content; analyse plurality of data included in web-page signature file to identify a modification in website; compare web-content stored in data repository with web-content displayed on website to determine additional web-content included in web-content displayed on website; use a machine learning algorithm to determine an importance value for additional web-content using a set of predefined parameters; crawl web-content stored in data repository based on additional web-content upon determining importance value to be greater than a predefined threshold value; and predict a time for crawling web-content using forecast module.