Guided Web Crawler With Pluggable Modules for Resource Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional web crawlers lack the ability to efficiently identify and verify specific webpage resources without manual effort, as they do not account for the unique layouts and data presentations of different websites, leading to inefficient data processing and resource identification.

Innovation Solution

A guided web crawler system utilizing pluggable coded modules that simulate user interactions, such as clicking and form filling, to automate the identification and verification of webpage resources, leveraging machine learning and user-configurable scripts to enhance efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional web crawlers are used to browse webpages, then basic webpage data can be collected, but manual effort is required to determine webpage elements and features which is time-consuming and resource-intensive

Engineering Contradiction:
Improvewebpage resource identification efficiencyVSAvoidmanual processing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system enables automated self-service by training the machine learning model to independently identify and verify webpage resources without manual intervention. The model learns from training data to automatically determine webpage elements, features, and resources, eliminating the need for manual processing while maintaining high accuracy in resource identification.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by pre-training the machine learning model with training data before actual webpage crawling. This preliminary training phase enables the model to possess the knowledge and capabilities needed for efficient automated resource identification, reducing the time required during actual crawling operations.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If conventional web crawlers browse all webpage data, then comprehensive data is collected, but unnecessary data processing occurs which reduces efficiency

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidcomputational energy
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The system extracts only the necessary webpage resources by using the trained machine learning model to identify and filter relevant elements. Instead of processing all webpage data, the model extracts specific resources of interest, eliminating unnecessary data processing and reducing computational energy consumption while maintaining comprehensive coverage of important elements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies local quality by treating different webpage elements with different processing priorities. The machine learning model identifies which specific webpage resources require attention and applies targeted processing only to those elements, rather than uniformly processing all webpage data, thereby improving efficiency and reducing energy waste.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If conventional web crawlers are used without website-specific knowledge, then general browsing is possible, but specific webpage resources cannot be properly identified and verified

Engineering Contradiction:
Improveresource identification accuracyVSAvoidcrawling system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system changes parameters by training the machine learning model with website-specific characteristics and patterns from training data. This enables the model to adapt to different website structures and layouts, improving resource identification accuracy without requiring separate configuration for each website. The model learns to recognize patterns and parameters specific to different webpage types.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system achieves universality by creating a multi-functional machine learning model that can handle multiple website types and layouts with a single trained system. The model is designed to work across diverse websites by learning common patterns and adapting to variations, eliminating the need for separate crawling systems for different websites while maintaining high identification accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250225189A1Guided web crawler for automated identification and verification of webpage resources
Publication Date: 2025.07.10 PAYPAL INC
  • US20250225189A1 patent drawing
  • US20250225189A1 patent drawing
  • US20250225189A1 patent drawing

AI summary

There are provided systems and methods for a guided web crawler for automated identification and verification of webpage resources. A service provider, such as an online transaction processor, may provide a guided web crawler and/or resources for such crawler for execution by computing devices of users. Users may load different pluggable modules to the guided web crawler, which are associated with specific web crawling tasks. Web crawling tasks may correspond to identification and verification of webpage resources on a webpage, such as a location, placement, use of, and/or number of appearances of the resource. The web crawler may use code from the pluggable module being executed to parse and/or crawl webpage data for a webpage and identify requested resources. Thereafter, the guided web crawler may automate resources to use, display, and/or interact with the identified and verified resource.