Script-Driven Browser for Automated Web Data Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional client-server communication techniques face challenges with scale and complexity when accessing and comparing data from multiple remote servers, especially when dealing with massive, ever-changing data sets, leading to inefficiencies in data retrieval and comparison processes.

Innovation Solution

A script-driven browsing application is developed that pairs scripts with a browser, utilizing XPath expressions on a DOM representation of web pages to facilitate efficient and comprehensive data extraction, enabling systematic and automated interaction with remote servers for data retrieval and comparison.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional client-server communication techniques are used to access data from multiple remote servers, then basic data retrieval is possible, but the system faces challenges with scale and complexity when dealing with massive, ever-changing data sets

Engineering Contradiction:
Improvedata retrieval efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a script-driven browsing application as an intermediary layer between the user and multiple remote servers. This application automates the interaction with web browsers, handles navigation, and manages data extraction processes, thereby reducing the complexity of directly managing multiple client-server connections while improving data retrieval efficiency through systematic automation

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If traditional methods are used for data extraction from web pages, then simple data retrieval is possible, but the process becomes inefficient when handling comprehensive data extraction from multiple web pages

Engineering Contradiction:
Improvedata extraction efficiencyVSAvoiddata retrieval time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent employs preliminary action by pre-defining XPath expressions that specify the exact location and structure of desired data elements within web page DOM representations. These expressions are prepared in advance, allowing the script-driven application to quickly locate and extract relevant data without manual navigation or analysis during the actual data retrieval process, thereby improving efficiency and reducing time loss

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The script-driven browsing application performs self-service by automatically navigating through multiple web pages, identifying data elements using pre-defined XPath expressions, and extracting the required information without human intervention. This automation eliminates manual data extraction processes, significantly improving productivity while reducing the time required for comprehensive data retrieval from multiple sources

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10579712B1Script-driven data extraction using a browser
Publication Date: 2020.03.03 TRAVELPORT INT OPERATIONS LTD
  • US10579712B1 patent drawing
  • US10579712B1 patent drawing
  • US10579712B1 patent drawing

AI summary

Various example embodiments are directed to a script-driven browsing application for automatically interacting with remote servers. Retrieval of data by virtue of a script-driven browsing application may be facilitated by pairing a script or scripts with a suitable browser, and employing XPath expressions operating on a DOM representation of one or more web pages such that the script-driven browsing application can iterate efficiently and comprehensively over the web pages.