Web Data Extraction via Seed Site Feature Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for extracting structured web data from various web sites require significant human effort and are not cost-effective, as they rely on manual labeling or template approaches, which are inefficient for handling large volumes of data across different verticals like books, cars, or restaurants.

Innovation Solution

The approach involves labeling text nodes on a seed site for each vertical, extracting features, learning vertical knowledge, and adapting this knowledge to new, unlabeled web sites to automatically identify and extract targeted attributes, reducing human involvement and enhancing scalability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual labeling approach is used to extract structured web data, then extraction accuracy can be maintained, but human time and effort requirements increase significantly

Engineering Contradiction:
Improveextraction accuracyVSAvoidhuman time and effort
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent creates a labeled seed site that serves as a template or copy of the labeling process. Instead of manually labeling every web page, the system extracts features from the labeled seed site and uses these features to automatically label and extract data from numerous other web sites, effectively copying the labeling action across multiple sites without requiring proportional human effort for each site.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables self-service by automatically extracting features from the labeled seed site and using these features to autonomously label and extract data from new web sites. The feature extraction and learning mechanisms allow the system to serve itself by automatically adapting to different web site structures without continuous human intervention.

Inventive Principle:
Principle #25Self-service

2Ease of manufacture

If template approach is used for web data extraction, then extraction process can be standardized, but cost effectiveness deteriorates due to high human effort requirements

Engineering Contradiction:
Improveextraction process standardizationVSAvoidcost effectiveness
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent uses a labeled seed site as a template that captures the essential labeling patterns for a given vertical. By extracting features from this template and applying them across multiple web sites, the system standardizes the extraction process while eliminating the need for proportional human effort for each site, thereby improving cost effectiveness.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the parameters of the extraction process by transitioning from manual human labeling to automated feature-based labeling. The feature extraction module transforms manual labeling into a computational process that can be replicated across numerous sites, maintaining standardization while dramatically reducing human resource requirements.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If conventional extraction methods are applied to numerous web sites, then data collection can be performed, but scalability is limited due to manual processing requirements

Engineering Contradiction:
Improvevolume of web dataVSAvoidmanual processing level
Core Design Contradiction:
Quantity of substanceVSExtent of automation

Solution Approach 1:

The patent enables scalability by extracting features from a single labeled seed site and copying these feature representations to automatically label and extract data from numerous other web sites. This approach allows the system to handle large volumes of data across multiple sites without requiring proportional increases in manual processing capacity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system achieves scalability through self-service automation where the feature extraction and learning mechanisms automatically process new web sites without human intervention. The system serves itself by using the learned features to autonomously extract data from numerous sites, enabling large-scale data collection while maintaining minimal manual processing.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8856129B2Flexible and scalable structured web data extraction
Publication Date: 2014.10.07 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8856129B2 patent drawing
  • US8856129B2 patent drawing
  • US8856129B2 patent drawing

AI summary

This document describes techniques that label text nodes of a seed site for each of a plurality of verticals. Once a seed site is labeled for a given vertical, the techniques extract features from the labeled text nodes of the seed site. The techniques learn vertical knowledge for the seed site based on the human labels and the extracted features, and adapt the learned vertical knowledge to a new web site to automatically and accurately identify attributes and extract attribute values targeted within a given vertical for structured web data extraction.