Web Browser Structured Data Extraction for Social Product Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current socially curated networks face challenges in indexing, searching, and comparing products due to unstructured data, lacking meta-data, and incomplete product information, which hinders efficient product comparison and search results across different retailers and manufacturers.

Innovation Solution

A method and system for creating templates to extract structured data from web pages, categorize, normalize, and index product records, allowing for efficient comparison and integration with user social graphs, enabling detailed product searches and analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If socially curated networks store unstructured data from web pages, then data capture is simple, but indexing and searching become inefficient

Engineering Contradiction:
Improvedata capture simplicityVSAvoidindexing and searching efficiency
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent applies preliminary action by extracting and structuring data from web pages at the time of capture, rather than attempting to structure unstructured data later. The system pre-processes web page content into structured formats (JSON, XML, CSV) with defined schemas during the initial data collection phase, enabling efficient subsequent indexing and searching without compromising capture simplicity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary data structure layer between raw web page content and the social network database. This intermediary structured format acts as a mediator that preserves the simplicity of web scraping while enabling efficient database operations. The structured data includes standardized fields for product information, pricing, and metadata that facilitate both easy capture and efficient retrieval.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If complete product records are extracted from web pages, then product comparison capability improves, but data processing complexity increases

Engineering Contradiction:
Improveproduct comparison accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing product records into distinct, standardized components such as product identification fields, pricing fields, specification fields, and metadata fields. Each segment is processed and validated independently according to its specific schema, reducing overall processing complexity while maintaining complete product information for accurate comparison.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms unstructured web page data into structured parameters with defined data types, validation rules, and relationships. By changing the parameter representation from free-form text to standardized data structures with explicit schemas, the system enables precise product comparison while managing complexity through consistent parameter handling across all products.

Inventive Principle:
Principle #35Parameter changes

3Stability of the object's composition

If structured data templates are created and applied, then data organization improves, but initial setup time increases

Engineering Contradiction:
Improvedata organizationVSAvoidtemplate creation time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent applies copying by creating reusable data templates that can be replicated across multiple product extractions. Once a structured data template is established for a particular product type or website format, it can be copied and applied to similar extractions without recreating the structure each time, significantly reducing setup time while maintaining consistent data organization.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent creates universal data templates that can handle multiple product types and website formats through flexible schema design. These templates incorporate optional fields and adaptive structures that work across different data sources, allowing a single template framework to serve multiple functions and reducing the need to create separate templates for each product category.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9606970B2Web browser device for structured data extraction and sharing via a social network
Publication Date: 2017.03.28 DATA RECORD SCI INC
  • US9606970B2 patent drawing
  • US9606970B2 patent drawing
  • US9606970B2 patent drawing

AI summary

A method and system for implementing a browser based information extraction and transmission method. A method and system for identifying, extracting, and transmitting predefined structured information from web pages browser interface. The extracted information is then added to a user profile on a social network and a database. The information is shared with other users who can comment, copy, vote on, or go to the original information source. The information can be combined with other extracted information to form collections for the purposes of voting on one or more items in the collection, combining multiple items to form a useful kit, saving information for later use, adding addition information such as dates and purchase location for personal inventory purposes, and for saving bookmarks to structured data.