Web Browser Structured Data Extraction for Social Product Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current socially curated networks face challenges in indexing, searching, and comparing products due to unstructured data, lacking meta-data, and incomplete product information, which hinders efficient product comparison and search results across different retailers and manufacturers.
Innovation Solution
A method and system for creating templates to extract structured data from web pages, categorize, normalize, and index product records, allowing for efficient comparison and integration with user social graphs, enabling detailed product searches and analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If socially curated networks store unstructured data from web pages, then data capture is simple, but indexing and searching become inefficient
Solution Approach 1:
The patent applies preliminary action by extracting and structuring data from web pages at the time of capture, rather than attempting to structure unstructured data later. The system pre-processes web page content into structured formats (JSON, XML, CSV) with defined schemas during the initial data collection phase, enabling efficient subsequent indexing and searching without compromising capture simplicity.
Solution Approach 2:
The patent introduces an intermediary data structure layer between raw web page content and the social network database. This intermediary structured format acts as a mediator that preserves the simplicity of web scraping while enabling efficient database operations. The structured data includes standardized fields for product information, pricing, and metadata that facilitate both easy capture and efficient retrieval.
2Measurement precision
If complete product records are extracted from web pages, then product comparison capability improves, but data processing complexity increases
Solution Approach 1:
The patent applies segmentation by dividing product records into distinct, standardized components such as product identification fields, pricing fields, specification fields, and metadata fields. Each segment is processed and validated independently according to its specific schema, reducing overall processing complexity while maintaining complete product information for accurate comparison.
Solution Approach 2:
The patent transforms unstructured web page data into structured parameters with defined data types, validation rules, and relationships. By changing the parameter representation from free-form text to standardized data structures with explicit schemas, the system enables precise product comparison while managing complexity through consistent parameter handling across all products.
3Stability of the object's composition
If structured data templates are created and applied, then data organization improves, but initial setup time increases
Solution Approach 1:
The patent applies copying by creating reusable data templates that can be replicated across multiple product extractions. Once a structured data template is established for a particular product type or website format, it can be copied and applied to similar extractions without recreating the structure each time, significantly reducing setup time while maintaining consistent data organization.
Solution Approach 2:
The patent creates universal data templates that can handle multiple product types and website formats through flexible schema design. These templates incorporate optional fields and adaptive structures that work across different data sources, allowing a single template framework to serve multiple functions and reducing the need to create separate templates for each product category.
Data Source
AI summary
A method and system for implementing a browser based information extraction and transmission method. A method and system for identifying, extracting, and transmitting predefined structured information from web pages browser interface. The extracted information is then added to a user profile on a social network and a database. The information is shared with other users who can comment, copy, vote on, or go to the original information source. The information can be combined with other extracted information to form collections for the purposes of voting on one or more items in the collection, combining multiple items to form a useful kit, saving information for later use, adding addition information such as dates and purchase location for personal inventory purposes, and for saving bookmarks to structured data.


