Media Guidance Application Data Aggregation via Path Attributes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional media guidance systems face difficulties in aggregating information from multiple sources and mediums due to the need for manual data input and the inaccuracies caused by web crawlers that only scan rendered text without examining source code.

Innovation Solution

A media guidance application that automatically collects information by identifying path attributes associated with specific types of data, such as actor names, within webpages, and uses verified training data to determine accurate mappings for data retrieval, thereby reducing manual input and improving data accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If web crawlers merely search rendered text without examining source code, then the process is simpler, but data accuracy deteriorates

Engineering Contradiction:
Improvesimplicity of crawling processVSAvoiddata accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

Instead of searching for text content and assuming its meaning, the patent inverts the approach by searching for structural path attributes (like HTML tags and hierarchical positions) and using them to identify and extract data. This structural-based approach accurately distinguishes between different data types even when text appears in unexpected locations.

Inventive Principle:
Principle #13The other way round (Inversion)

2Measurement precision

If users manually input data into database format, then data accuracy can be controlled, but user effort and time increase significantly

Engineering Contradiction:
Improvedata accuracyVSAvoidtime for data entry
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs self-service by automatically analyzing webpage structure, identifying data fields through path attributes, extracting information, and formatting it for database storage without requiring manual user intervention. The crawler autonomously completes the entire data collection and preparation pipeline.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent establishes predetermined path attribute mappings based on verified training data before actual data collection. These pre-established structural patterns enable the crawler to automatically recognize and extract relevant information without needing manual guidance during the crawling process.

Inventive Principle:
Principle #10Preliminary action

3Extent of automation

If web crawlers are used to automate data collection, then user effort is reduced, but data accuracy deteriorates due to inability to distinguish data types

Engineering Contradiction:
Improveautomation of data collectionVSAvoiddata type identification accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent applies local quality by examining specific structural characteristics (path attributes) at different locations within the webpage hierarchy. Each data type is identified by its unique structural context rather than relying on global text patterns, enabling accurate differentiation of data types throughout the document.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system uses verified training data to establish accurate path attribute mappings and continuously refines its data extraction capabilities. By comparing extracted data against known correct values, the system learns and improves its ability to identify data types accurately across different webpage structures.

Inventive Principle:
Principle #23Feedback

4Loss of information

If multiple data sources and mediums are aggregated, then information completeness improves, but system complexity increases

Engineering Contradiction:
Improveinformation completenessVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent creates a universal data extraction framework that works across multiple data sources and mediums by focusing on common structural patterns (HTML path attributes). This single approach can extract data from various webpage formats and structures, eliminating the need for separate handling mechanisms for different sources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10650065B2Methods and systems for aggregating data from webpages using path attributes
Publication Date: 2020.05.12 ROVI PRODUCT CORP
  • US10650065B2 patent drawing
  • US10650065B2 patent drawing
  • US10650065B2 patent drawing

AI summary

Methods and systems that improve the ability of a media guidance application to aggregate information from one or more sources and one or more mediums. For example, the media guidance application may automatically collect information based on attributes associated with information of a particular type. Specifically, the media guidance application may determine based on comparison with verified training data that one source or medium typically associates information of a particular type, for example, “Actor,” with one or more path attributes, for example, a location in a directory structure. The media guidance application may then search the source or medium for the one or more path attributes. Upon detecting the one or more path attributes, the media guidance application may designate any sub-set of information associated with the one or more path attributes as a particular type of information.