Media Guidance Application Data Aggregation via Path Attributes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional media guidance systems face difficulties in aggregating information from multiple sources and mediums due to the need for manual data input and the inaccuracies caused by web crawlers that only scan rendered text without examining source code.
Innovation Solution
A media guidance application that automatically collects information by identifying path attributes associated with specific types of data, such as actor names, within webpages, and uses verified training data to determine accurate mappings for data retrieval, thereby reducing manual input and improving data accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If web crawlers merely search rendered text without examining source code, then the process is simpler, but data accuracy deteriorates
Solution Approach 1:
Instead of searching for text content and assuming its meaning, the patent inverts the approach by searching for structural path attributes (like HTML tags and hierarchical positions) and using them to identify and extract data. This structural-based approach accurately distinguishes between different data types even when text appears in unexpected locations.
2Measurement precision
If users manually input data into database format, then data accuracy can be controlled, but user effort and time increase significantly
Solution Approach 1:
The system performs self-service by automatically analyzing webpage structure, identifying data fields through path attributes, extracting information, and formatting it for database storage without requiring manual user intervention. The crawler autonomously completes the entire data collection and preparation pipeline.
Solution Approach 2:
The patent establishes predetermined path attribute mappings based on verified training data before actual data collection. These pre-established structural patterns enable the crawler to automatically recognize and extract relevant information without needing manual guidance during the crawling process.
3Extent of automation
If web crawlers are used to automate data collection, then user effort is reduced, but data accuracy deteriorates due to inability to distinguish data types
Solution Approach 1:
The patent applies local quality by examining specific structural characteristics (path attributes) at different locations within the webpage hierarchy. Each data type is identified by its unique structural context rather than relying on global text patterns, enabling accurate differentiation of data types throughout the document.
Solution Approach 2:
The system uses verified training data to establish accurate path attribute mappings and continuously refines its data extraction capabilities. By comparing extracted data against known correct values, the system learns and improves its ability to identify data types accurately across different webpage structures.
4Loss of information
If multiple data sources and mediums are aggregated, then information completeness improves, but system complexity increases
Solution Approach 1:
The patent creates a universal data extraction framework that works across multiple data sources and mediums by focusing on common structural patterns (HTML path attributes). This single approach can extract data from various webpage formats and structures, eliminating the need for separate handling mechanisms for different sources.
Data Source
AI summary
Methods and systems that improve the ability of a media guidance application to aggregate information from one or more sources and one or more mediums. For example, the media guidance application may automatically collect information based on attributes associated with information of a particular type. Specifically, the media guidance application may determine based on comparison with verified training data that one source or medium typically associates information of a particular type, for example, “Actor,” with one or more path attributes, for example, a location in a directory structure. The media guidance application may then search the source or medium for the one or more path attributes. Upon detecting the one or more path attributes, the media guidance application may designate any sub-set of information associated with the one or more path attributes as a particular type of information.


