Targeted Crawler for Media Catalog Accuracy and Resource Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The rapid growth and frequent changes in media content make it difficult for user devices to obtain and maintain accurate information about available media content across multiple content providers, including determining which providers offer specific content and when it is available.
Innovation Solution
A targeted crawling system that uses an electronic program guide (EPG) data receiver and a media content catalog enhancer to identify new media content, causing a web crawler to retrieve information from source websites and store it in a searchable database, allowing end users to access media content through their devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If continuous crawling of all content provider websites is performed to maintain accurate media content information, then database accuracy is improved, but resource consumption and access denials increase
Solution Approach 1:
The system performs preliminary actions by crawling content provider websites at or near the time new media content becomes available, rather than continuously crawling all content. This preliminary timing ensures accurate capture of new content information while avoiding unnecessary resource consumption from repeated crawling of existing content.
Solution Approach 2:
Instead of continuous crawling, the system employs periodic action by crawling content provider websites at specific intervals or events (when new content is expected to be available). This periodic approach maintains database accuracy for new content while significantly reducing overall resource consumption compared to continuous crawling.
2Loss of time
If frequent crawling of content provider websites is performed to update media content information, then database timeliness is improved, but access denials from content providers increase
Solution Approach 1:
The system performs crawling at or near the time new media content becomes available, which is the critical moment for database timeliness. This preliminary timing ensures the database is updated exactly when needed without triggering excessive access requests that would cause content providers to block the system.
Solution Approach 2:
The system applies local quality by focusing crawling efforts only on content that is newly available or expected to be available, rather than uniformly crawling all content providers at all times. This selective approach maintains timeliness for new content while minimizing overall access frequency to avoid denials.
3Adaptability or versatility
If comprehensive crawling of all content providers is performed to ensure complete media content coverage, then content availability is improved, but system complexity and access denials increase
Solution Approach 1:
The system segments the content provider website crawling task by focusing only on content that is newly available or expected to be available, rather than uniformly crawling all content providers at all times. This segmentation maintains comprehensive coverage for new content while significantly reducing system complexity and access denial risks.
Solution Approach 2:
The system performs preliminary identification of new media content through EPG data and other sources before crawling content provider websites. This preliminary action filters the crawling scope to only relevant content, ensuring complete coverage of new content while reducing overall system complexity and access request volume.
Data Source
AI summary
A system is described that includes an electronic program guide (EPG) data receiver and a media content catalog enhancer. The EPG receiver is configured to receive EPG data from an EPG data provider. The media content catalog enhancer is configured to determine that an item of media content identified by the EPG data comprises new media content and, in response to determining that the item of media content identified by the EPG data comprises new media content, to cause a web crawler to crawl a source website associated with the new media content to obtain information about the new media content and to store the obtained information about the new media content in a database, the database comprising a catalog of media content that is searchable by an end user to identify and access content for playback via an end user device.


