Visual Leaf Page Identification via Hub Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current search engines face challenges in effectively identifying and ranking visual leaf pages, which are terminal web pages prominently featuring images or videos, leading to suboptimal search results in image-based searches.
Innovation Solution
A method and system that identify visual leaf pages by analyzing image data prominence, semantic relatedness, and structural features, generating cluster data, and classifying web pages as visual leaf pages based on feature values and model scores, allowing for improved ranking in search results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional search engine methods are used to identify and rank web pages, then general search functionality is maintained, but visual leaf pages cannot be effectively identified and ranked, leading to suboptimal search results for image-based searches
Solution Approach 1:
The patent segments web pages into different categories (visual leaf pages, text leaf pages, hub pages) based on their characteristics. This segmentation allows the search engine to apply different ranking strategies to different page types, improving the accuracy of visual leaf page identification while maintaining manageable system complexity through modular classification rules.
Solution Approach 2:
The patent introduces specific parameters for identifying visual leaf pages, such as the prominence of image data relative to other content, the presence of image-based links from hub pages, and the terminal nature of the page. By changing the parameters used for page evaluation from text-only metrics to include visual prominence metrics, the system can accurately identify visual leaf pages without requiring overly complex classification mechanisms.
2Measurement precision
If manual annotation methods are used to classify visual leaf pages, then classification accuracy can be improved, but the process requires significant human effort and time
Solution Approach 1:
The patent implements a self-service classification system where web pages automatically classify themselves as visual leaf pages, text leaf pages, or hub pages based on their inherent characteristics. The system uses automated algorithms to evaluate image prominence, link structures, and content types, eliminating the need for manual human annotation while maintaining high classification accuracy through objective, consistent criteria.
Solution Approach 2:
The patent replaces the mechanical human annotation process with an automated computational system. Instead of relying on human reviewers to manually classify pages, the system uses algorithmic analysis of page features (image data prominence, hub page link structures, content ratios) to automatically and accurately classify visual leaf pages, significantly reducing time loss while maintaining or improving classification precision.
3Productivity
If the system analyzes multiple features and generates cluster data for visual leaf pages, then search result relevance is enhanced, but the processing complexity and computational resources increase
Solution Approach 1:
The patent performs preliminary analysis by identifying hub pages and their image-based links before classifying visual leaf pages. This preliminary action organizes the data structure in advance, creating a framework that simplifies subsequent classification processing. By pre-identifying the hierarchical relationships and image link structures, the system reduces the complexity of the main classification task while enhancing search result relevance through comprehensive feature analysis.
Solution Approach 2:
The patent adds a new dimension to page analysis by incorporating visual prominence metrics alongside traditional text-based features. Instead of only analyzing text content, the system evaluates the prominence of image data, the structure of image-based links, and the visual characteristics of pages. This dimensional expansion enriches the feature set for classification, improving search result relevance while managing complexity through structured multi-dimensional analysis.
Data Source
AI summary
In some implementations, a method includes, for each of multiple hosts: identifying visual leaf pages hosted by the host that are each a web page including image data defining an image or a video that is prominently displayed relative to all other content of the web page, identifying a set of hub pages hosted by the host that each link to at least one of the visual leaf pages through an image-based link, and for each hub page, generating cluster data representing the visual leaf pages to which the hub page links by determining, for each visual leaf page, a set of feature values that each indicate pre-defined features of the visual leaf page, and generating, from the sets of feature values, a set of central feature values as the cluster data for the hub page that indicate a central tendency of each respective pre-defined feature.


