Primary Product Object Identification via Statistical Probability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional search algorithms on the Internet rely on keyword-based indexing, failing to accurately account for product-related information such as images and titles, leading to low accuracy in search results for online shopping.
Innovation Solution
A method and system that identifies primary product objects on web pages by extracting features and computing probabilities using a statistical model, selecting objects based on these probabilities to enhance search accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional keyword-based search algorithms are used, then the search process is simple and fast, but the accuracy of search results is low
Solution Approach 1:
The patent segments the web page into multiple objects (product images, titles, prices, descriptions) and evaluates each object independently using different features. This segmentation allows the system to focus on specific product-related elements rather than processing the entire page as a single unit, thereby improving search accuracy while managing complexity through modular evaluation.
Solution Approach 2:
The patent changes the parameters used for evaluation from simple keyword matching to multiple product-specific features including image attributes (size, position, file name), text attributes (title, description, price), and contextual attributes. This parameter transformation enables more accurate identification of primary product objects by considering multiple dimensions of product information.
2Measurement precision
If multiple product-related features are analyzed, then the accuracy of identifying primary product objects improves, but the computational complexity increases
Solution Approach 1:
The patent extracts and isolates specific product-related features from the complex web page structure, focusing only on relevant attributes such as image dimensions, file names, titles, prices, and descriptions. By extracting only the necessary features rather than analyzing all page elements, the system achieves high identification accuracy while reducing unnecessary computational overhead.
Solution Approach 2:
The system uses the inherent structure and attributes of product-related objects themselves to identify primary products. For example, it leverages the natural positioning, size, and textual content of product images and descriptions to determine their importance, allowing the data to inform its own classification without requiring excessive external computational resources.
3Loss of information
If comprehensive product information is extracted from web pages, then the relevance of search results improves, but the time required for processing increases
Solution Approach 1:
The patent performs preliminary extraction and evaluation of product-related features during the indexing phase, preparing product object data in advance. By pre-processing and organizing product information (images, titles, prices, descriptions) before actual search queries, the system reduces processing time during live searches while maintaining comprehensive product information availability.
Data Source
AI summary
Methods, systems and computer program products for identifying primary product objects on a web page. A primary product object is the object that shows the best view of the product the web page is detailing. A set of features is extracted for one or more objects on the web page. The primary product objects are identified by computing the probabilities of one or more objects on the web page being a primary product object, the probabilities indicating the likelihood of the one or more objects being the primary product object. The probabilities are computed by querying a statistical model.


