Image Clustering Using Co-Click Data and Extrinsic Information
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Search engines face challenges in effectively clustering digital images based on their subject matter, leading to irrelevant results being presented to users, as existing methods rely heavily on user-provided tags and relevance scores that may not accurately reflect the image content.
Innovation Solution
The system associates extrinsic image-related information, including co-click data and labels, to cluster images, allowing for the identification of clusters that match user queries and providing cluster results that include descriptive text and representative images, thereby exploiting users' search histories to categorize images without requiring explicit tags.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If user-provided tags and relevance scores are used to cluster images, then the clustering process is simple to implement, but the accuracy of image subject matter categorization deteriorates
Solution Approach 1:
The patent introduces co-click data as an intermediary signal to bridge the gap between user behavior and image subject matter. Instead of directly using user tags or relevance scores, the system uses co-click patterns (when users click on multiple images in sequence) as an indirect indicator of semantic relationships between images, thereby improving categorization accuracy without requiring complex manual tagging
Solution Approach 2:
The system implements feedback loops where user interaction data (co-clicks) is continuously collected and used to refine cluster assignments. The clustering algorithm iteratively improves by incorporating feedback from user behavior patterns, allowing the system to learn and adapt to actual user preferences and image semantics over time
2Reliability
If images are clustered based on multiple extrinsic information sources, then the relevance of search results improves, but the system complexity increases
Solution Approach 1:
The patent merges multiple extrinsic information sources (co-click data, labels, relevance scores) into a unified clustering framework. Instead of treating each data source separately, the system combines them into a comprehensive similarity metric that leverages the strengths of each individual source while reducing overall system complexity through integration
Solution Approach 2:
The clustering system is designed to be universal by accepting multiple types of extrinsic information inputs and processing them through a common framework. This multi-functional approach allows the same system to handle diverse data sources (user tags, automated labels, click behavior) without requiring separate processing pipelines for each type
3Loss of information
If co-click data is collected and processed to identify image relationships, then the understanding of image subject matter improves, but the data processing requirements increase
Solution Approach 1:
The system extracts only the essential co-click signals needed for clustering, filtering out redundant interaction data. Instead of processing all user interactions equally, the system identifies and extracts meaningful co-click patterns (sequential clicks on multiple images) that directly inform subject matter relationships, reducing unnecessary data processing volume
Data Source
AI summary
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for clustering images. In one aspect a system includes one or more computers configured to, for each of a plurality of digital images, associate extrinsic image-related information with each individual image, the extrinsic image-related information including text information and co-click data for the individual image, assign images from the plurality of images to one or more of the clusters of images based on the extrinsic information associated with each of the plurality of images, receive in the search system a user query from a user device, identify by operation of the search system one or more clusters of images that match the query, and provide one or more cluster results, where each cluster result provides information about an identified cluster.


