Visual Object Graphs for Cross-Scene Object Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing volume of digital content makes it difficult for users to discover relevant visual content, particularly objects or items represented in visual scenes.
Innovation Solution
A system that processes a corpus of scenes to segment individual objects, forms object clusters based on visual similarity, and links scenes that include these objects, allowing users to select query objects and receive matching scenes that include the same or similar objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If users search for visual content using traditional methods, then they can access digital content, but it becomes increasingly difficult to discover relevant content as the volume of content expands
Solution Approach 1:
The system segments visual content into individual objects by detecting and extracting objects from scenes. Each object is represented as a separate entity with its own visual features and scene context, enabling independent indexing and retrieval. This segmentation allows users to search for specific objects rather than navigating through entire scenes, resolving the contradiction between content volume and discoverability.
Solution Approach 2:
The patent introduces an intermediary representation layer between raw visual content and user queries. Object graphs and visual embeddings serve as intermediaries that transform scene data into structured representations, enabling efficient matching between user queries and content. This intermediary layer abstracts the complexity of large-scale visual content, making discovery easier despite the expanding volume of content.
2Measurement precision
If users search for specific objects in visual content, then they can find relevant objects, but traditional image search requires the entire scene to match which limits search precision
Solution Approach 1:
The system segments scenes into individual objects, allowing users to search for specific objects independent of the entire scene. Each object is extracted with its visual features and associated scene context, enabling precise object-level search while maintaining the ability to retrieve scenes containing those objects. This resolves the contradiction by decoupling object search from scene-level matching constraints.
Solution Approach 2:
The patent transitions from two-dimensional image search to a multi-dimensional object graph representation. Objects are indexed with multiple attributes including visual features, spatial relationships, and scene context, creating a richer search space. This dimensional expansion allows precise object search while maintaining scene flexibility, as the system can match objects across different scene contexts.
Data Source
AI summary
Disclosed are implementations that enable the linking or connection of objects and different scenes in which those objects are represented. For example, a corpus of scenes (e.g., digital images) that include a representation of one or more objects may be processed using the disclosed implementations to segment from those scenes the individual objects represented in those scenes. The disclosed implementations may further determine clusters of visually similar object segments and form object clusters for those object segments. The scenes that include those object segments are also linked to the object cluster. With scenes linked to different object clusters, a user may select one or more query objects or a query scene and be presented with other scenes that include visually similar objects, even though the overall scenes may be visually different.


