Visual Object Graphs for Cross-Scene Object Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing volume of digital content makes it difficult for users to discover relevant visual content, particularly objects or items represented in visual scenes.

Innovation Solution

A system that processes a corpus of scenes to segment individual objects, forms object clusters based on visual similarity, and links scenes that include these objects, allowing users to select query objects and receive matching scenes that include the same or similar objects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If users search for visual content using traditional methods, then they can access digital content, but it becomes increasingly difficult to discover relevant content as the volume of content expands

Engineering Contradiction:
Improvevolume of digital contentVSAvoiddifficulty to discover relevant content
Core Design Contradiction:
Quantity of substanceVSEase of operation

Solution Approach 1:

The system segments visual content into individual objects by detecting and extracting objects from scenes. Each object is represented as a separate entity with its own visual features and scene context, enabling independent indexing and retrieval. This segmentation allows users to search for specific objects rather than navigating through entire scenes, resolving the contradiction between content volume and discoverability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary representation layer between raw visual content and user queries. Object graphs and visual embeddings serve as intermediaries that transform scene data into structured representations, enabling efficient matching between user queries and content. This intermediary layer abstracts the complexity of large-scale visual content, making discovery easier despite the expanding volume of content.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If users search for specific objects in visual content, then they can find relevant objects, but traditional image search requires the entire scene to match which limits search precision

Engineering Contradiction:
Improvesearch precision for objectsVSAvoidflexibility in scene matching
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system segments scenes into individual objects, allowing users to search for specific objects independent of the entire scene. Each object is extracted with its visual features and associated scene context, enabling precise object-level search while maintaining the ability to retrieve scenes containing those objects. This resolves the contradiction by decoupling object search from scene-level matching constraints.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from two-dimensional image search to a multi-dimensional object graph representation. Objects are indexed with multiple attributes including visual features, spatial relationships, and scene context, creating a richer search space. This dimensional expansion allows precise object search while maintaining scene flexibility, as the system can match objects across different scene contexts.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250291840A1Visual object graphs
Publication Date: 2025.09.18 PINTEREST INC
  • US20250291840A1 patent drawing
  • US20250291840A1 patent drawing
  • US20250291840A1 patent drawing

AI summary

Disclosed are implementations that enable the linking or connection of objects and different scenes in which those objects are represented. For example, a corpus of scenes (e.g., digital images) that include a representation of one or more objects may be processed using the disclosed implementations to segment from those scenes the individual objects represented in those scenes. The disclosed implementations may further determine clusters of visually similar object segments and form object clusters for those object segments. The scenes that include those object segments are also linked to the object cluster. With scenes linked to different object clusters, a user may select one or more query objects or a query scene and be presented with other scenes that include visually similar objects, even though the overall scenes may be visually different.