Multimodal Scene Graph Editing for Structural Media Variation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems lack intuitive and practical mechanisms to create, edit, and share multimedia content that encompasses different varieties, forms, and media types, limiting the ability to generate different variations of media with structural changes.

Innovation Solution

The use of multimodal scene graphs to process visual information, recognize objects, and generate media elements such as images and videos, allowing for editing and personalization through scene managers and trained machine learning models to create variations with replaced avatars, backgrounds, and captions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional predefined filters and constructs are used for media creation, then the system is simple to operate, but the media variation and personalization capability is limited

Engineering Contradiction:
Improvemedia variation capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments media content into structured scene graphs with discrete components (objects, attributes, relationships) that can be independently manipulated. This allows users to create varied media by reconfiguring individual elements rather than relying on predefined filters, resolving the contradiction between media variation capability and system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The scene graph structure serves as a universal framework that can represent multiple media types (images, videos, augmented reality content) and enable various operations (creation, editing, transformation) through a single consistent interface. This multi-functional approach increases adaptability without proportionally increasing complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If predefined filters are applied to captured media, then the operation is simple, but the ability to create transformed content with structural changes is limited

Engineering Contradiction:
Improveease of media editingVSAvoidcontent transformation capability
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary action by automatically generating scene graphs from captured media, pre-structuring the content into editable components before the user begins editing. This automation reduces the complexity of structural transformations while maintaining ease of operation, as users work with pre-parsed scene graphs rather than raw media.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The scene graph acts as an intermediary representation between the captured media and the final transformed content. Users edit the intermediate scene graph structure rather than directly manipulating raw media, enabling complex structural transformations while maintaining user-friendly operation through standardized editing interfaces.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If a multimodal scene graph system is implemented to generate varied media elements, then the adaptability and personalization are enhanced, but the system complexity increases

Engineering Contradiction:
Improvepersonalization capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements self-service by automatically generating scene graphs, extracting objects and attributes, and organizing relationships without requiring manual user input for these complex tasks. This automation handles the complexity internally while presenting a simplified interface to users, enabling personalization capabilities without exposing the underlying system complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system enables personalization through parameter changes in the scene graph structure, allowing users to modify attributes, relationships, and component properties. This approach provides high adaptability by changing parameters within an existing framework rather than requiring complex system reconfiguration, balancing personalization capability with manageable system complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250316000A1Multimodal Scene Graph for Generating Media Elements
Publication Date: 2025.10.09 META PLATFORMS TECHNOLOGIES LLC
  • US20250316000A1 patent drawing
  • US20250316000A1 patent drawing
  • US20250316000A1 patent drawing

AI summary

Aspects of the present disclosure are directed to generating media element(s) using a multimodal scene graph. A scene manager can process visual information, such as video, images, and/or a recorded artificial relay scene, and generate a multimodal scene graph that comprises components and metadata generated via the processing. The scene manager can utilize the multimodal scene graph to generate social media elements, such as images, video, and/or artificial reality scenes. For example, a video of a user can be converted to a multimodal scene graph, which can be used to generate one or more images (e.g., memes, animated images, stickers, etc.), such as an image that represents the user via an avatar of the user. This generated media can be shared with other social platform users, and the stored multimodal scene graph can be accessed by the others to generate variations of the media.