Memory Video Generation From Natural-Language Media Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Managing and navigating large media libraries efficiently to retrieve contextually-relevant content is cumbersome and time-consuming, particularly in battery-operated devices, due to inefficient user interfaces and manual navigation.
Innovation Solution
A method utilizing machine-learning models to generate a video from media assets based on a natural language description of a memory, involving multiple machine-learning adapters to process query tokens, traits, and story outlines, enabling efficient retrieval and arrangement of media items into a cohesive video.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If users manually navigate through media libraries to retrieve contextually-relevant content, then they can access their media assets, but the process becomes time-consuming and cumbersome
Solution Approach 1:
The system automatically retrieves and organizes media content based on natural language descriptions without requiring manual user navigation. The machine learning model processes the description, searches the media library, and assembles relevant content automatically, eliminating the need for users to manually browse through directories.
Solution Approach 2:
The patent replaces manual mechanical navigation through user interfaces with automated machine learning-based content retrieval. Instead of requiring users to press buttons or navigate menus, the system uses natural language processing and machine learning algorithms to automatically find and present relevant media content.
2Measurement precision
If existing techniques use complex user interfaces to retrieve contextually-relevant content, then they can filter and search media assets, but the process requires multiple key presses and wastes user time and device energy
Solution Approach 1:
The system replaces complex mechanical user interface operations with automated machine learning-based search. The machine learning model processes natural language descriptions and automatically retrieves relevant media content without requiring multiple key presses or manual navigation steps.
Solution Approach 2:
The system changes the search parameters from manual user inputs to machine-generated queries based on natural language processing. The machine learning model converts descriptive language into search parameters, automatically filtering and retrieving relevant content without user intervention.
3Productivity
If users manually search through large media libraries, then they can find specific content, but the process is inefficient and wastes device resources
Solution Approach 1:
The system performs self-service by automatically executing the search and retrieval process without requiring active user participation. The machine learning model independently processes descriptions, searches the media library, and assembles results, eliminating the need for continuous user input and reducing overall system energy consumption.
Solution Approach 2:
The system performs preliminary actions by pre-processing media assets and organizing them in a way that enables rapid retrieval. The machine learning model analyzes and indexes media content in advance, so when a search is needed, the system can quickly retrieve relevant content without requiring extensive real-time processing.
Data Source
AI summary
The present disclosure generally relates to generating a video corresponding to a memory (e.g., an event or context) from media assets on a device. In some embodiments, the device receives user inputs requesting a video based on a natural language description of a memory. The device sends information of the natural language description to a first machine-learning (ML) model, and receives query tokens, which are used to find media items on the device that match the query tokens. The device sends information representing the found media items to another ML model that determines traits from the media items. These traits are sent to a third ML model to generate a story outline, and the video is generated by comparing the descriptions of shots in the story outline to visual embeddings of the found media assets to curate and arrange them into the video consistent with the story outline.


