Proxy-Based Video Asset Generation Platform

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video production requires diverse technical expertise and is hindered by the complexity of machine learning models, which are difficult to control, and existing systems treat videos as indivisible entities, limiting user interaction and creativity.

Innovation Solution

An interactive platform that allows users to guide machine-learning asset enhancement modules via proxy elements in a video production workspace, enabling the transformation and generation of multimedia assets such as text, images, and audio, using pre-trained ML modules that can produce assets in various styles.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If complex machine learning models are used for video generation, then the quality and versatility of video assets are improved, but the difficulty of controlling and operating the system increases

Engineering Contradiction:
Improvevideo asset generation capabilityVSAvoiduser control difficulty
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent introduces a workspace environment with proxy elements that serve as intermediaries between users and machine learning models. Users manipulate simple proxy representations (images, text, audio clips) in the workspace rather than directly controlling complex ML models. The system automatically translates these proxy manipulations into appropriate ML model inputs and parameters, bridging the gap between user capability and model complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates simplified copies (proxies) of the actual video assets and model parameters that users can manipulate. Instead of directly editing complex model weights or parameters, users work with visual proxies like thumbnail images, text descriptions, and audio clips that represent the underlying assets. These proxies are automatically translated into the appropriate format for the machine learning models.

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If video production systems include multiple specialized modules for different video elements, then the quality and professionalism of output is improved, but the system complexity and learning curve increase

Engineering Contradiction:
Improvevideo production qualityVSAvoidsystem structure complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple specialized video production modules (image generation, text-to-speech, video synthesis, etc.) into a single unified workspace environment. Users interact with all these different functionalities through a common interface and the same proxy-based interaction model, rather than learning separate tools for each function. The system automatically routes user actions to the appropriate specialized modules.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The workspace environment is designed as a universal platform that can handle multiple types of video production tasks through a single set of interaction principles. The same proxy manipulation mechanisms work for images, text, audio, and video elements, providing multi-functionality without requiring users to learn different interaction modes for each asset type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of manufacture

If machine learning models treat video as an indivisible entity, then the simplicity of processing is maintained, but the flexibility and creativity of video editing is limited

Engineering Contradiction:
Improveprocessing simplicityVSAvoidvideo editing flexibility
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent segments video content into distinct editable elements (characters, props, environments, actions, dialogue) that can be independently manipulated as separate proxy elements in the workspace. Each segment can be individually selected, modified, enhanced, or replaced using the same proxy-based interaction model, enabling fine-grained control over video components rather than treating the entire video as a single unit.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12136442B1System, method, and computer program for providing an interactive platform for video generation in which users are able to interact with machine-learning asset enhancement modules via proxy elements in a video production workspace
Publication Date: 2024.11.05 GOANIMATE INC
  • US12136442B1 patent drawing
  • US12136442B1 patent drawing
  • US12136442B1 patent drawing

AI summary

This disclosure relates to a system, method, and computer program for enabling an interactive process for video generation in which a user is able to guide the output of machine-learning asset enhancement modules to produce assets for a video. The system enables the user to interact with the asset enhancement modules via proxy elements in a video production workspace. The system provides a novel way to produce and edit video. The system enables any asset added to a video production workspace to serve as a proxy asset that a user can leverage to guide the output of one or more machine learning models trained to generate a type of multimedia asset. Specifically, a user can select an asset in the video production workspace, provide input on what attributes the user would like the asset to have, and the system then uses one or more machine-learning modules to generate an asset with the attributes requested by the user. The machine-generated asset may visually replace the selected asset at the same time and location in the video as the selected asset. A user is able to transform any asset to any other asset. Even a simple shape can serve as a proxy for a much more complex asset, even an asset of different multimedia type. A user may transform assets independently of each other and the video as a whole, or in conjunction with each other and the video as a whole.