Video Editing Component Embeddings Using Cross-Attention Guidance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video representation learning methods struggle to encode information from diverse editing components due to their lack of clear semantics, subjective action, or context, and existing datasets fail to support research on learning universal representations for major types of video editing components.

Innovation Solution

A novel embedding guidance architecture and a large-scale video editing components dataset are used to train a machine learning model, utilizing a first sub-model with spatial and temporal encoders and a second sub-model with cross-attention mechanisms, along with a contrastive learning loss and dynamic embedding queues to distinguish editing components from raw materials.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing video representation learning methods are used, then general video processing is possible, but they struggle to encode information from diverse editing components due to lack of clear semantics

Engineering Contradiction:
Improveencoding precision of editing componentsVSAvoidadaptability to diverse editing component types
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments video editing components into six distinct categories (video effect, animation, transition, filter, sticker, text) and develops specialized encoding mechanisms for each type. The editing component recognition model includes type-specific sub-models that process different editing component categories separately, enabling precise encoding while maintaining adaptability across diverse types.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal editing component recognition framework that can handle multiple types of editing components through a single system. The model architecture includes shared layers and type-specific layers that work together, allowing the system to process various editing component types (video effects, animations, transitions, filters, stickers, text) using a unified approach while maintaining specialized处理能力 for each type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If pixel-level supervision is used for training, then model accuracy improves, but it requires extensive labeled data and increases complexity

Engineering Contradiction:
Improvemodel accuracy in distinguishing editing componentsVSAvoidtraining complexity and data requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary contrastive learning mechanism that operates at the editing component level rather than requiring pixel-level annotations. The contrastive loss function uses editing component embeddings as intermediaries to guide the recognition model, enabling accurate differentiation without the need for extensive pixel-level labeled data, thus reducing training complexity while maintaining model accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If a large-scale dataset with atomic editing components is created, then research on single editing components is enabled, but data generation and processing complexity increases

Engineering Contradiction:
Improveprecision of atomic editing component separationVSAvoidease of dataset creation and processing
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent applies preliminary action by pre-processing videos to isolate and separate editing components before creating the dataset. The editing component separation model performs atomic decomposition of editing components from raw videos in advance, generating a structured dataset where each editing component is individually identified and categorized. This preliminary separation simplifies subsequent research and analysis while maintaining high precision in component identification.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12394445B2Generating representations of editing components using a machine learning model
Publication Date: 2025.08.19 LEMON INC(GB)
  • US12394445B2 patent drawing
  • US12394445B2 patent drawing
  • US12394445B2 patent drawing

AI summary

The present disclosure describes techniques for generating representations of editing components using a machine learning model. Images and guidance tokens are input into a first sub-model of the machine learning model. The machine learning model is trained to distinguish the editing components from raw materials and generate the representations of the editing components. Tokens corresponding to the images are generated by the first sub-model based on the images and the guidance tokens. The tokens corresponding to the images and the guidance tokens are input into a second sub-model of the machine learning model. The second sub-model comprises a cross-attention mechanism. An embedding indicative of at least one editing component is generated based on the tokens corresponding to the images and the guidance tokens by the second sub-model.