Multimodal Video Effect Rendering for Automated Content Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video effect processing methods are cumbersome and limited, often resulting in suboptimal enhancement outcomes that fail to enhance the visual and auditory experience of packaged videos.

Innovation Solution

A method involving the extraction and encoding of text, audio, and video frame sequences to generate feature information, followed by effect enhancement inference and rendering to create a more comprehensive and enhanced video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If manual effect addition is used for video processing, then the process is simple to implement, but the effect enhancement outcome is limited and complicated

Engineering Contradiction:
Improveease of video processingVSAvoideffect enhancement outcome
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent replaces manual mechanical effect addition with an automated AI-based system. The AI model automatically analyzes video content, generates appropriate effects, and applies them without human intervention, thereby maintaining ease of operation while significantly improving effect enhancement outcomes and versatility

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service video processing where the AI model autonomously performs effect selection, generation, and application based on video content analysis. The system serves itself by automatically determining the best effects to apply without requiring user guidance or manual configuration

Inventive Principle:
Principle #25Self-service

2Device complexity

If single effect type processing is used, then the processing method is simple, but the visual and auditory experience is not enhanced sufficiently

Engineering Contradiction:
Improveprocessing method complexityVSAvoidvisual and auditory experience enhancement
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent implements a universal processing system that can handle multiple effect types including visual effects, audio effects, and combined effects. The AI model is designed to universally analyze video content and generate appropriate multi-type effects, thereby enhancing visual and auditory experience while maintaining a unified simple processing interface

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges multiple effect generation capabilities into a single integrated processing pipeline. The AI model combines visual effect generation, audio effect generation, and effect coordination into one unified process, allowing simultaneous enhancement of both visual and auditory experiences through a single processing operation

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250378612A1Video processing method, device and storage medium
Publication Date: 2025.12.11 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20250378612A1 patent drawing
  • US20250378612A1 patent drawing
  • US20250378612A1 patent drawing

AI summary

Embodiments of the present disclosure provide a video processing method, device and storage medium. The method includes: extracting text content, audio content and a video frame sequence included in an original video; encoding the text content, the audio content and the video frame sequence to obtain text feature information, audio feature information and video frame feature information, respectively; performing effect enhancement inference on the original video to obtain effect enhancement description information according to the text feature information, the audio feature information and the video frame feature information, where the effect enhancement description information includes an effect enhancement position description and a corresponding effect enhancement element description; and performing effect rendering on the original video using the effect enhancement description information to obtain an effect enhanced video of the original video.