Real-time interactive feedback system for partial re-rendering of initially generated ai videos

WO2026197454A1PCT designated stage Publication Date: 2026-09-24KIM HONG JEON
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/003603
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2026-09-24

Smart Images

  • Figure KR2025003603_24092026_PF_FP_ABST
    Figure KR2025003603_24092026_PF_FP_ABST
Patent Text Reader

Abstract

The invention proposes an innovative interactive feedback system enabling users of artificial intelligence (AI)-based video platforms to selectively and partially re-render desired segments or objects of initially generated AI videos in real-time without regenerating the entire video. Unlike existing systems, this invention significantly reduces GPU resource consumption and enhances user convenience, speed, and video quality by enabling immediate, intuitive modifications via natural language inputs. The invention fundamentally addresses inefficiencies in existing AI video generation technologies and is expected to contribute significantly to the industrialization and popularization of AI video technologies.
Need to check novelty before this filing date? Find Prior Art

Description

REAL-TIME INTERACTIVE FEEDBACK SYSTEM FOR PARTIAL RE-RENDERING OF INITIALLY GENERATED AI

[0001] The present invention relates to the field of artificial intelligence (AI)-based video generation systems, specifically focusing on interactive feedback systems enabling selective and partial real-time re-rendering of specific video segments based on natural language user requests.

[0002] Recent advancements in AI have enabled video generation based on user-provided text prompts, exemplified by platforms such as OpenAI's "Sora" and Google DeepMind's "Veo 2". However, these existing systems inherently require users to repeatedly revise their initial prompts and regenerate entire videos if initial outputs do not match user intentions. Such repetitive regeneration consumes excessive GPU resources, prolongs video production, and limits the immediate and precise reflection of user-requested modifications. Similarly, existing post-production editing technologies, like Adobe’s natural language-based editing solutions, provide limited editing functionalities on finalized videos without enabling instantaneous, detailed adjustments.

[0003] The existing AI video generation and editing systems require complete video regeneration or provide limited editing capabilities after initial generation, leading to unnecessary GPU resource consumption, delays, and difficulties in accurately and immediately reflecting detailed modifications requested by users

[0004] The present invention provides a real-time interactive feedback system for partial re-rendering within AI-based video platforms. It allows users to directly and intuitively select specific video segments or objects within initially generated videos via a timeline interface and enter natural language modification requests. The system immediately analyzes user inputs, selectively re-renders only the specified segments or objects, and instantly presents modified results, maintaining overall video continuity.

[0005]

[0006] The system comprises:

[0007] - A Natural Language Processing (NLP) module, analyzing user natural language requests in real-time to precisely identify video frames, objects, or visual elements requiring modifications.

[0008] - An Interactive Feedback Processing Engine (IFPE), converting NLP-analyzed inputs into specific modification commands related to video attributes such as brightness, color, facial expression, or object shape.

[0009] - A Video Selective Partial Re-rendering Module, performing selective and immediate re-rendering of specified frames or objects according to modification requests.

[0010] This invention significantly reduces unnecessary GPU resource usage by avoiding repetitive regeneration of the entire video. It enhances user experience, providing immediate, precise reflection of user requests, thus greatly improving efficiency, accuracy, and user satisfaction in AI-generated video production and editing processes.

[0011] Figure 1 is a block diagram illustrating an overview of the system according to the present invention, describing the overall configuration and process in which a user reviews an initially generated AI video and performs selective partial re-rendering of specific video parts through real-time natural language-based interactive feedback.

[0012]

[0013] Figure 2 is a process flowchart illustrating the real-time interactive feedback process of the present invention, detailing an interaction sequence where the user selects a specific video segment on a timeline from the initially generated AI video, enters a re-rendering request via natural language, the system immediately analyzes the request, performs partial re-rendering, and provides the modified results to the user in real-time.

[0014]

[0015] Figure 3 is an embodiment of independent claim 1, illustrating a process where a user requests partial re-rendering of specific video segments or objects using multimodal inputs such as text, image, voice, or video. Additionally, it includes the process where, based on an initially provided natural language request, the AI automatically and repeatedly performs creative and randomized remixes on the already partially re-rendered video without further natural language inputs from the user, generating diverse video variations reflecting the user's intentions (refer to dependent claims 1-6).

[0016] The best mode for carrying out the invention involves:

[0017] 1. Users viewing initially generated AI videos and selecting segments or objects needing modification via an intuitive timeline interface.

[0018] 2. Entering natural language commands such as "Make the character’s expression serious" or "Change background to darker sunset."

[0019] 3. Real-time NLP analysis and immediate IFPE command generation, leading to selective, instant partial re-rendering of specific video segments.

[0020] Detailed examples include changing specific objects (e.g., converting a character's weapon from a gun to a sword) instantly through user natural language inputs. Furthermore, preliminary AI-generated modification suggestions can be automatically provided based on the user's segment selection, enabling rapid and intuitive application of chosen modifications.

[0021] This invention significantly improves the industrial applicability and practicality of AI-based video production by reducing production costs and GPU resource usage, thereby making AI-generated video technologies more accessible and efficient for widespread industrial and consumer applications.

[0022] Not applicable

Claims

1.[Independent Claim]1. A real-time interactive feedback method for partial re-rendering in an artificial intelligence (AI)-based video platform, comprising:- allowing a user to view an initially generated AI video in real-time and directly select a specific video segment or object within the video via a timeline that requires re-rendering;- inputting detailed modification requests for the selected part through natural language;- analyzing the user's natural language input in real-time to precisely identify the specific segment or object within the video;- selectively partially re-rendering only the identified segment or object while maintaining the remainder of the video data unchanged from its initially generated state to preserve overall video continuity, thereby effectively reducing unnecessary GPU resource usage and immediately reflecting the user's requests in the re-rendered video (Refer to Figures 1 and 2).Dependent Claim 1.The method according to claim 1, wherein the natural language feedback process comprises:- a natural language processing (NLP) module that analyzes the user's natural language input to identify specific video frames, time intervals, objects, or visual elements requiring modification;- an Interactive Feedback Processing Engine (IFPE) that converts results from the NLP analysis into specific video attribute modification commands, including brightness, color, object shape, facial expression, and background, and transmits these commands to the video generation and selective partial re-rendering module (Refer to Figures 1 and 2).Dependent Claim 2.The method according to claim 1, wherein interactions between the user and the system may include casual conversations or general dialogue topics beyond specific video modification requests, and wherein partial re-rendering functionality can be immediately invoked if a relevant need for video re-rendering is identified during such general dialogue.Dependent Claim 3.The method according to claim 1, wherein the user's input may be provided in multimodal forms, including but not limited to text, images, audio, or video, and wherein the system comprehensively analyzes these multimodal inputs to accurately identify user intentions and perform selective re-rendering of specified parts accordingly.Dependent Claim 4.The method according to claim 1, wherein upon a user's request for remixing the selectively re-rendered video result without additional natural language input, the system repeatedly and autonomously generates randomized or creative variations solely on the same video segment or object initially specified by the user's previous natural language input, continuously providing diverse video outcomes aligned with the user's original modification intent.Dependent Claim 5.The method according to claim 1, wherein after the user selects a specific video segment or object but before providing explicit natural language modification requests, the AI-based system preliminarily analyzes the user's selection to generate pre-suggested potential modification outcomes, presents these suggestions to the user, and upon the user's selection of any pre-suggested outcome, immediately applies the chosen result through selective partial re-rendering.Dependent Claim 6.The method according to claim 1, wherein the system continuously accumulates and learns from interactive natural language inputs and partial re-rendering interaction data between the user and the system, subsequently enhancing re-rendering speed and accuracy for similar video segments by referencing accumulated data and previous rendering results when identical or similar natural language inputs are received in the future.