Video editing system based on natural language and sketch

The video editing system addresses the steep learning curve of existing tools by enabling natural language and sketch-based editing, facilitating easier and more convenient video editing through AI-assisted command parsing and categorization.

WO2025143439A1PCT designated stage expired Publication Date: 2025-07-03KOREA ADVANCED INST OF SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/013968
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-09-13
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing video editing tools have a steep learning curve, making them difficult for beginners to use, and previous support services for expressing editing requests are limited in interaction options.

Method used

A video editing system that supports expression of editing requests through natural language text and video sketches, utilizing an AI model to parse and classify commands into temporal, spatial, and editing task references, and process them into predefined categories for easier video editing.

Benefits of technology

Reduces the difficulty of video editing for beginners by allowing intuitive interaction through natural language and sketches, enhancing user convenience and ease of use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024013968_03072025_PF_FP_ABST
    Figure KR2024013968_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A video editing method according to an embodiment of the present invention comprises: a preprocessing step; and a natural language parsing step of parsing an input natural language command by using a conversational natural language processing model and classifying same into any one of a temporal reference, a spatial reference that stores information on a natural language command capable of referring to the location or area of a video frame, an editing task reference that stores information on a natural language command capable of indicating a selectable editing task, and an editing parameter reference that stores information on a natural language command capable of referring to a specific parameter of a selected editing task.
Need to check novelty before this filing date? Find Prior Art

Description

A natural language and sketch-based video editing system

[0001] The present invention relates to a video editing system, and more particularly, to a system for editing a video based on natural language and image sketches.

[0002] Producing informational videos involves a complex process that includes planning, scripting, filming, and editing.

[0003] Especially in the editing stage of informational videos, careful composition of the video, removal of unnecessary parts, and integration of additional media assets are especially important. Popular commercial video editing tools provide all the necessary tools, but their steep learning curve can make them difficult for beginners to use.

[0004] In these automated editing tools or video generating systems, in order to provide convenience to users when editing videos, support services for expressing editing requests were previously provided. However, the support services provided previously only provide a limited set of interactions for expressing editing requests, and thus have limitations.

[0005] The present invention has been devised to solve the above problems, and an object of the present invention is to provide a video editing system that supports expression of a video editing request through natural language text and video sketch on a frame.

[0006] According to one embodiment of the present invention for achieving the above object, a natural language and sketch-based video editing method includes a preprocessing step of performing a preprocessing operation for extracting metadata in units of frames and preset clips from input video data; and a natural language parsing step of parsing an input natural language command using an interactive natural language processing model and classifying it into any one of a temporal reference in which information about a natural language command that can refer to a segment of a video is stored, a spatial reference in which information about a natural language command that can refer to a location or area of ​​a video frame is stored, an editing task reference in which information about a natural language command that can indicate a selectable editing task is stored, and an editing parameter reference in which information about a natural language command that can refer to a specific parameter of a selected editing task is stored.

[0007] And, according to one embodiment of the present invention, a natural language and sketch-based video editing method may further include a temporal interpretation step of interpreting natural language commands classified by temporal reference and classifying them into preset temporal categories and processing them.

[0008] Additionally, the predefined temporal categories may include a location category where information specifying a separate time code or an abstract time is stored, a script-based category where information directly or indirectly about the script of the video is stored, and a video-based category where information about a description indicating actions or visual descriptions related to the video is stored.

[0009] And the temporal interpretation step may include a step of selecting a preset number of related segments in order of high similarity by comparing them with the labels corresponding to "transcript" or "video" in the metadata extracted in units of frames and preset clips, and a step of performing a compilation task on all segments in the video data that match the selected related segments using an interactive natural language processing model.

[0010] In addition, a natural language and sketch-based video editing method according to one embodiment of the present invention may further include a spatial interpretation step of interpreting a natural language command classified into a spatial reference and classifying and processing it into a preset spatial category.

[0011] The predefined spatial categories may include a visual content-dependent category in which information specific to an object, element, or region within a video frame is stored, and a visual content-independent category in which information about a specific location or position relative to a frame is stored, but which does not vary depending on the visual content of the video.

[0012] Additionally, the spatial interpretation step may include, for a segment that includes information belonging to a visual content dependent category, a step of extracting a representative frame from the center of a time code range for the segment; a step of extracting instance crop data to be encoded into an embedding space from metadata of each frame constituting the segment; and a step of encoding the extracted instance crop data into the embedding space.

[0013] And, according to one embodiment of the present invention, a natural language and sketch-based video editing method may further include an editing task and editing parameter interpretation step of interpreting natural language commands classified into editing task references and editing parameter references and classifying them into preset editing categories and processing them.

[0014] Additionally, the preset edit categories may include explicit parameter categories that store information about directly specified parameters, relative parameter categories that store information about parameters that are adjusted based on the current settings, and abstract parameter categories that store information about general directives that do not have specific values ​​for associated parameters.

[0015] Meanwhile, according to another embodiment of the present invention, a natural language and sketch-based video editing system includes: a communication unit into which video data is input; and a processor that performs a preprocessing operation for extracting metadata in units of frames and preset clips from the input video data, parses the input natural language commands using an interactive natural language processing model, and classifies the input natural language commands into any one of a temporal reference in which information about a natural language command that can refer to a segment of a video is stored, a spatial reference in which information about a natural language command that can refer to a location or area of ​​a video frame is stored, an editing task reference in which information about a natural language command that can indicate a selectable editing task is stored, and an editing parameter reference in which information about a natural language command that can refer to a specific parameter of a selected editing task is stored.

[0016] As described above, according to embodiments of the present invention, by supporting expression of video editing requests through natural language text and video sketches on frames, even if a user participating in editing an informational video is a beginner, it is possible to reduce difficulties caused by the steep learning curve of a video editing tool and support video editing more easily and conveniently.

[0017] FIG. 1 is a drawing provided to explain the configuration of a video editing system according to one embodiment of the present invention;

[0018] FIG. 2 is a drawing provided for the description of a user interface provided by a video editing system according to one embodiment of the present invention;

[0019] FIG. 3 is a flowchart provided to explain a video editing method according to one embodiment of the present invention;

[0020] FIG. 4 is a drawing provided to further explain the preprocessing step of the video editing method according to one embodiment of the present invention.

[0021] FIG. 5 is a drawing provided to further explain the natural language parsing step of the video editing method according to one embodiment of the present invention.

[0022] FIG. 6 is a drawing provided to further explain the temporal interpretation step of the video editing method according to one embodiment of the present invention; and

[0023] FIG. 7 is a drawing provided to further explain the spatial interpretation step, the editing operation, and the editing parameter interpretation step among the video editing methods according to one embodiment of the present invention.

[0024] Hereinafter, the present invention will be described in more detail with reference to the drawings.

[0025] FIG. 1 is a diagram provided to explain the configuration of a video editing system according to one embodiment of the present invention.

[0026] The video editing system according to the present embodiment can enhance the natural language and video sketch interface through a backend computation pipeline that utilizes an AI model.

[0027] Referring to FIG. 1, a video editing system according to the present embodiment may include a communication unit (110), a processor (120), a storage unit (130), an input unit (140), and an output unit (150).

[0028] The communication unit (110) is equipped with a communication means that is connected to an external server or the like through a communication network, and can obtain video data to be edited from the outside.

[0029] The storage unit (130) is a storage medium that stores programs and data necessary for the processor (120) to operate.

[0030] The input unit (140) may be equipped with an input interface device that receives user input, such as a mouse or keyboard, and the output unit (150) is a display means that outputs information to be output by the processor (120) on the screen.

[0031] The processor (120) is provided to process all aspects of the video editing system.

[0032] Specifically, the processor (120) can perform a preprocessing task to extract metadata in units of frames and preset clips from video data acquired through the communication unit (110).

[0033] And the processor (120) can parse the input natural language command using an interactive natural language processing model (e.g., GPT-4, Generative Pre-trained Transformer 4) and classify it into one of a temporal reference, a spatial reference, an editing task reference, and an editing parameter reference.

[0034] Thereafter, the processor (120) can interpret natural language commands classified by temporal reference, classify them into preset temporal categories, and process them, interpret natural language commands classified by spatial reference, classify them into preset spatial categories, and then interpret natural language commands classified by editing task reference and editing parameter reference, classify them into preset editing categories, and process them.

[0035] At this time, the temporal reference may store information about a natural language command that can refer to a segment of a video, the spatial reference may store information about a natural language command that can refer to a location or area of ​​a video frame, the editing task reference may store information about a natural language command that can indicate a selectable editing task, and the editing parameter reference may store information about a natural language command that can refer to a specific parameter of a selected editing task.

[0036] FIG. 2 is a drawing provided for explaining a user interface provided by a video editing system according to one embodiment of the present invention.

[0037] The video editing system according to the present embodiment can be implemented as a multi-mode system that allows a user to edit a video using natural language and video sketches, and can provide a user interface such as the screen illustrated in FIG. 2.

[0038] Specifically, in the edit list panel (a), the user can layer video edits or edit specific segments (a2) by clicking the add button (a1).

[0039] And you can input editing commands using natural language in the editing description box (b1) in the editing description panel (b), and use the sketch pad (b2) to specify where to apply the editing contents within the video frame.

[0040] The inspection panel (c) can analyze the user's natural language prompt and indicate which parts of the user input correspond to the user's intended editing action, editing action parameters, spatial location, and temporal location.

[0041] The Parameters panel (e) allows the user to manually adjust the parameters and perform result editing operations (d).

[0042] The video player (f) displays the edited content overlaid on the video frame and allows the user to directly adjust the spatial position of the edited content by clicking and dragging.

[0043] The video control panel (g) allows users to quickly navigate through the video timeline (g2) as well as quickly navigate through the editing suggestions suggested by the system and accept or reject them (g1).

[0044] The timeline (h) can show the temporal location of each edit and edit suggestion, and the transcript (i) shows an alternative view of the temporal location of edits and edit suggestions, allowing the user to precisely match the timing of edits to the speaker's words.

[0045] FIG. 3 is a flowchart provided to explain a video editing method according to one embodiment of the present invention.

[0046] The video editing method according to the present embodiment can be executed by the video editing system described above with reference to FIGS. 1 and 2.

[0047] Specifically, the video editing system can perform a preprocessing task to extract metadata in units of frames and preset clips from video data acquired through the communication unit (110) using the processor (120) (S310).

[0048] The video editing system can use a processor (120) to parse an input natural language command using an interactive natural language processing model (e.g., GPT-4, Generative Pre-trained Transformer 4) and classify it into one of a temporal reference, a spatial reference, an editing task reference, and an editing parameter reference (S320).

[0049] Thereafter, the video editing system can interpret natural language commands classified by temporal reference, classify them into preset temporal categories, and process them (S330), interpret natural language commands classified by spatial reference, classify them into preset spatial categories, and process them (S340), and interpret natural language commands classified by editing task reference and editing parameter reference, classify them into preset editing categories, and process them (S350).

[0050] FIG. 4 is a drawing provided to explain in more detail a preprocessing step of a video editing method according to one embodiment of the present invention.

[0051] As described above, the video editing system can execute a preprocessing step (Stage 0) that performs a preprocessing task to extract metadata in units of frames and preset clips from acquired video data using a processor (120).

[0052] Specifically, the video editing system may perform multiple stages of preprocessing to extract frame-by-frame and clip-by-clip metadata from a given input video, which is used to infer video context during interaction, to facilitate various real-time system interactions.

[0053] For example, in the preprocessing stage, once frame-level metadata has been extracted, frames can be sampled from the video every second and automatic instance segmentation can be performed at two levels of granularity using the Segment Anything model. Subsequently, a threshold for segmentation can be set based on the size of the instance crop data to remove instance masks that are too small.

[0054] And in the preprocessing stage, once the metadata of the preset clip units is extracted, activity recognition for 10-second clips can be performed using OpenGVLab's video learning model InternVideo.

[0055] Additionally, in the preprocessing stage, an image captioning model (e.g., BLIP-2) can be used to generate text descriptions of video content, and the most prominent objects in the scene are listed every second at preset intervals (e.g., every 10 seconds). Then, GPT-3.5 can be used to summarize the resulting captions of clips at preset intervals (e.g., every 10 seconds) to remove redundant information about the visual content.

[0056] FIG. 5 is a drawing provided to further explain the natural language syntax analysis step of the video editing method according to one embodiment of the present invention.

[0057] The video editing system can execute a natural language parsing step (Stage 1) that parses input natural language commands (requests) using a conversational natural language processing model (e.g., GPT-4) and classifies them into one of a temporal reference, a spatial reference, an editing task reference, and an editing parameter reference.

[0058] As described above, the temporal reference stores information about a natural language command that can refer to a segment of a video, the spatial reference stores information about a natural language command that can refer to a location or region of a video frame, the editing operation reference stores information about a natural language command that can indicate a selectable editing operation, and the editing parameter reference stores information about a natural language command that can refer to a specific parameter of a selected editing operation.

[0059] And the video editing system, in the natural language parsing process, can classify the editing task closest to the input natural language command by utilizing GPT-4.

[0060] For example, an editing operation could be any of "text", "image", "shape", "blur", "cut", "crop", or "zoom".

[0061] FIG. 6 is a drawing provided to further explain the temporal interpretation step of a video editing method according to one embodiment of the present invention.

[0062] The video editing system can execute a temporal interpretation stage (Stage 2) that interprets natural language commands classified by temporal reference and processes them by classifying them into preset temporal categories.

[0063] Here, the preset temporal categories may include a location category where information specifying a separate time code or an abstract time (e.g., "intro", "ending") is stored, a script-based category where direct or indirect information about the script of the video is stored, and a video-based category where information about a description indicating an action or visual description related to the video is stored.

[0064] In the temporal interpretation step, if the metadata extracted in frame units and preset clip units contains a label corresponding to "transcript" or "video," a preset number (e.g., 10) of related segments are selected in order of high similarity compared to the label, and a conversational natural language processing model is used to perform a compilation task on all segments in the video data that match the selected related segments.

[0065] That is, in the temporal interpretation step, if the metadata extracted in units of frames and preset clips contains labels corresponding to “transcript” or “video,” the top 10 relevant segments with the highest similarity can be selected by comparing them with the corresponding labels, and a compilation task can be performed on all segments in the video data that match the 10 selected relevant segments using GPT-4.

[0066] FIG. 7 is a drawing provided to further explain the spatial interpretation step, the editing operation, and the editing parameter interpretation step among the video editing methods according to one embodiment of the present invention.

[0067] The video editing system can execute a spatial interpretation step (Stage 3) that interprets natural language commands classified by spatial reference and processes them by classifying them into preset spatial categories.

[0068] Predefined spatial categories may include visual content-dependent categories, which store information specific to objects, elements, or regions within a video frame, and visual content-independent categories, which store information about a particular location or position relative to a frame but that does not depend on the visual content of the video.

[0069] In the spatial interpretation step, for a segment that includes information belonging to a visual content dependent category, a representative frame is extracted from the center of the time code range for the segment, and instance crop data to be encoded into the embedding space is extracted from the metadata of each frame constituting the segment, so that the extracted instance crop data can be encoded into the embedding space.

[0070] At this time, in the spatial interpretation step, only text references and instance crop data can be encoded into the embedding space if there is no user-provided image sketch.

[0071] And, in the spatial interpretation step, the spatial location within the frame that most closely aligns with the parsed command and / or image sketch can be determined by extracting the instance crop data that has the highest cosine similarity to the user input.

[0072] Additionally, in the spatial interpretation step, if there is no information corresponding to a visual content-dependent category to be encoded into the embedding space, if a video sketch is provided, the provided video sketch can become a candidate reference to be encoded into the embedding space.

[0073] If both the information corresponding to the visual content-dependent category and the image sketch are unavailable, the entire frame can be encoded into the final spatial location.

[0074] And in the spatial interpretation stage, if the natural language command contains information corresponding to a visual content-independent category, GPT-4 can be used to segment and resize the region of interest based on the command and frame boundaries.

[0075] Meanwhile, the video editing system can execute an editing task and editing parameter interpretation step (Stage 4) that interprets natural language commands classified into editing task references and editing parameter references after the spatial interpretation step is executed, classifies them into preset editing categories, and processes them.

[0076] Predefined edit categories can include explicit parameter categories, which store information about directly specified parameters (e.g., "12px" or "intro"), relative parameter categories, which store information about parameters that are adjusted relative to the current setting (e.g., "5 seconds longer" or "10% less"), and abstract parameter categories, which store information about general directives that do not have specific values ​​for associated parameters (e.g., "shorter" or "longer").

[0077] In the editing task and editing parameter interpretation phase, we can provide instructions that guide the creation process, the corresponding context, and related video content, with a particular focus on tasks involving text and images.

[0078] That is, for text, it can help determine the text display within the video, and for images, it can help create appropriate search terms to extract the images needed for the video.

[0079] Meanwhile, it goes without saying that the technical idea of ​​the present invention can also be applied to a computer-readable recording medium containing a computer program that performs the functions of the device and method according to the present embodiment. In addition, the technical idea according to various embodiments of the present invention can be implemented in the form of computer-readable code recorded on a computer-readable recording medium. The computer-readable recording medium can be any data storage device that can be read by a computer and store data. For example, the computer-readable recording medium can be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical disk, a hard disk drive, etc. In addition, the computer-readable code or program stored on the computer-readable recording medium can be transmitted through a network connected between computers.

[0080] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.

Claims

1. A preprocessing step for performing preprocessing work to extract metadata of frame units and preset clip units from input video data; and A natural language and sketch-based video editing method, comprising: a natural language parsing step of parsing an input natural language command using an interactive natural language processing model, and classifying the natural language command into any one of a temporal reference in which information about a natural language command that can refer to a segment of a video is stored, a spatial reference in which information about a natural language command that can refer to a location or area of ​​a video frame is stored, an editing task reference in which information about a natural language command that can indicate a selectable editing task is stored, and an editing parameter reference in which information about a natural language command that can refer to a specific parameter of a selected editing task is stored.

2. In claim 1, A natural language and sketch-based video editing method, characterized by further comprising a temporal interpretation step of interpreting natural language commands classified by temporal reference and classifying and processing them into preset temporal categories.

3. In claim 2, The predefined temporal categories are: A natural language and sketch-based video editing method characterized by including a location category in which information specifying a separate time code or an abstract time is stored, a script-based category in which direct or indirect information about a script of a video is stored, and a video-based category in which information about a description indicating actions or visual descriptions related to the video is stored.

4. In claim 2, The temporal interpretation step is, If the metadata extracted in frame units and preset clip units contains a label corresponding to "transcript" or "video", a step of selecting a preset number of related segments in order of high similarity by comparing them with the corresponding label; and A natural language and sketch-based video editing method, characterized by comprising the step of performing a compilation task on all segments in video data that match the selected relevant segments using a conversational natural language processing model.

5. In claim 1, A natural language and sketch-based video editing method, characterized by further comprising a spatial interpretation step of interpreting natural language commands classified into spatial references and classifying them into preset spatial categories and processing them.

6. In claim 5, The predefined space categories are: A natural language and sketch-based video editing method characterized by including a visual content-dependent category in which information specific to an object, element or region within a video frame is stored, and a visual content-independent category in which information about a location or position relative to a frame is stored, but which does not vary depending on the visual content of the video.

7. In claim 6, The spatial interpretation step is, For a segment containing information belonging to the visual content dependent category, a step of extracting a representative frame from the center of the time code range targeting the segment; A step of extracting instance crop data to be encoded into the embedding space from the metadata of each frame composing the segment; and A natural language and sketch-based video editing method, characterized by including a step of encoding extracted instance crop data into an embedding space.

8. In claim 1, A natural language and sketch-based video editing method, characterized in that it further includes an editing task and editing parameter interpretation step of interpreting natural language commands classified into editing task references and editing parameter references and classifying them into preset editing categories and processing them.

9. In claim 8, The preset editing categories are: A natural language and sketch-based video editing method, characterized in that it includes an explicit parameter category in which information about directly specified parameters is stored, a relative parameter category in which information about parameters that are adjusted based on current settings is stored, and an abstract parameter category in which information about general directives without specific values ​​of associated parameters is stored.

10. A communication unit into which video data is input; and A natural language and sketch-based video editing system, comprising: a processor for performing preprocessing for extracting metadata of frames and preset clips from input video data, parsing input natural language commands using an interactive natural language processing model, and classifying them into any one of a temporal reference in which information on natural language commands that can refer to segments of a video is stored, a spatial reference in which information on natural language commands that can refer to locations or regions of video frames is stored, an editing task reference in which information on natural language commands that can indicate selectable editing tasks is stored, and an editing parameter reference in which information on natural language commands that can refer to specific parameters of selected editing tasks is stored.

Citation Information

Patent Citations

  • Video management system

    KR1020140079775A

  • Shot boundary detection method and apparatus using multi-classing

    KR102285039B1

  • Video classification method, information processing method, and server

    KR102392943B1

  • Apparatus for retrieval of semantic monent in video and method using the same

    KR102591314B1

  • Dictation that allows editing

    US20170263248A1