Video editing method, system, device, medium and program product
By receiving editing intent information in natural language within video editing tools and performing semantic recognition, the system automatically performs video adjustments, solving the problems of complexity and inefficiency of existing tools and enabling efficient video editing for non-professional users.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-12
- Publication Date
- 2026-03-13
AI Technical Summary
Existing video editing tools are complex in design, requiring users to spend a lot of time learning and operating them. They are inefficient in editing, have limited automation functions, and cannot meet the fast editing needs of non-professional users.
A video editing method is provided that receives editing intent information in natural language on the editing page, uses semantic recognition technology to automatically identify the user's editing intent, performs corresponding adjustment operations, and generates the target video, thus simplifying the editing process.
Users can automatically adjust videos by expressing their editing intentions without any professional skills, which improves editing efficiency, simplifies the operation process, and enhances the quality of video editing.
Smart Images

Figure CN121665066A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video processing technology, and in particular to a video editing method, a video editing system, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] With the explosive growth of short video and live streaming platforms, the demand for video editing is booming among individual creators, businesses, students, and non-professional users. For example, businesses need to efficiently and quickly produce product promotional videos that highlight product features; dancers want to showcase their dance performances through short videos and quickly edit out exciting clips; teachers and science popularization bloggers need to create short videos for teaching or demonstration content to better disseminate knowledge; and movie fans want to compile and share classic or high-energy clips from movies with their audiences.
[0003] However, all of the above videos require users to edit them themselves. Most existing professional video editing tools are complex in design, requiring users to spend a significant amount of time learning and operating them. While some editing tools offer automation features, these are usually limited to simple cropping and splicing. If users want more advanced effects, they often need to perform the editing manually, resulting in low video editing efficiency. Summary of the Invention
[0004] One objective of this invention is to provide a video editing method to improve video editing efficiency. The specific technical solution is as follows: In a first aspect of this invention, a video editing method is provided, comprising: acquiring a pre-edited video for display on an editing page, the editing page including an input control; receiving editing intent information, the editing intent information responding to a trigger determination of the input control; determining an adjustment operation through semantic recognition of the editing intent information; adjusting the pre-edited video according to the adjustment operation to generate a target video; and sending the target video to display the target video on the editing page.
[0005] In a second aspect of the present invention, a video editing method is also provided, comprising: displaying a pre-edited video on an editing page, the editing page including an input control; acquiring editing intent information in response to triggering the input control; sending the editing intent information to adjust the pre-edited video based on the editing intent information, the adjustment being performed based on an adjustment operation determined by semantic recognition of the editing intent information; and receiving and displaying the adjusted target video on the editing page.
[0006] In a third aspect of this invention, a video editing system is also provided, comprising: a server and a client; wherein the server acquires a pre-edited video for display on an editing page; receives editing intent information; determines an adjustment operation through semantic recognition of the editing intent information, adjusts the pre-edited video according to the adjustment operation to generate a target video; and sends the target video for display on the editing page; the client displays the pre-edited video on the editing page, the editing page including an input control; acquires editing intent information in response to triggering the input control; sends the editing intent information; and receives the adjusted target video and displays it on the editing page.
[0007] In another aspect of the present invention, an electronic device is also provided, characterized in that it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to execute the program stored in the memory to implement the steps of the video editing method as described in the embodiments of the present invention.
[0008] In another aspect of the present invention, a computer-readable storage medium is also provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform any of the video editing methods described above.
[0009] In another aspect of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the video editing methods described above.
[0010] The video editing method provided in this invention, after pre-editing a video, receives editing intent information, which is editing intent expressed as natural language text. Adjustment operations are determined through semantic recognition of the editing intent information; that is, the user only needs to express their editing ideas for the automatic semantic recognition and determination of corresponding adjustment operations. Then, the pre-edited video is adjusted according to the adjustment operations. No professional editing skills are required to complete the adjustment of the pre-edited video, resulting in the target video. The user only needs to express their editing intent for the automatic generation of editing quality and execution of editing processing, which improves video editing efficiency compared to manual editing. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0012] Figure 1 This is a flowchart illustrating the steps of an embodiment of the video editing method of the present invention; Figure 2 This is a flowchart illustrating the steps of a video pre-editing example in an embodiment of the present invention; Figure 3 This is a flowchart illustrating an example of extracting video segments in an embodiment of the present invention; Figure 4 This is a flowchart illustrating the steps of splicing and generating a pre-edited video in an embodiment of the present invention. Figure 5 This is a flowchart illustrating another example of video pre-editing in an embodiment of the present invention; Figure 6 This is a flowchart illustrating the steps of a dialogic adjustment example for a pre-edited video in an embodiment of the present invention. Figure 7 This is a flowchart illustrating an example of generating editing instructions based on editing intent information in an embodiment of the present invention. Figure 8 This is a flowchart illustrating an example of identifying editing intent information in an embodiment of the present invention; Figure 9 This is a flowchart illustrating the steps for locating the segment to be adjusted in an embodiment of the present invention. Figure 10 This is a flowchart illustrating the steps of a dialogic fine-tuning example in an embodiment of the present invention; Figure 11 This is a flowchart illustrating the steps of another embodiment of the video editing method of the present invention; Figure 12 This is a flowchart illustrating the steps of adjusting a segment to be edited in an embodiment of the present invention; Figure 13 This is an interactive schematic diagram of the editing system in an embodiment of the present invention; Figure 14 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0013] The technical solutions of the present invention will now be described with reference to the accompanying drawings in the embodiments of the present invention.
[0014] This invention provides a video editing method that displays an editing page in a browser, application (APP), or other client. The editing page offers conversational editing functionality. For a pre-edited video, the user inputs their editing intent via text, voice, or other dialogue methods. During server-side editing, the system semantically recognizes the user's intent and performs corresponding adjustments to the pre-edited video. Through one or more dialogues, the pre-edited video can be adjusted one or more times to obtain the target video. The user only needs to express their editing intent for the editing quality to be automatically generated and adjustments performed, significantly improving video editing efficiency compared to manual editing.
[0015] Reference Figure 1 This diagram illustrates a flowchart of an embodiment of a video editing method according to this application. Applied to the server side, it obtains the user's editing intent through interaction with the client and performs editing processing, thereby improving editing efficiency.
[0016] Step 102: Send the pre-edited video.
[0017] Pre-edited videos are videos that have been edited in advance based on predetermined goals. These goals can be determined based on editing requirements, such as the purpose or style of the editing. Pre-edited videos can be created by editing and splicing one or more videos.
[0018] Step 104: Receive editing intent information.
[0019] After video pre-editing is completed, fine-tuning can be performed via dialogue, which can be done through text, voice, or other means. The system receives corresponding editing intent information, such as text input from the user or voice input, and converts the voice data into text information, which serves as the editing intent information. This editing intent information represents the user's needs for video adjustments. The editing intent information can be expressed in natural language, such as "shorten the first five seconds," "this shot has a warmer tone," "add slow motion here," "slow down the subtitles," or "brighten the person in the red shirt."
[0020] Step 106: Determine the adjustment operation by semantic recognition of the editing intent information, and adjust the pre-edited video according to the adjustment operation to generate the target video.
[0021] The system performs natural language understanding on the received editing intent information to identify the semantic information it conveys and determine the adjustments the user intends to make to the pre-edited video, such as adjusting brightness, adding effects, deleting segments, or speeding up or slowing down playback. The corresponding adjustment is then executed on the pre-edited video. Adjustments to the pre-edited video can be made to the entire video (e.g., brightening, adding filters) or to a specific segment within the video.
[0022] During the fine-tuning process after pre-editing the video, one or more editing intent messages may be received. Based on the received editing intent messages, one or more adjustments are performed on the pre-edited video. After confirming that the fine-tuning process has been completed, the adjusted segments replace the segments to be adjusted before fine-tuning, generating the target video.
[0023] Step 108: Send the target video to display the target video on the editing page.
[0024] The target video is sent to the user's terminal device, and the target video is displayed on the editing page of the terminal device to complete the video editing process.
[0025] In summary, after a pre-edited video is received, editing intent information is received. This editing intent information is expressed as natural language text. Adjustment operations are determined through semantic recognition of this intent information. In other words, the user only needs to express their editing ideas for the system to automatically recognize the semantics and determine the corresponding adjustment operations. Then, the pre-edited video is adjusted according to these operations. No professional editing skills are required to complete the adjustment of the pre-edited video and obtain the target video. The system automatically generates editing quality and performs editing processing simply by the user expressing their editing intent, significantly improving video editing efficiency compared to manual editing.
[0026] Based on the above embodiments, the video editing of this invention includes multiple stages: a pre-editing stage, a dialog-based fine-tuning stage, and a multi-platform standardized export stage. The pre-editing stage is used to perform preliminary editing of the video to be edited according to an editing strategy. The dialog-based fine-tuning stage is used to fine-tune the pre-edited video in a dialog-based manner, refining the granularity of the video editing and improving the quality of the video editing. The multi-platform standardized export stage is used to output the edited target video according to the format requirements of each platform, facilitating the publication of the target video on the corresponding platform.
[0027] For the pre-editing stage: In this embodiment of the invention, the user uploads one or more videos to be edited on the editing page. The editing page provides a pre-editing function with a selection control offering multiple editing strategies. Each editing strategy is a set of video editing methods, and different types of videos can correspond to different editing strategies. The editing processing end performs pre-editing based on the videos to be edited and the editing strategies.
[0028] In one alternative embodiment, such as Figure 2 As shown, the pre-editing step includes the following steps: Step 202: Receive the video to be edited and the editing strategy.
[0029] Step 204: Detect the video to be edited and extract multiple video segments that conform to the editing strategy.
[0030] The process involves a global analysis and inspection of the video to be edited, dividing it into multiple video segments. This segmentation can be based on various methods such as shot composition and object selection, and multiple video segments are then selected according to editing strategies. These editing strategies control the video's rhythm, narrative order, transitions, sound effects, and more. For example, for product introduction videos, the editing strategy emphasizes product-related content, thus selecting video segments highlighting the product's main features. Similarly, for sports highlight videos, the editing strategy focuses on selecting high-profile moments, such as goal-scoring moments or high-difficulty stunts. The editing strategies in this embodiment include various types, such as deduplication strategies, highlight extraction strategies, and template editing strategies. Among them, the deduplication editing strategy is suitable for scenarios such as dance demonstrations and teaching demonstrations, automatically removing duplicate or similar shots; the highlight extraction strategy is suitable for editing sports events, high-energy movie clips, etc., automatically extracting visually impactful and fast-paced scenes; the template editing strategy has preset multiple editing style templates, such as fast-paced editing templates, slow-paced continuous editing templates, and brand promotion intro style templates, which are suitable for content such as voice-overs and brand promotions.
[0031] Step 206: Splice the video segments together to generate a pre-edited video.
[0032] The selected video clips can be spliced together, and special effects can be added, such as filters or scene transition effects, to generate corresponding pre-edited videos.
[0033] This allows for the rapid pre-editing of videos to be edited using editing strategies, generating a pre-edited video that initially meets the requirements, thus improving video editing efficiency.
[0034] In one alternative embodiment, such as Figure 3As shown, step 204 involves detecting the video to be edited and extracting multiple video segments that conform to the editing strategy, including the following sub-steps: Sub-step 2042: Detect the video to be edited and extract video segments.
[0035] The video to be edited is inspected, segmented according to relevant rules, and video clips are extracted. For example, video clips can be segmented according to scenes, characters, shots, etc.
[0036] The process of detecting and extracting video segments from the video to be edited includes: performing shot boundary detection on the video to be edited, and extracting video segments according to the start and end frames of each shot. This involves detecting video frames in the video to be edited, determining the video frames for different shot transitions based on the content of the video frames, thereby identifying the start and end frames of each shot, and extracting the video segments between the start and end frames. Shot boundary detection can also be performed using an algorithm model. In one example, the video to be edited is input into a Convolutional Neural Network (CNN) model, which detects the video to be edited and outputs the start and end frames corresponding to each shot. A shot detection algorithm based on a CNN can be used to automatically identify the start and end frames of each shot. For example, a Deep Residual Network (ResNet) model and / or a Long Short-Term Memory (LSTM) model can be used for shot recognition.
[0037] When extracting video clips, keyframes (I-frames) can be used as starting frames. An I-frame, also known as an intrapicture, is typically the first frame of each Group of Pictures (GOP). This allows for subsequent editing, sorting, and splicing based on the keyframes.
[0038] Sub-step 2044: Determine the tag of the video segment.
[0039] Video segments are detected, and tags are determined for each segment. The tags represent the content category of the video segment, describe its characteristics, and each video segment may correspond to one or more tags.
[0040] The tags include at least one of the following types: action, target object, emotion, scene, text, and repetition. Action tags describe the actions of moving objects in the video clip, such as playing ball, running, turning, etc. Target object tags describe target objects in the video clip, such as people, animals, objects, etc. Emotion tags describe the emotions and atmosphere expressed in the video clip. Scene tags describe the scene and background of the video clip, such as indoors, outdoors, night scene, slow motion, etc. Text tags describe the text associated with the video clip; this text can be described through subtitles or other text descriptions in the video clip, or through audio, background music, etc., and audio-related content can be recognized as text. Repetition tags describe the repetition of video clips; for example, multiple ball-playing clips can be tagged with a repetition tag, and repetitive video clips can be tagged with the same identifier, with different repetition tags for different clips.
[0041] Content recognition is performed on the video segments, and tags are determined based on the recognition results. Content recognition of the video segments can be achieved through various detection methods, such as image classification detection, object detection, action recognition, object tracking, emotion recognition, and scene recognition. Tags are then added to the video segments based on the recognition results. In one example, the video segments are input into a large-scale artificial intelligence (AI) model, and semantic tags are added to each video segment using various detection methods. For example, spatiotemporal action recognition models such as the SlowFast model and the Inflated 3D (I3D) model based on 3D editing are used to quantify the intensity and trajectory of actions of people or objects in the video, obtaining corresponding tags. For example, classification and detection models like YOLO (You Only Look Once), such as YOLOv8, and multi-object tracking models like DeepSORT (Deep Simple Online and Realtime Tracking), can track and label all people and objects in a video in real time, generating corresponding tags. Through global video feature extraction and scene recognition algorithms, such as combining Vision Transformer models with time-series clustering, the system can automatically determine the scene category of video clips, such as sports, meetings, or travel. This provides a basis for subsequent video clip selection.
[0042] Audio semantic understanding: Combining Automatic Speech Recognition (ASR) and Natural Language Processing (NLP) technologies, keywords and semantic emotions are extracted from video audio. Corresponding editing strategies can be set up based on semantics, such as switching shots according to the rhythm of dialogue or adjusting the editing density according to emotional fluctuations.
[0043] One type of label is shown in Table 1 below:
[0044] Table 1 In this embodiment of the invention, each video segment may correspond to one or more tags. For example, the tags for video segment A may include basketball shooting, passion, repeated shooting, red jersey, etc. For two or more video segments with the same tag, similarity detection can be performed, and duplicate tags can be set based on the similarity value.
[0045] Sub-step 2046: Filter tags according to the editing strategy and extract multiple video clips whose tags match the editing strategy.
[0046] Tags are filtered based on editing strategies. For example, a deduplication strategy will not filter duplicate tags, a highlight extraction strategy will filter action tags (e.g., for sports highlight videos, tags for jumping and running are primarily selected), and a brand promotional intro style template strategy will filter text tags (e.g., for product introduction videos, tags containing keywords related to the product and its features are selected). The corresponding video clips for each tag are then extracted.
[0047] It can divide the video to be edited into multiple video segments and set tags, and select appropriate video segments according to the editing strategy to improve editing efficiency.
[0048] In one alternative embodiment, such as Figure 4 As shown, step 206 involves splicing the video segments to generate a pre-edited video, including the following sub-steps: Sub-step 2062: Sort the video segments according to the editing strategy, and splice the video segments according to the sorting results.
[0049] Sub-step 2064: Add transition effects between two adjacent video clips to generate a pre-edited video.
[0050] After selecting video clips according to the editing strategy, the clips can be sorted according to that strategy. In one example, the clips can be sorted by the timestamps of their start and / or end frames. Other examples sort them by the editing template corresponding to the editing strategy. Alternatively, the sorting order can be determined based on factors such as tempo, content relevance, etc., corresponding to the editing strategy.
[0051] The video clips are spliced together according to the sorting results, and transition effects such as fade-in / fade-out, jump cuts, and slides can be added between adjacent video clips to generate corresponding pre-edited videos.
[0052] In this embodiment of the invention, the editing strategy includes at least: deduplication editing, highlight extraction, and template editing. The goal of deduplication editing is to retain the most representative content with high movement diversity, automatically deleting duplicate segments from performance / teaching videos. The strategy logic is as follows: video segments corresponding to the same tag are clustered, and those with a similarity greater than a threshold are considered duplicate segments. Alternatively, video segments with repeated movements can be directly filtered according to the duplicate tag, forming a group of duplicate video segments. The highest-scoring video segment is retained from each group, for example, the segment with the highest movement score or the segment with the highest emotion score in each group. Unselected duplicate segments are deleted, or their weight is reduced. This removes duplicate and similar video segments, and the selected video segments are spliced together to form a pre-edited video for use in dance demonstrations, teaching presentations, etc.
[0053] Highlight extraction aims to capture "excitement points," "climaxes," and "goals / kills" in sports or entertainment videos. The strategy involves: assigning motion intensity scores to video clips tagged with sports (or determining and recording these scores when setting video tags); prioritizing clips with motion intensity scores above a threshold; and considering emotion tags such as intense, exciting, and joyful. For example, a clip tagged with "goal" or "excitement" and having a motion intensity score above the threshold can be selected as a highlight. When splicing video clips, factors such as frame rate of change, background transitions, volume, and subtitles are considered to accelerate the editing process, resulting in a pre-edited video with a high frame rate of change, significant background transitions, and abrupt changes in volume or subtitles.
[0054] The editing goal of the editing template is to conform to a certain preset rhythmic structure, such as the rhythm and emotional change patterns of video blogs (Vlogs), brand promotional videos, and short dramas. Its strategy logic is to combine the content set in the editing template with the selected segments to generate a pre-edited video that conforms to the strategy. In one optional embodiment, the step of filtering tags according to the editing strategy and extracting multiple video segments whose tags conform to the editing strategy includes: determining the editing template corresponding to the editing strategy, matching tags according to the editing template, and extracting the video segments corresponding to the tags. The step of splicing the video segments to generate a pre-edited video includes: sorting and splicing the video segments according to tags and the editing template to generate a pre-edited video.
[0055] The editing templates include settings such as: the duration distribution of each segment (introduction, development, climax, and conclusion); a sequence of emotional tags, such as "gentle - gradually tense - intense - soothing at the end"; and background transition rhythms. Indirect templates also include tags to be filled, used to match video clips with corresponding tags to fill template gaps. Some examples also allow setting up a media library; if the required video clip is not found, similarly tagged video clips can be searched from the library to supplement it, improving the data quality and user experience of the pre-edited video.
[0056] In this embodiment, the server can set up a corresponding model to perform the above-mentioned pre-editing process. Based on the pre-trained model and template, it outputs a pre-edited video that meets the user's needs for subsequent fine-tuning.
[0057] In one alternative embodiment, highlight extraction is used as the editing strategy to illustrate the pre-editing process for the video to be edited, such as... Figure 5 As shown: Step 502: Receive the video to be edited and the editing strategy.
[0058] Users upload videos to be edited and select editing strategies on the editing page of their terminal devices. The terminal devices can then send the videos and editing strategies to the server, such as a Python-based web server like nginx, uWSGI, or FastAPI. The server stores the files in an appropriate storage location, such as a specified database or temporary storage. For example, the received editing strategy might be highlight extraction. Uploaded videos can be one or more videos.
[0059] The server can trigger a Python task, such as a decoding task of multiple videos to be edited in parallel using the Celery framework and a Redis Queue. Celery is a distributed task execution framework that can support the concurrent execution of a large number of tasks.
[0060] In one example, multi-threaded decoding can be performed as follows: import av # PyAV container=av.open(" / path / to / uploaded.mp4") stream=container.streams.video[0] stream.thread_type='AUTO' # Automatic multi-threaded decoding Step 504: Perform shot boundary detection on the video to be edited, and extract video segments according to the start and end frames of each shot.
[0061] The server can read the video to be edited frame by frame or by group of images (GOPs). Then, it performs shot boundary detection based on convolutional neural network models such as CNN and LSTM, and extracts video segments according to the start and end frames of each shot. The start frame of a shot can be a keyframe, i.e., an I-frame.
[0062] In one example, shot boundary detection and video clip extraction can be performed as follows: for frame in container.decode(stream): img=frame.to_ndarray(format="bgr24") # OpenCV image # Input the image into the CNN+LSTM model to detect if it is a "camera transition point". is_cut=shot_detector.predict(img) if is_cut: record_shot_boundary(frame.pts*ts_scale) Step 506: Perform content recognition on the video segment and determine the tag of the video segment based on the recognition result.
[0063] Step 508: Filter tags according to the editing strategy and extract multiple video clips whose tags match the editing strategy.
[0064] By performing various methods to identify the content of video clips, one or more tags can be determined for each clip. Specifically, for video clips tagged with motion or emotion, intensity detection can be performed during the identification process, and the corresponding intensity scores can be added to the tag. This intensity score can then be used to assist in filtering video clips during the selection process.
[0065] In one example, video clip filtering can be performed as follows: #Assuming slowfast_model is a PyTorch model that has already been loaded. for shot in shot_list: clip_frames=sample_frames(video,shot.start,shot.end, sample_rate=8) score=slowfast_model.predict(clip_frames) shot_scores[shot]=score # Sort by action intensity score and select the Top-N "highlight shots," which are the segments to be retained in the initial cut. N is a positive integer.
[0066] Step 510: Sort the video segments according to the editing strategy, and then splice the video segments according to the sorting results.
[0067] The selected video clips can be sorted according to time or other orders to obtain a list of start and end timestamps for each clip. One example already provides a list of start and end timestamps for several "highlight shots," such as: highlights=[ (start=15.2, end=20.4), (start=38.8, end=45.0), (start=72.5, end=80.1), ... ] Step 512: Add transition effects between two adjacent video clips to generate a pre-edited video.
[0068] In one example, the FFmpeg command line or PyAVAPI is used to concat these segments, and simple transition filters such as fade in and fade out are inserted between adjacent shots.
[0069] The above embodiments use highlight extraction editing strategies as an example to describe how a two-factor matching mechanism combining tags and motion scores is used to filter and splice pre-edited videos. If the user selects other editing strategies, the process is the same as described above, except that different video segments are selected and spliced based on different editing strategies.
[0070] After pre-editing is completed and a pre-edited video is generated, video metadata is generated by combining the tag data of each video segment in the pre-edited video. This video metadata includes tags and the corresponding time information. For example, based on object-type tags, the time information corresponding to each target object in the pre-edited video can be determined, thereby assisting in the location of video segments in subsequent dialogue-based fine-tuning, such as generating the position information of each object in each frame based on the characteristics and time information of the target object. Similarly, for action-type tags, the time information corresponding to each action, such as the timestamp range, can be recorded. Scene-type tags determine the time information corresponding to each scene. Based on the video metadata, subsequent location and extraction of video segments can be assisted, eliminating the need for repeated video inspection and improving processing efficiency.
[0071] During the video pre-editing stage, various types of tags are generated, including object-type tags based on YOLOv8 and DeepSORT models, action-type tags based on SlowFast models, and scene-type tags based on MiniCPM-V 2.6 models. The system also supports interactive editing intent and analysis of positioning information through natural language, click selection boxes, timestamps, and other methods.
[0072] Dialogue-based fine-tuning phase: This application embodiment can provide conversational fine-tuning, allowing users to express editing intentions in natural language through text or voice dialogue. Input controls are set in the editing page. In one example, this input control is a text input control such as an input box or input pop-up, which receives the input text data as editing intention information. In another example, this input control is a voice input control, through which voice data is input, and voice recognition is performed on the voice data, using the recognized text data as editing intention information. That is, the editing intention information includes the input editing intention text and / or the editing intention text converted from the input voice data, such as "cut the first five seconds shorter," "this shot has a warmer tone," "add some slow motion here," "slow down the subtitles," etc.
[0073] By using multimodal intent understanding and context-aware dialogue management, the system records multiple editing intents of users during the conversational fine-tuning process. Combined with contextual analysis of editing intents, the system enables users to directly control the fine-tuning of pre-edited videos using natural language or voice.
[0074] In one alternative embodiment, such as Figure 6 As shown, step 106 involves determining the adjustment operation through semantic recognition of the editing intent information, and adjusting the pre-edited video according to the adjustment operation, including the following sub-steps: Sub-step 1062 involves semantic recognition of the editing intent information to generate corresponding editing instructions, which include adjustment operations and target positioning information.
[0075] The editing intent information is semantically recognized, for example, through word segmentation and keyword recognition, or based on a semantic recognition model, to identify the adjustment operations and target positioning information within the editing intent information, and generate corresponding editing instructions. The adjustment operations are the adjustments to be performed on the video, such as adding filters, adjusting brightness, or adding effects. The target positioning information refers to the location (or segment) in the pre-edited video that needs adjustment. This can be the time range of the video segment to be adjusted, or the image information of a video frame, such as a specific person or scene.
[0076] Sub-step 1064: Match the segment to be adjusted in the pre-edited video according to the target positioning information, and perform the adjustment operation on the segment to be adjusted.
[0077] Then, according to the target location information in the editing instruction, the segment to be adjusted can be matched in the pre-edited video. For example, if the target location information is a timestamp of 1:15–1:25, then the 10-second video segment from 1 minute 15 seconds to 1 minute 25 seconds is determined as the segment to be adjusted. Then, the corresponding adjustment processing is performed on the segment to be adjusted according to the adjustment operation, such as adjusting the brightness. As another example, if the target location information is "a person in red clothes", "a person in red clothes" can be mapped to pedestrians in red clothing detected in several frames, thereby determining the corresponding segment to be adjusted.
[0078] This allows the system to recognize natural language as editing instructions, facilitating adjustments to the edited video and improving both editing efficiency and video quality.
[0079] In one optional embodiment of this application, such as Figure 7 As shown, sub-step 1062 involves semantic recognition of the editing intent information to generate corresponding editing instructions, including: Sub-step 702 involves performing natural language recognition on the editing intent information using a preset dictionary to determine target location information and adjustment operations. The adjustment operations include operation type and operation parameters.
[0080] Sub-step 704: Generate editing instructions using the adjustment operation and target positioning information.
[0081] Natural language processing is used to identify and process the editing intent information. In some examples, the editing intent information can be segmented into words, and then the segmentation results can be matched with a dictionary to determine the part of speech of the words and remove invalid words, such as prepositions and auxiliary words, to obtain multiple keywords. Based on the keywords, the target location information, operation type, and operation parameters of the segment to be adjusted can be determined. Some editing intent information explicitly indicates the time and image information of the segment to be adjusted. For example, from "brighten the person in red clothes", the target location information related to the image is "person in red clothes". Or, from "cut the first five seconds", the target location information related to the editing time is "first five seconds" or 0-5 seconds. Some editing intent information does not explicitly indicate the information of the segment to be adjusted, such as "this shot has a warm tone" or "add some slow motion here". The target location information can be determined by combining semantics with auxiliary information. For example, the position of the pointer on the timeline of the pre-edited video on the editing page, or the current playback position of the pre-edited video, which indicates the playback time, can be determined by combining the location information with semantics.
[0082] If the target location information is time information, i.e., the time range of the segment to be adjusted, then the segment to be adjusted can be obtained based on this time range. If the target location information is text description information, i.e., text information used to describe the content of the screen displayed in the segment to be adjusted, such as keywords, phrases, etc., such as "person in red clothes" or "roller coaster," then it is necessary to locate the segment to be adjusted in conjunction with the video footage. Thus, users can use text descriptions such as "second dance move," directly select the target in the preview screen, or use voice markers such as "from 1:15 to 1:30" to locate the precise spatiotemporal segment in one go, greatly facilitating the rapid target selection in complex video scenes.
[0083] In another alternative embodiment, such as Figure 8 As shown, sub-step 702 involves performing natural language recognition on the editing intent information using a preset dictionary to determine target location information and adjustment operations, including: Sub-step 7022: Input the editing intent information into the natural language model and output the corresponding adjustment operation. The natural language model includes a preset dictionary.
[0084] Sub-step 7024: Input the editing intent information and the pre-edited video into the multimodal model and output the target positioning information.
[0085] A natural language understanding (NLU) model can also be trained based on NLU algorithms. This model can then understand the semantics of editing intent information, identify the target location information, operation type, and operation parameters of the segment to be adjusted, and generate corresponding editing instructions. This NLU model can be a large model or a combination of one or more NLU-based models. In one example, a hierarchical NLU architecture can be used to decompose the editing intent information into structured editing instructions. One processing flow is as follows: def parse_command(user_interaction, video_metadata): # First level: Operation type classification, such as operation type is editing, color grading, adding effects, subtitles, etc. target = detect_target(text, video_metadata) # Second layer: parameter quantization, such as quantizing into brightness value, speed coefficient, effect type, subtitle content, etc. action = classify_action(text) # Third layer: Target positioning, which can be based on elements such as people, objects, camera angles, and timestamps. params = extract_params(text) return { 'action_type': action, # Example: Color mapping is for brightness adjustment 'parameters': params# Example: Brightness +20% 'target_anchor': target, # Example: Character ID = 3 (red clothes) + timestamp 1:15–1:25 } `classify_action(text)`: Action type classification can be based on pre-trained language models (such as BERT, RoBERTa, etc.) and custom action vocabularies to identify action types from editing intent information, such as categorizing them as editing, color grading, adding effects, subtitles, etc. BERT (Bidirectional Encoder Representation from Transformers) is a deep learning model based on the Transformer model. RoBERTa (A Robustly Optimized BERT Pretraining Approach) is an optimized BERT-based pretraining method.
[0086] extract_params(text): Parses numerical values, degree, modifiers, and other numerical terms in the editing intent information, such as warmer, half, faster, etc., and quantifies the numerical terms to obtain the corresponding numerical parameters.
[0087] `detect_target(text, video_metadata)` performs multimodal recognition on editing intent information and pre-edited video. For example, it performs semantic recognition on editing intent information based on Named Entity Recognition (NER) and relation extraction processing to obtain corresponding text features. It also performs image feature analysis on the video metadata of the pre-edited video. Furthermore, it can combine the tags of previously analyzed video segments to filter video segments and determine video metadata. For instance, it combines object class tags to determine the location of the video segment containing "the person in red," or, for "the second dance move," it searches for all "dance" tags in the action class tags of the pre-edited segment and locates the corresponding timestamp interval of the second segment according to their order of appearance. This reduces the amount of data analyzed and improves processing efficiency. It accurately locates the operation object or time interval through multimodal features such as text and images to obtain target location information.
[0088] This allows for the analysis of editing intent information using one or more models, identification of corresponding editing instructions, and improvement of processing efficiency.
[0089] In one optional embodiment, sub-step 1064, matching the segment to be adjusted in the pre-edited video according to the target positioning information, includes: determining a video frame in the pre-edited video that matches the target positioning information, and obtaining the segment to be adjusted based on the video frame. The target positioning information can be combined with the images of each video frame of the pre-edited video, and matching can be performed using multimodal information such as text and images to determine the video frame corresponding to the image described by the target positioning information, and obtaining the segment to be adjusted based on the video frame. For example, the segment between the first and last video frames containing the image described by the target positioning information can be used as the segment to be adjusted; or, based on the video frame containing the image described by the target positioning information, segments within a predetermined time period before and after it can be used as the segment to be adjusted.
[0090] In one alternative embodiment, such as Figure 9 As shown, the step of determining the video frame that matches the target positioning information in the pre-edited video, and obtaining the segment to be adjusted based on the video frame, includes the following sub-steps: Sub-step 902: Determine the text features of the target location information.
[0091] If the target location information is a screen description, feature extraction is performed on the screen description to determine the corresponding text features.
[0092] Sub-step 904: Determine the image features of the video frames in the pre-edited video.
[0093] For each video frame in the pre-edited video, feature extraction is performed on each video frame to determine the corresponding image features.
[0094] Sub-step 906: Calculate the similarity between the text features and the image features, and filter video frames that meet the matching conditions based on the similarity.
[0095] The similarity between the text features and image features is calculated. For example, after mapping the text features and image features to a semantic space, feature distances such as cosine similarity are calculated to obtain the corresponding similarity. A similarity threshold can be applied, and video frames with a similarity greater than the threshold are considered as meeting the matching criteria. One or more video frames that meet the matching criteria can be matched.
[0096] Sub-step 908: Based on the timestamps of video frames that meet the matching conditions, determine the segments to be adjusted.
[0097] Based on the timestamps of video frames that meet the matching criteria, such as using the segment between the timestamps of the first and last video frames as the segment to be adjusted, or using the timestamp of the video frame as a reference to extract segments before and after a predetermined time as the segment to be adjusted, etc.
[0098] In another optional embodiment, determining video frames in the pre-edited video that match the target positioning information, and obtaining the segment to be adjusted based on the video frames, includes: inputting the video frames of the pre-edited video and the target positioning information into a multimodal model, and outputting the timestamps of the matching video frames; and obtaining the segment to be adjusted based on the timestamps of the matching video frames.
[0099] In this embodiment, multimodal models can also be trained, such as Contrastive Language–Image Pre-training (CLIP) models and Multimodal Large Language Models (MLLMs), which are models capable of processing features from multiple modalities, including text, images, and videos. The multimodal model matches text features with video frames in the pre-edited video, and based on the timestamps of the matched video frames, locates the corresponding segment to be adjusted.
[0100] Taking the pre-trained CLIP model as an example, other similar alignment models can also be used to match target localization information with visual features in video frames. For example, "a person in red clothes" can be mapped to pedestrians detected in red clothing in several frames, the timestamps corresponding to the mapped video frames can be determined, and the start and end timestamp intervals can be determined as outputs, such as "1:15–1:25". Then, the segment to be adjusted can be extracted based on this start and end timestamp interval.
[0101] In this embodiment of the invention, a multimodal model, such as the CLIP model, is pre-trained using multimodal features from text and vision. This model matches the user's natural language description of their editing intent, such as "the person in red," with the visual features of video frames, achieving cross-modal semantic anchor point localization and locating the segment to be adjusted. A hierarchical NLU architecture is employed to decompose instructions into target location information, operation type, and operation parameters to generate editing instructions, accurately analyzing the user's semantics to obtain the editing instructions. Therefore, it eliminates the need for specific annotations for each object, such as people or objects. Based on the knowledge learned by the pre-trained model from large-scale image and text data, the system can quickly and accurately locate the area of the image the user wants to edit, significantly improving the accuracy and efficiency of instruction parsing.
[0102] In this embodiment of the invention, performing the adjustment operation on the segment to be adjusted includes: adjusting the segment to be adjusted according to the operation type and operation parameters, generating an incremental adjustment segment and caching it, so as to replace the segment to be adjusted after the adjustment is completed. After determining the adjustment operation, the adjustment processing is performed on the segment to be adjusted according to the operation type and operation parameters, such as increasing the brightness by 20% or shortening it by 2 seconds, etc. The adjustment is only performed on the segment to be adjusted, and the adjusted segment is the incremental adjustment segment and is cached. Other segments are not adjusted. If the user confirms that the adjustment of this segment is not a problem, no further adjustment is needed, and the adjusted segment can be used to replace the corresponding segment in the video. In an optional embodiment, rendering parameters are determined according to the operation type and operation parameters, and the incremental adjustment segment is rendered according to the rendering parameters to generate preview data and provide feedback, so as to display the adjusted content of the video in the client. For the incremental adjustment segment, the effect after the adjustment is completed still needs to be previewed by the user to determine whether it meets the user's needs and whether further adjustment is needed. Therefore, the rendering parameters are determined according to the operation type and parameters. For example, executable rendering parameters could be brightness +20%, speed 0.5×, and filter type "cool tone". The incremental adjustment segment is then rendered according to these rendering parameters, generating corresponding preview data and providing feedback. When displayed on the client, only the adjustment segment (i.e., the preview data) can be shown, or the preview data can replace the segment to be adjusted, thus displaying the entire video to the user.
[0103] In this embodiment of the invention, after parsing the user's intent, the server completes partial rendering and returns a preview within a few seconds. The user can view the current adjustment effect in real time and continue to issue new commands for fine-tuning, achieving a closed-loop multi-turn dialogue. For unadjusted video segments, the rendering result of the previous version is directly read from the cache, eliminating the need for recalculation.
[0104] After each user submits an editing intent, for adjustments to a specific segment, the server no longer re-renders the entire pre-edited video; instead, it incrementally renders only the matched segment. In one example, the adjustment and rendering process for the segment is as follows: VideoClip clip = getOriginalClip(); Adjustment adjustment = parseAdjustment(command); / / Parse the command to obtain the operation type and operation parameters. FrameRange range = adjustment.getTargetFrames(); / / Determine the frame range of the segment to be adjusted, such as 1:15–1:25 clip.applyAdjustment(range, adjustment); / / Applies the adjustment type and parameters to the range of this frame, such as increasing or decreasing brightness by 20%, adding a specified effect, or increasing playback speed by 30%. PreviewGenerator.generateIncrementalPreview(clip, range); / / Generates only incrementally adjusted clips as preview data. etOriginalClip(): Gets the video data structure of the current initial clip or the result of the last rendering.
[0105] parseAdjustment(command): Converts the editing command into executable rendering parameters, such as brightness +20%, speed 0.5×, filter type "cool tone", etc.
[0106] In this embodiment of the invention, when the server renders a video (video clip), it can prioritize calling the front-end real-time preview, such as WebGL (Web Graphics Library), or the back-end batch rendering, such as CUDA (Compute Unified Device Architecture), for efficient image processing.
[0107] In another optional embodiment, rendering can be performed on the client side. After the server determines the rendering parameters according to the operation type and operation parameters, it sends the rendering parameters to the client. The client renders the incremental adjustment segment according to the rendering parameters, generates preview data, and provides feedback to display the adjusted video content on the client. For the incremental adjustment segment, the effect after adjustment still needs to be previewed to the user to determine whether it meets the user's needs and whether further adjustments are needed. Therefore, the rendering parameters are determined according to the operation type and operation parameters. For example, executable rendering parameters could be brightness +20%, speed 0.5×, and filter type "cool tone," etc. The incremental adjustment segment is rendered according to these rendering parameters, generating corresponding preview data and displaying it on the client. Only the adjustment segment (the preview data) can be displayed, or the preview data can replace the segment to be adjusted, thus displaying the entire video to the user.
[0108] Incremental rendering is performed only on the user-specified segments to be adjusted, using technologies such as WebGL and CUDA acceleration. Other unadjusted segments reuse the cache, avoiding repetitive rendering of the entire video and reducing computational load. Through this method, user-interactive fine-tuning operations, such as color correction, slow motion, and subtitle placement, can generate previewable videos within 3–5 seconds at most common resolutions. This ensures a smooth, lag-free editing process, significantly improving the interactive experience and user satisfaction.
[0109] Based on the above embodiments, conversational fine-tuning can be achieved through the following steps, such as... Figure 10 As shown: Step 1002: Receive editing intent information.
[0110] Step 1004: Combine the preset dictionary to perform natural language recognition on the editing intent information to determine the target positioning information and adjustment operation.
[0111] The adjustment operation includes an operation type and operation parameters. Specifically, the editing intent information is input into a natural language model, which outputs the corresponding adjustment operation; the natural language model includes a preset dictionary. The editing intent information and the pre-edited video are input into a multimodal model, which outputs target location information.
[0112] Step 1006: Generate editing instructions using the adjustment operation and target positioning information.
[0113] If the target location information is time information, proceed to step 1008; if the target location information is text description information, proceed to step 1010.
[0114] Step 1008: Extract the segment to be adjusted from the pre-edited video according to the time information.
[0115] Step 1010: Determine the video frame in the pre-edited video that matches the target positioning information, and obtain the segment to be adjusted based on the video frame.
[0116] In one optional embodiment, the text features of the target location information are determined, and the image features of the video frames in the pre-edited video are determined; the similarity between the text features and the image features is calculated, and video frames that meet the matching conditions are filtered based on the similarity; based on the timestamps of the video frames that meet the matching conditions, the segments to be adjusted are determined.
[0117] In another optional embodiment, the video frames and target positioning information of the pre-edited video are input into a multimodal model, and the timestamps of the matching video frames are output; based on the timestamps of the matching video frames, the segment to be adjusted is obtained.
[0118] Step 1012: Adjust the segment to be adjusted according to the operation type and operation parameters to generate an incremental adjusted segment.
[0119] Step 1014: Replace the segment to be adjusted with the incremental adjustment segment to generate the target video.
[0120] During conversational fine-tuning, context-aware dialogue management is implemented, recording the user's editing intentions across multiple rounds. This allows for continuous, "inherited" modifications, reducing repetitive confirmations and redundant input. For example, if a user says "Add a filter to the first shot" and then "Reduce the intensity by half," the system can automatically correlate the context to accurately determine if the adjustment is for the same segment. When instructions are ambiguous, the system automatically returns candidate segments for user confirmation and provides error correction options to address vague intentions, preventing accidental operations.
[0121] In this embodiment of the invention, the context state representation can be improved. It not only records the user's last operation but also models user preferences, such as frequently used editing styles and filter types, as well as historical aesthetic feedback like likes and ratings for videos, to achieve "preference memory" and "style adaptation" functions. For example, if a user is satisfied with the "cool tone" filter multiple times, similar parameters can be automatically recommended next time.
[0122] Multi-platform standardized export stage: Once the user is satisfied with the video quality, the server can automatically perform format conversion and packaging based on the target platform, reducing the user's manual adaptation work. Therefore, it can obtain the output platform, determine the format information based on the output platform, and perform format conversion on the target video based on that format information. Then, it sends the target video, including sending the format-converted target video.
[0123] The format information is determined based on the output platform. Different output platforms have different output format requirements. The format information for the platform to which the video will be output can be determined. This format information includes resolution, aspect ratio, video encoding format, and audio encoding specifications. Therefore, the target video can be converted to meet the requirements of the output platform based on this format information. For example, using low-level libraries such as FFmpeg, the already rendered target video can be transcoded and packaged according to the aforementioned format information to generate a platform-specific video file.
[0124] In one example, the output platform's format information is configured as follows: Resolution and aspect ratio: For platform A, 1080×1920 (9:16) is recommended; for platform B, 1280×720 (16:9) or 1920×1080 (16:9) are optional. Video encoding format: H.264 / H.265 standard parameters (bitrate, frame rate); Audio encoding standards: AAC / MP3, etc.
[0125] The target video can also be configured with a cover image for display on the output platform. The cover specifications for the output platform are determined; target image frames are selected from the target video to serve as the cover, and these frames are converted according to the specified cover specifications to generate the cover image. The target image frames can also be provided to the user through an editing page, allowing the user to adjust the cover image as needed. The cover specifications refer to the dimensions of the cover as defined by the output platform. These specifications can also include information such as the output platform's requirements for the content, thus allowing the selection of target image frames from the target video based on these specifications.
[0126] The embodiments of the present invention can also automatically generate the title and publishing tags of the target video based on the content of the target video. For example, the video title can be automatically generated based on the editing strategy of the target video, the title of the video segment, etc. The same video can generate different video titles for different output platforms.
[0127] It can also generate and embed subtitle files according to the subtitle specifications of the output platform, such as generating .srt format subtitles to add to the target video, and generating or optimizing cover images, such as automatically adding watermarks, logos, text descriptions, etc.
[0128] If the user has enabled the Application Programming Interface (API) or third-party publishing permissions for the corresponding output platform, the server can publish the video, cover, title, publishing tags, etc. to the target output platform with one click, eliminating the need for subsequent manual operations.
[0129] This invention incorporates platform standards for resolution, encoding, cover art, and metadata across multiple platforms, enabling one-click matching and generation of compliant video files, supporting simultaneous publishing. It eliminates the tedious steps of manually adjusting resolution, bitrate, cover art size, and tag format, achieving end-to-end automation from editing to publishing, significantly improving operational efficiency and reducing error risks.
[0130] This invention provides a method for rapid pre-editing and dialog-based fine-tuning, enabling the rapid generation of target videos for easy distribution on appropriate output platforms. Building upon the above embodiments, this invention also provides a video editing method that performs editing processing via an editing page on a terminal device.
[0131] Reference Figure 11 The diagram illustrates a flowchart of another embodiment of the video editing method of the present invention.
[0132] Step 1102: Display the pre-edited video on the editing page, which includes input controls.
[0133] A video editing page is displayed on the terminal device. This page provides various editing-related functions. After instructing the user to perform pre-editing via this page, the pre-edited video is received and displayed on the editing page. The editing page includes input controls for entering editing intent information describing the intended adjustments.
[0134] Step 1104: In response to the triggering of the input control, obtain the editing intent information.
[0135] When a user views a pre-edited video and wants to make adjustments, they can trigger an input control. In response to this input control, the system retrieves the user's editing intent. Based on this input control, interactive fine-tuning of the pre-edited video can be performed; that is, the user describes the desired adjustments using natural language, providing the editing intent information without requiring manual adjustments to the pre-edited video.
[0136] Step 1106: Send the editing intent information.
[0137] The editing intent information is sent to the server, which then adjusts the pre-edited video based on this information. The adjustment process is described in the above-described embodiment. The editing intent information is the user-provided adjustment intention described in natural language during the dialogic fine-tuning process. Throughout the fine-tuning process, the user can input the editing intent information multiple times to perform multiple rounds of adjustments on the pre-edited video until the adjustments are complete and the adjusted target video is generated.
[0138] Step 1108: Receive the adjusted target video and display it on the editing page.
[0139] Receive the adjusted target video, display the target video on the editing page, and complete the video editing.
[0140] In summary, after pre-editing a video on the editing page, users can provide editing intent information via natural language. The server can then parse this intent information based on natural language understanding and adjust the pre-edited video accordingly, enabling conversational fine-tuning of the video. This eliminates the need for users to learn professional editing skills and tools, lowering the barrier to entry for video editing and improving editing efficiency.
[0141] Dialogue-based fine-tuning eliminates the cumbersome traditional click-based or parameter form-based operations. With just one sentence, you can complete various fine-tuning tasks such as color correction, speed adjustment, special effects overlay, and subtitle layout, greatly improving editing efficiency and ease of use.
[0142] In this embodiment of the invention, the input control includes a text input control and / or a voice input control. The text input control is used to input text information describing the adjustment intention, and the voice input control is used to input voice information describing the adjustment intention. Step 104, in response to triggering the input control, acquiring editing intention information, including at least one of the following: in response to triggering the text input control, using the received text information as editing intention information; in response to triggering the voice input control, receiving voice data and performing voice recognition on the voice data, using the recognized text information as editing intention information.
[0143] In response to a trigger on the text input control, the received text information is used as editing intent information to obtain the editing intent described in natural language text. In response to a trigger on the voice input control, voice data is received, and then voice recognition is performed on the voice data to obtain text information, which is used as editing intent information. Thus, editing intent information can be input through dialogue methods such as text and voice.
[0144] During the dialog-based fine-tuning process, adjustments to the pre-edited video can be performed incrementally. The server can provide the user with the segment to be adjusted for confirmation to ensure the accuracy of the adjustments. Figure 12 As shown, it also includes the following steps: Step 1202: Receive the segment to be adjusted, which is obtained by matching the pre-edited video based on target positioning information, and the target positioning information is determined based on semantic recognition of the editing intent information.
[0145] Step 1204: Display the segment to be adjusted on the editing page.
[0146] The server locates the segment to be adjusted based on the semantic location of the editing intent information, sends the segment to be adjusted, and the terminal device displays the segment to be adjusted on the editing page, for example, through a preview box.
[0147] Step 1206: Receive confirmation results for the segment to be adjusted, and send the confirmation results to determine the accuracy of the segment to be adjusted based on the confirmation results.
[0148] If a user confirms that the segment to be adjusted is correct, they can input a confirmation instruction via text or voice using the input controls. Confirmation controls can also be provided to issue this confirmation. If a user confirms that the segment to be adjusted is incorrect, they can input editing intent information via text or voice using the input controls to adjust the segment. The aforementioned confirmation instructions and editing intent information can be sent to the server as confirmation results. If the server confirms there are no problems, it can execute the adjustment process. If the server confirms there are problems, it can adjust the segment to be edited based on the editing intent information, and then make the adjustment only after confirming that everything is correct.
[0149] Before dialog-based fine-tuning, the video can be pre-edited. The editing page includes an upload control for uploading the video to be edited. In response to triggering the upload control, the video to be edited is received and sent to the server. The editing page also includes a strategy selection control for selecting an editing strategy. In response to triggering the strategy selection control, the selected editing strategy is determined and sent to the server.
[0150] After the editing page is launched, upload controls and strategy selection controls can be displayed on the editing page. After the user makes the selection, the video to be edited and the editing strategy are sent to the server as the first step of editing, and then conversational fine-tuning is performed.
[0151] The editing page includes a platform selection control for selecting at least one output platform. In response to triggering the platform selection control, the output platform is determined and sent to the server. The server can determine format information based on the output platform and perform format conversion on the target video according to the format information. It can also determine the cover specifications, subtitle specifications, etc., of the output platform and generate the cover and subtitles required by the output platform.
[0152] Based on the above embodiments, this invention provides a video editing system, including a terminal device and a server. The terminal device includes a front-end user interaction module, providing an editing page for user interaction via an application (APP), browser, etc. The server can include a Natural Language Understanding (NLU) module, a multimodal semantic understanding module, a video processing and rendering module, and a dialogue management module.
[0153] The overall process of the system is as follows: 1) Users input their editing intent information in natural language via text or voice on the editing page of the terminal device.
[0154] 2) The server uses the Natural Language Intent Recognition (NLU) module to perform layered parsing of the instructions, breaking them down into target location information and adjustment operations. The adjustment operations include operation type and operation parameters.
[0155] 3) The multimodal semantic understanding module uses a pre-trained model (such as CLIP) to perform semantic matching between user descriptions and video frames, thereby achieving spatiotemporal localization of semantic anchor points and locating the segments to be adjusted.
[0156] 4) The video processing module performs local adjustments and incremental rendering on the segment to be adjusted based on the parsed adjustment operations.
[0157] 5) The dialogue management module records the interaction status and supports multi-round instruction and intent clarification.
[0158] 6) The server returns a preview of the rendered video, which the user can then adjust or export.
[0159] Reference Figure 13 The diagram illustrates an interactive schematic of a video editing system according to this application.
[0160] Step 1302: The server sends the page data of the clipped page to the terminal device.
[0161] Step 1304: The user terminal renders the page data and displays the clip page.
[0162] The editing page includes: an upload control and a strategy selection control.
[0163] Step 1306: In response to the triggering of the upload control, receive the video to be edited.
[0164] Step 1308: In response to the triggering of the strategy selection control, determine the selected clipping strategy.
[0165] On the editing page, users select "Initial Edit" and upload the video to be edited using the upload controls, such as uploading a basketball game video. They then select an editing strategy using the strategy selection controls, such as choosing the "Highlight Extraction" strategy.
[0166] This invention provides various automatic editing strategies such as deduplication, highlight extraction, and editing templates. Based on deep learning models for shot transition detection, keyframe extraction, and motion intensity analysis, it generates coherent pre-edited videos with a single click. Users only need to select an editing strategy, and the server can complete scene segmentation, keyframe selection, and rhythm splicing of complex videos within seconds, significantly lowering the initial editing threshold and enabling non-professional users to quickly obtain high-quality rough cuts.
[0167] Step 1310: Send the video to be edited and the editing strategy to the server.
[0168] Step 1312: The server generates a pre-edited video based on the video to be edited and the editing strategy.
[0169] The pre-editing step includes: receiving the video to be edited and an editing strategy; detecting the video to be edited and extracting multiple video segments that conform to the editing strategy; and splicing the video segments together to generate a pre-edited video.
[0170] Taking the pre-editing of basketball game videos according to the highlight extraction strategy as an example, the server analyzes the basketball game video and detects multiple offensive shots, dunks and other highlight video clips. It automatically cuts out the highlight video clips such as 00:02:15–00:02:30 and 00:05:40–00:05:55 according to the strategy, and generates a 30-second pre-edited video.
[0171] On the server side, video decoding and frame processing can be handled using audio / video processing engines such as the FFmpeg framework, combined with corresponding tool libraries such as PyAV and OpenCV. FFmpeg is responsible for "decoding" and "encoding" video, and is also used to synthesize video files.
[0172] The process involves frame-by-frame reading and writing using PyAV (a Python binding for FFmpeg) or OpenCV: Open the original video with PyAV, decode it by keyframe (I-frame) or a specified frame rate (e.g., extract N frames per second) to obtain a Python-operable frame sequence (av.VideoFrame), and convert each frame into a NumPy array or an OpenCV image (cv2 format) for processing by the backend deep learning model (CLIP, YOLOv8, etc.).
[0173] The process of detecting the video to be edited and extracting multiple video segments that conform to the editing strategy includes: detecting the video to be edited, extracting video segments and determining the tags of the video segments; filtering tags according to the editing strategy, and extracting multiple video segments whose tags conform to the editing strategy.
[0174] The process involves identifying video segments based on shot transition detection, such as using lightweight networks based on ResNet and bidirectional LSTM models to extract video segments. For video segment labeling, object detection, object tracking, and action recognition can be used to determine the labels. Object detection is implemented using the YOLOv8 or Detectron2 model, object tracking using the DeepSORT model, and action recognition using models such as SlowFast, I3D, or TSM.
[0175] Step 1314: The server sends the pre-edited video to the terminal device.
[0176] Step 1316: Display the pre-edited video on the editing page.
[0177] Step 1318: In response to the triggering of the input control, obtain the editing intent information.
[0178] When a user previews a pre-edited video on the editing page, they find that the athlete's shooting motion is too dark. They then use voice input to "brighten the first highlight segment of the basketball shooting motion by 30%". The corresponding text information is obtained through voice recognition and used as the editing intent information, namely, "brighten the first highlight segment of the basketball shooting motion by 30%".
[0179] Step 1320: Send the editing intent information to the server.
[0180] Step 1322: The server determines the adjustment operation and locates the segment to be adjusted by semantic recognition of the editing intent information.
[0181] The system uses natural language understanding and video for multimodal intent understanding, mapping the "first highlight segment" to target location information: 00:02:15–00:02:30, with the adjustment operation being a +30% increase in brightness. Furthermore, it aligns text and video images based on multimodal information, such as using the CLIP model implemented in PyTorch by OpenAI.
[0182] In this embodiment of the invention, each model is implemented in a Python environment, such as PyTorch, TensorFlow, Keras, etc., to extract structured metadata from video frames, including object location, action category, scene label, etc.
[0183] Step 1324: Send the segment to be adjusted.
[0184] Step 1326: Display the segment to be adjusted on the editing page.
[0185] Step 1328: Receive and send the confirmation result for the segment to be adjusted.
[0186] If the user confirms that the clip is correct, they can directly provide feedback that it is accurate. If the user confirms that it is incorrect, they can re-enter the editing intent information to adjust the clip, such as providing a more accurate description of the clip to be adjusted.
[0187] Step 1330: Perform the adjustment operation on the segment to be adjusted to generate the adjusted segment.
[0188] The rendering engine performs incremental rendering, re-rendering color correction only for the specified time period, and returning a new preview after 3 seconds. Incremental rendering can utilize the trim, concat, and filter functions of the FFmpeg framework to process only the segments to be adjusted.
[0189] For example, the video can be divided into three segments: "[0,t_start)|[t_start,t_end)|(t_end,duration]". Filters such as color adjustment filters and speed adjustment filters can be applied only to the middle segment [t_start,t_end). The segments before and after the middle segment can be directly copied (copy codec) or the cache can be reused.
[0190] For user-specified segments requiring color correction, effects, or slow motion, simple brightness / contrast adjustments are implemented using OpenCV and NumPy, or custom filters accelerated by CUDA are invoked, such as GPUTensor operations based on PyTorch, before being re-encoded back into video format using PyAV or FFmpeg.
[0191] Server-side incremental rendering is combined with GPU acceleration, such as CUDA, WebGL, and Vulkan, to achieve acceleration. For example, on the server side, if the hardware supports NVIDIA GPUs, CUDA acceleration can be enabled.
[0192] In particular, PyTorch's custom color transformation operators, such as brightness / contrast operations on GPU tensors, enable the adjustment of segments to be adjusted.
[0193] Adjustments to slow motion / fast motion are actually frame interpolation / extraction logic for frame rate and timestamps, which can be achieved through cuVID decoding and NVENC encoding in the NVIDIA Video Codec SDK for extremely fast GPU-side processing.
[0194] Step 1332: The server sends the adjustment fragment.
[0195] Users can continue to input editing intent information to further adjust the pre-edited video. For example, a user might input "add slow motion to the first two seconds, and add the subtitle 'Excellent Shot' to the last two seconds" as their editing intent. The server parses the editing intent information, obtaining editing instructions of speed 0.5 applied to 00:02:15–00:02:17, and subtitle text and time periods mapped to 00:02:23–00:02:25. A partial rendering is then performed, and the updated preview is returned.
[0196] This can be achieved by enabling CUDA acceleration via FFmpeg to achieve sub-second rendering. During the preview in the front-end WebGL editing page, to ensure users see the preview within seconds, lightweight incremental rendering can be performed in the front-end browser using WebAssembly combined with FFmpeg.wasm or WebGL. Server-side adjusted segments, such as those in H.264 / H.265 format, are pushed to the front-end app's editing page for display. The front-end app uses WebGL Shaders to make real-time adjustments to color, contrast, speed, etc., and draw them onto the canvas. When the user sends a new command, the server is requested to continue the adjustments. This avoids the user having to wait for the backend to repackage the entire video each time, instead loading only "several seconds" of footage, improving processing efficiency.
[0197] Step 1334: Display the adjusted clip on the editing page.
[0198] Step 1336: Send confirmation instruction.
[0199] In response to the confirmation control being triggered, a confirmation instruction is generated and sent.
[0200] Step 1338: Generate the target video and return it to the terminal device.
[0201] Step 1340: In response to the triggering of the platform selection control, determine the output platform and send it to the server.
[0202] Step 1342: Determine the format information based on the output platform, and perform format conversion on the target video based on the format information.
[0203] Once the user is satisfied, they click the confirmation instruction, such as the "Export" control. Then, they select the target platform, such as Platform A, using the platform selection control. The content can be published directly to the user's account on Platform A, or it can be submitted to the editing page and then published to the user's account on Platform A from there.
[0204] The server automatically transcodes the target video to meet the format requirements of platform A, such as converting it to a resolution of 1080×1920 and H.264 encoding, and generates a cover image. It can also be "published" to the user's Douyin account with one click.
[0205] The embodiments of this invention, through processes such as initial pre-editing, conversational fine-tuning, and multi-platform publishing, utilize innovative technologies such as multimodal understanding of videos, semantic tagging, incremental rendering, and contextual dialogue analysis to achieve industry-leading video editing efficiency and experience.
[0206] In high-concurrency scenarios, this invention utilizes container orchestration platforms such as Kubernetes and Docker Swarm to automatically scale and load balance rendering services, avoiding queuing or long waiting times during peak editing periods.
[0207] This invention also enables version control and rollback functions, providing version management capabilities. Each "dialogue-based fine-tuning" generates a new version adjustment segment, allowing users to rewind to any historical video node at any time or compare the differences between different video versions, such as comparing the color histograms and length time-series curves of frames before and after rendering, thereby improving the quality of edited videos.
[0208] This invention also enables a professional integrated interface for film and television post-production, connecting with professional editing / color grading software such as DaVinci Resolve, Adobe Premiere Pro, and Final Cut Pro to form plugins or APIs. This allows for the embedding of features such as dialog-based fine-tuning and incremental rendering into existing workflows, meeting the highly customized needs of professional film and television post-production.
[0209] This invention supports the "live streaming with editing" scenario, accelerating the system to zero latency, allowing live streaming content to be edited and displayed in real time during the live stream, and enabling functions such as "viewers giving instructions to edit and replay during the live stream" and "automatically extracting highlights from the live stream and inserting bullet comments".
[0210] This invention also provides an electronic device, such as... Figure 14 As shown, it includes a processor 141, a communication interface 142, a memory 143, and a communication bus 144, wherein the processor 141, the communication interface 142, and the memory 143 communicate with each other through the communication bus 144. Memory 143 is used to store computer programs; When processor 141 executes a program stored in memory 143, it performs the following steps: Send a pre-edited video for display on an editing page, which includes input controls; Receive editing intent information, the editing intent information being in response to a trigger determination of the input control; The adjustment operation is determined by semantic recognition of the editing intent information, and the pre-edited video is adjusted according to the adjustment operation to generate the target video; Send the target video to display it on the editing page.
[0211] In another alternative embodiment, when the processor 141 executes the program stored in the memory 143, it performs the following steps: The pre-edited video is displayed on the editing page, which includes input controls; In response to the triggering of the input control, the editing intent information is obtained; Send the editing intent information to adjust the pre-edited video based on the editing intent information, wherein the adjustment is performed based on the adjustment operation determined by semantic recognition of the editing intent information; Receive the adjusted target video and display it on the editing page.
[0212] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0213] The communication interface is used for communication between the aforementioned terminal and other devices.
[0214] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0215] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0216] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform the video key point extraction method and the key point-based video playback method described in any of the above embodiments.
[0217] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the video key point extraction method and the key point-based video playback method described in any of the above embodiments.
[0218] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0219] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0220] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0221] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A video editing method, characterized in that, The method includes: Send a pre-edited video for display on an editing page, which includes input controls; Receive editing intent information, the editing intent information being in response to a trigger determination of the input control; The adjustment operation is determined by semantic recognition of the editing intent information, and the pre-edited video is adjusted according to the adjustment operation to generate the target video; Send the target video to display it on the editing page.
2. The method according to claim 1, characterized in that, The editing intent information includes the input editing intent text, and / or the editing intent text obtained by converting input speech data; The step of determining the adjustment operation through semantic recognition of the editing intent information, and adjusting the pre-edited video according to the adjustment operation, includes: The editing intent information is semantically recognized to generate corresponding editing instructions, which include adjustment operations and target positioning information. Match the segment to be adjusted in the pre-edited video according to the target positioning information, and perform the adjustment operation on the segment to be adjusted.
3. The method according to claim 2, characterized in that, The step of semantically recognizing the editing intent information and generating corresponding editing instructions includes: The editing intent information is analyzed using natural language recognition based on a preset dictionary to determine target location information and adjustment operations, wherein the adjustment operations include operation type and operation parameters; Editing instructions are generated using the aforementioned adjustment operations and target positioning information.
4. The method according to claim 3, characterized in that, The step of combining a preset dictionary with natural language recognition to determine target location information and adjustment operations for the editing intent information includes: The editing intent information is input into a natural language model, and the corresponding adjustment operation is output. The natural language model includes a preset dictionary. The editing intent information and the pre-edited video are input into the multimodal model, and the target localization information is output.
5. The method according to claim 2, characterized in that, The step of matching the segment to be adjusted in the pre-edited video according to the target positioning information includes: If the target location information is time information, then the segment to be adjusted is extracted from the pre-edited video according to the time information; If the target location information is text description information, then a video frame matching the target location information is determined in the pre-edited video, and the segment to be adjusted is obtained based on the video frame.
6. The method according to claim 5, characterized in that, The step of determining the video frame in the pre-edited video that matches the target positioning information, and obtaining the segment to be adjusted based on the video frame, includes: Determine the text features of the target location information and the image features of the video frames in the pre-edited video; Calculate the similarity between the text features and the image features, and filter video frames that meet the matching conditions based on the similarity. Based on the timestamps of video frames that meet the matching criteria, the segments to be adjusted are determined.
7. The method according to claim 5, characterized in that, The step of determining the video frame in the pre-edited video that matches the target positioning information, and obtaining the segment to be adjusted based on the video frame, includes: The video frames and target positioning information of the pre-edited video are input into the multimodal model, and the timestamps of the matching video frames are output. Based on the timestamps of the matched video frames, the segment to be adjusted is obtained.
8. The method according to claim 3, characterized in that, Performing the adjustment operation on the segment to be adjusted includes: According to the operation type and operation parameters, the segment to be adjusted is adjusted, an incremental adjustment segment is generated and cached, so as to replace the segment to be adjusted after the adjustment is completed.
9. The method according to claim 8, characterized in that, Also includes: The rendering parameters are determined according to the operation type and operation parameters. The incremental adjustment segment is rendered according to the rendering parameters, preview data is generated and fed back, so as to display the video adjustment content in the client.
10. The method according to claim 1, characterized in that, It also includes a pre-editing step: Receive the video to be edited and the editing strategy; The video to be edited is detected, and multiple video segments that conform to the editing strategy are extracted; The video clips are spliced together to generate a pre-edited video.
11. The method according to claim 10, characterized in that, The step of detecting the video to be edited and extracting multiple video segments that conform to the editing strategy includes: The video to be edited is inspected, video segments are extracted, and the tags of the video segments are determined; Based on the editing strategy, select tags and extract multiple video clips whose tags match the editing strategy.
12. The method according to claim 11, characterized in that, The process of detecting the video to be edited, extracting video segments, and determining the tags of the video segments includes: Perform shot boundary detection on the video to be edited, and extract video segments according to the start and end frames of each shot; The video segment is subjected to content recognition, and the video segment is tagged based on the recognition results. The tags represent the content category of the video segment.
13. The method according to claim 10, characterized in that, The process of splicing together the video segments to generate a pre-edited video includes: The video segments are sorted according to the editing strategy, and then spliced together according to the sorting results. Add transition effects between two adjacent video clips to generate a pre-edited video.
14. The method according to claim 1, characterized in that, Also includes: The format information is determined based on the output platform, and the target video is converted based on the format information; Sending the target video includes sending the target video after format conversion.
15. A video editing method, characterized in that, The method includes: The pre-edited video is displayed on the editing page, which includes input controls; In response to the triggering of the input control, the editing intent information is obtained; Send the editing intent information to adjust the pre-edited video based on the editing intent information, wherein the adjustment is performed based on the adjustment operation determined by semantic recognition of the editing intent information; Receive the adjusted target video and display it on the editing page.
16. The method according to claim 15, characterized in that, The input control includes a text input control and / or a voice input control. The step of obtaining editing intent information in response to triggering the input control includes at least one of the following: In response to the triggering of the text input control, the received text information is used as editing intent information; In response to the triggering of the voice input control, voice data is received and voice recognition is performed on the voice data, and the recognized text information is used as editing intent information.
17. The method according to claim 15, characterized in that, Also includes: Receive a segment to be adjusted, the segment to be adjusted being matched in the pre-edited video based on target positioning information, the target positioning information being determined based on semantic recognition of the editing intent information; The segment to be adjusted is displayed on the editing page; Receive confirmation results for the segment to be adjusted, and send the confirmation results to determine the accuracy of the segment to be adjusted based on the confirmation results.
18. The method according to claim 15, characterized in that, The editing page includes: an upload control and a strategy selection control; the method further includes: In response to the triggering of the upload control, the video to be edited is received; In response to the triggering of the strategy selection control, the selected editing strategy is determined; Send the video to be edited and the editing strategy to pre-edit the video based on the editing strategy to generate a pre-edited video.
19. The method according to claim 15, characterized in that, The editing page includes a platform selection control, and the method further includes: In response to the triggering of the platform selection control, the output platform is determined and sent to the server so that the server can perform format conversion on the target video based on the format information of the output platform.
20. A video editing system, characterized in that, The system includes: a server and a client; The server sends a pre-edited video to be displayed on the editing page; receives editing intent information; determines adjustment operations through semantic recognition of the editing intent information; adjusts the pre-edited video according to the adjustment operations to generate a target video; and sends the target video to be displayed on the editing page. The client displays a pre-edited video on an editing page, which includes an input control; in response to triggering the input control, it acquires editing intent information; sends the editing intent information; and receives the adjusted target video and displays it on the editing page.
21. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the method described in any one of claims 1-19.
22. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-19.
23. A computer program product comprising a computer program / computer-executable instructions, wherein, When the computer program / computer executable instructions are executed by a processor in an electronic device, they implement the method described in any one of claims 1-19.