Video editing method and related device

By splitting the video into sub-videos and generating plot scripts, the problem of low editing of long videos is solved, and an efficient and coherent video editing process is achieved, and the editing quality is improved.

CN120499474APending Publication Date: 2025-08-15BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510894209.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

In the prior art, when editing long videos is short videos, it is necessary to manually watch and write plot scripts, resulting in high time consumption and low efficiency. The model is burdened with long videos, slow in reasoning speed, and incomplete plot coverage when processing long videos.

Method used

Split the video to be edited into multiple sub-videos, generate text description information through the video description model, and generate plot scripts using the text generation model to guide the editing process and accurately filter video clips related to the topic.

Benefits of technology

It improves the degree of automation and efficiency of video editing, ensures the consistency and integrity of editing results, reduces the workload of post-adjustment, and significantly improves the overall quality of video editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499474A_ABST
    Figure CN120499474A_ABST
Patent Text Reader

Abstract

The invention discloses a video editing method and a related device, and the method comprises the steps: dividing a to-be-edited video into a plurality of sub-videos, carrying out the segmentation processing, generating text description information in combination with feature information, and further generating a plot script to guide an editing process. The problems of heavy model burden, low reasoning speed, incomplete plot coverage and the like caused by inputting the whole section of long video in the prior art are solved. Meanwhile, the editing strategy is guided through the plot script, so that the continuity and integrity of the editing result are improved, and the possibility of later patch repairing is remarkably reduced, thereby greatly improving the video editing efficiency on the premise of ensuring the video quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a video editing method and related devices. Background Art

[0002] Short videos, as a means of disseminating video content, have grown rapidly in recent years due to their fast-paced plots and numerous exciting moments. To cater to viewers' fast-paced viewing habits, editors often cut longer videos into shorter ones.

[0003] In related technologies, during the editing process, it is necessary to manually watch all long videos and sort out the context, and then manually select video clips that meet the editing requirements and edit them.

[0004] However, this video editing method requires manual work that consumes a lot of time in watching videos and writing plot scripts, resulting in low editing efficiency. Summary of the Invention

[0005] In response to the above problems, the present application provides a video editing method and related devices to improve editing efficiency.

[0006] Based on this, this application discloses the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides a video editing method, the method comprising:

[0008] Acquire a video to be edited and feature information of the video to be edited, wherein the video to be edited includes multiple sub-videos;

[0009] Generate text description information of the target sub-video based on the video frame and feature information of the target sub-video through a video description model, wherein the feature information and the video frame have a temporal correlation relationship, and the text description information is used to describe the plot content at different moments in the corresponding sub-video;

[0010] Taking each of the sub-videos as the target sub-video, and obtaining text description information of each of the sub-videos;

[0011] Based on the text description information of all the sub-videos, a text generation model is used to process the text description information to generate a plot script of the video to be edited, wherein the plot script is used to describe the plot content of the video to be edited at different moments;

[0012] According to the plot script of the video to be edited, a plurality of video segments associated with the plot script are determined, and a target video including each of the video segments is generated, where the duration of the target video is shorter than that of the video to be edited.

[0013] Optionally, the multiple sub-videos correspond to different scenes respectively, and the sub-videos include multiple shot segments, and the obtaining of the video to be edited and feature information of the video to be edited includes:

[0014] Acquire feature information of each shot segment in the video to be edited and each sub-video;

[0015] The step of processing the video frame of the target sub-video and the feature information of the target sub-video by a video description model to generate text description information of the target sub-video includes:

[0016] For the target shot segment of the target sub-video, processing the target shot segment by the video description model according to the video frame of the target shot segment and the feature information of the target shot segment to generate text description information of the target shot segment;

[0017] Taking each shot segment of the target sub-video as the target shot segment, and obtaining text description information of each shot segment in the target sub-video;

[0018] The text description information of all the shot segments in the target sub-video is processed by the text generation model to generate the text description information of the target sub-video.

[0019] Optionally, if the characteristic information includes subtitle information, the target shot segment of the target sub-video is processed by the video description model based on the video frame of the target shot segment and the characteristic information of the target shot segment to generate text description information of the target shot segment, including:

[0020] For the target shot segment of the target sub-video, processing the target shot segment through the video description model according to the video frame of the target shot segment, the feature information of the target shot segment, and the subtitle information of the target sub-video to generate text description information of the target shot segment;

[0021] The step of processing the text description information of all the sub-videos by a text generation model to generate a plot script of the video to be edited includes:

[0022] The text description information of all the sub-videos and the subtitle information of the video to be edited are processed by the text generation model to generate a plot script of the video to be edited.

[0023] Optionally, if the videos to be edited include multiple videos, and the contents of the videos to be edited are associated with each other, then the text description information of all the sub-videos is processed by a text generation model to generate a plot script of the videos to be edited, including:

[0024] Processing the text description information corresponding to all the videos to be edited by the text generation model to generate storyline information corresponding to each of the videos to be edited, wherein the storyline information is used to describe the global plot of each of the videos to be edited and the relationship between each of the videos to be edited;

[0025] According to the story line information and the text description information corresponding to each of the videos to be edited, the text generation model is used to process and generate a plot script for each of the videos to be edited.

[0026] Optionally, the object information includes an emotional parameter of an object in a video frame, the emotional parameter being used to identify an emotional state or intensity of the object in the video frame; the plot script includes a segment identifier, the segment identifier being used to indicate a time period in the sub-video during which the emotional parameter meets a preset high-energy segment condition; and determining, based on the plot script of the video to be edited, a plurality of video segments associated with the plot script comprises:

[0027] Determining, according to the plot script of the video to be edited, a plurality of candidate video segments associated with the plot script;

[0028] A candidate video segment including the segment identifier among the multiple candidate video segments is determined as the video segment associated with the plot script.

[0029] Optionally, the step of processing the text description information of all the sub-videos by a text generation model to generate a plot script of the video to be edited includes:

[0030] Generate a plot script for the main storyline of the video to be edited by processing the text description information of all the sub-videos and the first prompt word through the text generation model, wherein the first prompt word is used to indicate the generation of the plot script for the main storyline of the video to be edited;

[0031] The step of determining a plurality of video clips associated with the plot script of the video to be edited comprises:

[0032] According to the plot script of the main story line of the video to be edited, a plurality of video clips associated with the plot script of the main story line are determined.

[0033] Optionally, the step of processing the text description information of all the sub-videos by the text generation model to generate a plot script of the main story line of the video to be edited includes:

[0034] Generate a plot script for multiple story lines of the video to be edited based on the text description information of all the sub-videos and the second prompt word through the text generation model, wherein the second prompt word is used to indicate the generation of the plot script for multiple story lines of the video to be edited;

[0035] The step of determining, based on the plot script of the video to be edited, a plurality of video segments associated with the plot script, and generating the target video including the respective video segments, comprises:

[0036] According to the plot scripts of each story line of the video to be edited, multiple video clips associated with each story line are determined, and multiple target videos are generated. The multiple target videos correspond to different story lines respectively, and the target videos include multiple video clips associated with the corresponding story lines.

[0037] Optionally, the determining, based on the plot script of the video to be edited, a plurality of video segments associated with the plot script, and generating a target video including each of the video segments, comprises:

[0038] According to the plot script of the video to be edited and the text description information of all the sub-videos, the text generation model is used to process and generate a time information set of the video to be edited, and a target video including each of the video clips is generated based on the time information set, wherein the time information set includes multiple time information for indicating the start and end time of the video clips.

[0039] In a second aspect, an embodiment of the present application provides a video editing device, the device comprising:

[0040] An acquisition unit, configured to acquire a video to be edited and feature information of the video to be edited, wherein the video to be edited includes a plurality of sub-videos;

[0041] a processing unit configured to generate text description information of a target sub-video based on a video frame and feature information of the target sub-video by processing the target sub-video using a video description model, wherein the feature information is temporally associated with the video frame, and the text description information is used to describe the plot content at different moments in the corresponding sub-video;

[0042] The processing unit is further configured to use each of the sub-videos as the target sub-video and obtain text description information of each of the sub-videos;

[0043] The processing unit is further configured to process the text description information of all the sub-videos using a text generation model to generate a plot script of the video to be edited, wherein the plot script is used to describe the plot content of the video to be edited at different moments;

[0044] The editing unit is used to determine multiple video clips associated with the plot script of the video to be edited according to the plot script of the video to be edited, and generate a target video including each of the video clips, wherein the duration of the target video is shorter than that of the video to be edited.

[0045] In a third aspect, an embodiment of the present application provides a computer device, the computer device including a processor and a memory:

[0046] The memory is used to store a computer program and transmit the computer program to the processor;

[0047] The processor is configured to execute the method described in the first aspect above according to the computer program.

[0048] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method described in the first aspect above.

[0049] In a fifth aspect, an embodiment of the present application provides a computer program product comprising a computer program, which, when executed on a computer device, enables the computer device to execute the method described in the first aspect above.

[0050] It can be seen from the above technical solutions that this application has at least the following beneficial effects:

[0051] Obtain the video to be edited and its feature information, where the video to be edited includes multiple sub-videos. Based on the video frame of the target sub-video and the feature information of the target sub-video, process it through the video description model to generate text description information of the target sub-video. The feature information has a temporal correlation with the video frame. The text description information is used to describe the plot content at different moments in the corresponding sub-video, thereby avoiding the problem of excessive data volume caused by inputting the entire video to be edited into the model at one time, thereby effectively reducing the pressure on the model and improving the model processing efficiency. Take each sub-video as the target sub-video and obtain the text description information of each sub-video. By using the video description model to analyze the feature information in each sub-video, the core semantic content of the sub-video can be extracted more meticulously, thereby providing a high-quality semantic basis for the generation of subsequent plot scripts and avoiding incomplete plots due to missing information. Based on the text description information of all sub-videos, the text generation model is used to process them and generate a plot script for the video to be edited. The plot script is used to describe the plot content of the video to be edited at different moments. Thus, the plot script can achieve the theme summary and plot combing of the entire video to be edited, so that the subsequent editing process has a clear logical basis, improves the coherence and completeness of the editing results, and avoids the problems of arbitrary clip selection and lack of main line in traditional editing methods. According to the plot script of the video to be edited, multiple video clips associated with the plot script are determined, and a target video including each video clip is generated. The target video is shorter than the video to be edited. Based on the plot script, video clips that are highly relevant to the theme are accurately screened. This not only improves the automation level of video editing, but also significantly reduces the workload of manual supplementation and adjustment in the later stage, thereby greatly improving the overall efficiency and output quality of video editing.

[0052] By segmenting the video to be edited into multiple sub-videos, generating corresponding text descriptions, and further generating a plot script to guide the editing process, this method addresses existing issues such as heavy model burden, slow inference speed, and incomplete plot coverage caused by inputting a full, long video. Furthermore, by guiding the editing strategy with a plot script, the coherence and completeness of the editing results are improved, significantly reducing the possibility of post-production patching, thereby significantly improving video editing efficiency while maintaining video quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0054] Figure 1A schematic diagram of a video editing method according to an embodiment of the present invention;

[0055] Figure 2 A schematic diagram of a multi-level division of a video to be edited provided in an embodiment of the present application;

[0056] Figure 3 A schematic structural diagram of a video editing device provided in an embodiment of the present application;

[0057] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only and are not intended to limit the scope of protection of the present application.

[0059] In related technologies, a method of editing videos using a large model is proposed. By inputting the entire long video into the large model, the large model outputs the time information of each segment of the edited short video, thereby achieving efficient video editing.

[0060] However, this method requires the entire long video (which may contain tens of minutes or even longer content) to be directly input into the large model for processing. The large model has limited ability to process long sequence data, especially the huge amount of video frame data. This will not only lead to slow model inference speed, but also make it difficult for the output video clips to cover the important plots of the large model, resulting in information missing. After editing the short video, new video clips still need to be added to improve the completeness of the plot, resulting in low efficiency of video editing.

[0061] Based on this, the embodiments of the present application provide a video editing method and related apparatus. By splitting the video to be edited into multiple sub-videos, segmenting them and generating corresponding textual descriptions, and further generating a plot script to guide the editing process, this method addresses the existing problems of heavy model burden, slow inference speed, and incomplete plot coverage caused by inputting an entire long video. Furthermore, by guiding the editing strategy with a plot script, the coherence and completeness of the editing results are improved, significantly reducing the possibility of post-production patching, thereby significantly improving the efficiency of video editing while ensuring video quality.

[0062] The video editing method provided in this application can be applied to computer devices with data processing capabilities, such as terminal devices and servers. Specifically, the terminal devices may include desktop computers, laptop computers, mobile phones, and tablet computers. The server may be a standalone physical server, a server cluster composed of multiple physical servers, or a distributed system. The terminal devices and servers may be connected directly or indirectly via wired or wireless communication, and this application does not impose any restrictions thereon.

[0063] See also Figure 1 , which is a flow chart of the video editing method provided by the embodiment of the present application. For the sake of convenience, the following embodiment is introduced by taking the execution subject of the video editing method as the server as an example. Figure 1 As shown, the video editing method includes S101-S105.

[0064] S101: Obtain a video to be edited and feature information of the video to be edited.

[0065] All data collected by this application (such as videos to be edited and feature information) are collected with the consent and authorization of the subject to which the data belongs (such as users, institutions or enterprises), and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.

[0066] The video to be edited is a video that needs to be edited, such as a film clip, live broadcast recording, or documentary. The video to be edited includes multiple sub-videos, which are divided by time or content. For example, each sub-video corresponds to a different scene. For example, a 2-hour video to be edited can be divided into 10 sub-videos at 20-minute intervals.

[0067] The feature information is used to reflect the features of the video to be edited, for example, the feature information may be subtitle information, object information, etc.

[0068] Subtitles are textual information synchronized with the content of the video being edited. They can be textual information displayed in the video being edited, or they can be textual information not displayed in the video being edited and obtained through speech recognition technology. For example, subtitles could be lines from a character in a movie or commentary from a documentary. Subtitles provide semantic supplementary textual information, helping the model parse more accurately. Furthermore, subtitles can serve as an alternative source of information when audio or visual information is missing.

[0069] Object information is the characteristic information of objects appearing in the video to be edited. These objects can be people, anime characters, objects, scenes, and so on. For example, object information could be the name of a person in a video frame, such as {"time":100,"role":"Xia Ming"}, where time is a timestamp. Object information provides visual semantic information, enhancing the model's understanding of video content.

[0070] As a possible implementation method, optical character recognition (OCR) may be used to recognize each video frame of the video to be edited, thereby obtaining subtitle information and a corresponding timestamp of each video frame.

[0071] As a possible implementation method, the audio of the lines in the edited video can be clustered to identify the speaker (object) corresponding to each line, and then the objects in each video frame can be detected through image recognition technology to determine the names corresponding to the objects in each video frame.

[0072] The server can obtain feature information corresponding to each sub-video, and then perform subsequent processing on each sub-video.

[0073] S102: Processing the target sub-video frame and feature information of the target sub-video through a video description model to generate text description information of the target sub-video.

[0074] The video frames and the feature information have a temporal association relationship. For example, the video frames and the feature information can be associated through a timestamp.

[0075] The video description model processes information including video type and image type and generates corresponding textual information. For example, the video description model can be a pre-trained multimodal model that integrates sub-video feature information with video frames and generates textual description information corresponding to the sub-video. The textual description information is used to describe the plot content at different moments in the corresponding sub-video. The textual description information can include timelines, scenes, character actions, key events, etc.

[0076] For example, the text description information of sub-video 1 can be: "From 5 minutes 30 seconds to 6 minutes 20 seconds of the video to be edited, the protagonist A receives a mysterious call at home, and the content of the call involves a suspenseful incident." The text description content of sub-video 2 can be: "From 6 minutes 21 seconds to 8 minutes 20 seconds of the video to be edited, the protagonist A goes to the crime scene, finds the body and records the details of the scene."

[0077] The target sub-video is one of multiple sub-videos. Taking the target sub-video as an example, the video frame and feature information of the target sub-video can be used as input of the video description model, and processed by the video description model to obtain text description information of the target sub-video.

[0078] S103: Taking each sub-video as a target sub-video, and obtaining text description information of each sub-video.

[0079] By executing the above S102 for each sub-video as a target sub-video, text description information of each sub-video is obtained, which can retain semantic information related to the content of the video to be edited as much as possible, thereby providing a data basis for subsequent generation of a plot script.

[0080] S104: Processing the text description information of all sub-videos through a text generation model to generate a plot script for the video to be edited.

[0081] The plot script is a text description of the core plot, theme development and editing logic of the video to be edited, and is used to describe the plot content at different moments of the video to be edited.

[0082] For example, a plot script might read, "Protagonist A begins an investigation after receiving a mysterious letter (Cause), gradually uncovers the conspiracy behind the case (Development), and ultimately confronts the murderer on a rainy night and uncovers the truth (Climax)."

[0083] A text generation model processes text-based information and generates corresponding text-based information. For example, a text generation model can be a deep learning model based on natural language processing technology, which can generate coherent and logical text output based on input text information.

[0084] The text descriptions of each sub-video can then be used as input to a text generation model, which processes the information to generate the plot of the video to be edited. The contextual modeling capabilities of the text generation model can identify the causal relationships and temporal order between the sub-videos, ensuring the logical coherence of the plot.

[0085] S105: According to the plot script of the video to be edited, a plurality of video segments associated with the plot script are determined, and a target video including the respective video segments is generated.

[0086] The target video is the video obtained by editing the video to be edited. The target video's duration is shorter than the video to be edited. Video clips are selected from the video to be edited and are associated with the plot script. Video clips are segments of the video to be edited that are divided by the text generation model based on the plot distribution of the plot script. If a segment of the video to be edited has the same semantics or similarities to the plot script at a corresponding moment or time period, it is considered a video clip associated with the plot script. Video clips are typically short (e.g., 5-30 seconds) and are used to splice together to generate the target video.

[0087] For example, if the plot script is labeled "Protagonist A fights the enemy in the snow castle (climax)", the text generation model will determine that the sub-video includes video clips of "snow battle" and "explosion scene".

[0088] After determining the video segments for generating the target video, the video segments can be spliced together using video editing software to generate the target video.

[0089] In one possible implementation, the plot script of the video to be edited and the text description information of all sub-videos can be processed through a text generation model to generate a time information set of the video to be edited, and a target video including various video clips can be generated based on the time information set. The time information set includes multiple time information for indicating the start and end time of the video clips.

[0090] For example, a text generation model can take the text descriptions of all sub-videos, the plot of the video to be edited, and a prompt word template as input, and process them to generate time information for multiple video segments associated with the plot, thereby determining the position of each video segment in the video to be edited. The prompt word template is used to instruct the text generation model to output a time list of each video segment, which facilitates subsequent editing.

[0091] Therefore, there is no need to output the corresponding video clip. The video clip can be accurately located through the start and end time indicated by the time information set, which makes it easier to capture the corresponding video clip in the video to be edited and generate the subsequent target video, thereby improving editing efficiency.

[0092] When locating video clips through a text generation model, there may be issues with extremely short shots and dialogue truncation. To address this, we can set a minimum video duration threshold to discard video clips shorter than the threshold, or locate extremely short shots and replace them with video clips of a preset extended duration, either before or after the scene in which they occur. This ensures that the target video does not contain fragmented video clips that affect the viewing experience. Regarding dialogue truncation, we can identify the audio information of the dialogue, locate the start and end timestamps of the dialogue, and replace the video clips corresponding to the start and end timestamps with the video clips that appear during the dialogue phase.

[0093] As a possible implementation method, the text description information corresponding to each video clip can be processed through a text generation model to obtain the keywords between each video clip. According to the correspondence between the keywords and the video transition effects, the video transition effects between each video clip are determined and applied to the generation process of the target video, thereby obtaining a target video that conforms to the plot development logic and has a stronger visual effect.

[0094] As a possible implementation, a duration threshold of the target video may be obtained, and multiple video segments associated with the plot script may be determined based on the duration threshold, where the total duration of each video segment is less than the duration threshold.

[0095] As a possible implementation method, the plot script can be split according to the correspondence between the plot script and the video clips, and the narration subtitles of each video clip can be obtained and embedded into the target video. This not only generates a target video that meets the audience's quick viewing needs, but also provides corresponding explanation supplements to improve the content richness of the target video.

[0096] As can be seen from the above technical solution, the video to be edited and the feature information of the video to be edited are obtained, and the video to be edited includes multiple sub-videos. According to the video frame of the target sub-video and the feature information of the target sub-video, the video description model is used to process and generate text description information of the target sub-video. The feature information has a time correlation relationship with the video frame. The text description information is used to describe the plot content at different moments in the corresponding sub-video, thereby avoiding the problem of excessive data volume caused by inputting the entire video to be edited into the model at one time, thereby effectively reducing the pressure on the model and improving the model processing efficiency. Each sub-video is used as a target sub-video, and the text description information of each sub-video is obtained. By using the video description model to analyze the feature information in each sub-video, the core semantic content of the sub-video can be extracted more meticulously, thereby providing a high-quality semantic basis for the generation of subsequent plot scripts and avoiding incomplete plots due to missing information. Based on the text description information of each sub-video, it is processed through a text generation model to generate a plot script for the video to be edited. The plot script is used to describe the plot content of the video to be edited at different moments. Thus, the plot script can achieve the theme summary and plot combing of the entire video to be edited, so that the subsequent editing process has a clear logical basis, improves the coherence and completeness of the editing results, and avoids the problems of arbitrary clip selection and lack of main line in traditional editing methods. According to the plot script of the video to be edited, multiple video clips associated with the plot script are determined, and a target video including each video clip is generated. The target video is shorter than the video to be edited. Based on the plot script, video clips that are highly relevant to the theme are accurately screened out, which not only improves the automation level of video editing, but also significantly reduces the workload of manual supplementation and adjustment in the later stage, thereby greatly improving the overall efficiency and output quality of video editing.

[0097] By segmenting the video to be edited into multiple sub-videos, generating corresponding text descriptions, and further generating a plot script to guide the editing process, this method addresses existing issues such as heavy model burden, slow inference speed, and incomplete plot coverage caused by inputting a full, long video. Furthermore, by guiding the editing strategy with a plot script, the coherence and completeness of the editing results are improved, significantly reducing the possibility of post-production patching, thereby significantly improving video editing efficiency while maintaining video quality.

[0098] This embodiment of the application provides a further division method for editing videos, see Figure 2 , Figure 2This embodiment of the present application provides a schematic diagram of a multi-level division of a video to be edited. A video to be edited includes multiple sub-videos based on scene division. Scenes are divided into multiple independent scenes (such as chapters or continuous dialogue segments in a movie) based on plot logic or visual features. The sub-videos can be divided using a scene segmentation algorithm. In addition, each scene includes multiple shot clips. A shot clip is the smallest video unit in the sub-video continuously captured by a single camera. The sub-videos can be divided using a shot segmentation algorithm. For the specific implementation of obtaining text description information for the target sub-video based on this video hierarchical division method, see A1-A4:

[0099] A1: Obtain feature information of each shot segment in the video to be edited and each sub-video.

[0100] From the above, it can be seen that sub-videos are divided based on scenes, and different sub-videos correspond to different scenes. For the sub-videos of each scene, the characteristic information of each shot segment included in the sub-video is obtained, so that the shot segment is used as the smallest information acquisition unit, reducing the degree of information loss in the process of generating the plot script. In addition, splitting large sub-videos into shot segment-level processing units can also reduce the amount of data processed by a single model and improve processing efficiency.

[0101] A2: For the target shot segment of the target sub-video, the video description model is used to process the target shot segment based on the video frame and feature information of the target shot segment to generate text description information of the target shot segment.

[0102] The target shot segment is one of the multiple shot segments included in the target sub-video. Taking the target shot segment as an example, the video frame of the target shot segment and the feature information of the target shot segment are used as input, and are processed by the video description model to generate text description information of the target shot segment.

[0103] A3: Taking each shot segment of the target sub-video as a target shot segment, and obtaining text description information of each shot segment in the target sub-video.

[0104] Each shot segment of the target sub-video is used as a target shot segment, and A2 is executed to obtain text description information of each shot segment in the target sub-video.

[0105] A4: Based on the text description information of all shot segments in the target sub-video, the text description information of the target sub-video is generated by processing the text description information through the text generation model.

[0106] The text descriptions of all the shot segments in the target sub-video are fed into the text generation model to generate a unified text description for the sub-video. This approach semantically integrates the local shot segments, consolidating the shot-level descriptions into sub-video-level text descriptions to preserve the core plot. The text generation model identifies causal relationships between shot segments, ensuring the logicality of the text descriptions for the corresponding scenes in the sub-video.

[0107] Therefore, by further segmenting the sub-videos, the amount of data processed by a single model is reduced, improving processing speed. Parallel computing can further shorten overall processing time. By segmenting and generating corresponding text description information at the shot-by-shot level, we ensure that no details are missed and reduce the possibility of information loss due to the length of the sub-videos. In addition, the segmentation method based on scenes and shot segments conforms to the shooting logic and narrative logic of the video to be edited, thereby generating more logical text description information and further improving the accuracy of the plot script.

[0108] When using shot clips or sub-videos as model input, subtitles may cross scenes or shot clips. Based on this, in order to reduce the loss of semantic information caused by subtitles, this application provides an implementation method that uses complete subtitle information as auxiliary information in the process of generating text description information.

[0109] Specifically, taking a target shot segment as an example, the video description model processes the target shot segment's video frame, its feature information, and the subtitle information of the target subvideo to generate a text description of the target shot segment. For each shot segment's video frame and feature information as input, the subtitle information of the subvideo containing the shot segment is also fed into the text generation model as auxiliary information.

[0110] The text description information of all sub-videos and the subtitle information of the video to be edited are processed by the text generation model to generate the plot script of the video to be edited. In this case, when the text description information of each sub-video is used as input, the subtitle information of all sub-videos is also input into the text generation model as auxiliary information.

[0111] Therefore, by introducing more complete subtitle information as auxiliary input, the problem of semantic information loss caused by subtitles crossing scenes or shot segments during video editing is solved, and the accuracy and coherence of text description information and plot script generation are significantly improved.

[0112] If there are multiple videos to be edited, and the contents of each video to be edited are related, such as a multi-episode TV series, in this case, only inputting one video to be edited into the model will not be able to grasp the development context of the multiple related videos to be edited. Based on this, the embodiment of the present application provides a specific implementation method for introducing storyline information, see B1-B2:

[0113] B1: Based on the text description information corresponding to each video to be edited, the text generation model is used to process the text description information to generate the story line information corresponding to each video to be edited.

[0114] Since there is a correlation between the contents of each video to be edited, if the text description information corresponding to each video to be edited is directly used as the plot script, the overall development context of the story may be ignored, resulting in plot fragmentation between the plot scripts of each video to be edited.

[0115] Based on this, the text description information corresponding to each video to be edited can be processed again through the text generation model to generate the story line information corresponding to each video to be edited.

[0116] Among them, storyline information is information that describes the overall plot development of multiple content-related videos (such as multi-episode TV series). Storyline information can include the main events, cause and effect relationships, time sequence and theme orientation of each video to be edited. Storyline information is used to describe the global plot and the correlation relationship across videos.

[0117] B2: Based on the storyline information and the text description information corresponding to each video to be edited, the text generation model is used to process the information and generate the plot script of each video to be edited.

[0118] Since the text description information corresponding to each video to be edited is relatively independent and fragmented, the story line information used to describe the overall plot development context can be input into the text generation model, so that the text generation model can understand the relationship between the contents of each video to be edited, and rewrite the text description information corresponding to each video to be edited based on the story line information to generate a plot script for each video to be edited.

[0119] Therefore, in order to solve the problems of plot fragmentation, theme deviation and logical break caused by independent processing in the editing of multiple content-related videos to be edited, by generating storyline information and mapping it back to the text description information of each video to be edited, the plot script of each video to be edited is made to fit the overall development context and the overall storyline, which significantly improves the integrity and narrative coherence of the target video, avoids the subsequent rewriting of the plot script, and improves editing efficiency.

[0120] In one possible implementation, the object information may further include an emotional parameter of the object in the video frame. The emotional parameter is used to identify the emotional state or intensity of the object in the video frame. For example, the emotional parameter may be "anger" (intensity 0.9), which can be determined for each video frame using image recognition technology. The plot script includes a segment identifier, which is used to indicate a time period in the sub-video when the emotional parameter meets a preset high-energy segment condition. The preset high-energy segment condition is a preset condition for determining a high-energy segment. For example, the preset high-energy segment condition may be:

[0121] (1) The video segments in which the emotion parameters are greater than the preset parameter threshold exceed a fixed proportion.

[0122] (2) The video segments where the video frames whose emotion parameters are greater than the maximum parameter threshold are located.

[0123] (3) Video segments in which the difference in emotion parameters between adjacent video frames exceeds a preset difference threshold.

[0124] Since the object information includes emotional parameters, the preset high-energy segment conditions can be used as the input of the text generation model, so that in the process of generating the plot script, the video segments that meet the conditions can be identified, and the specific time period can be located and represented by segment identifiers.

[0125] Finally, according to the plot script of the video to be edited, a plurality of video segments associated with the plot script and corresponding to the segment identifiers are determined.

[0126] Specifically, based on the plot script of the video to be edited, multiple candidate video segments associated with the plot script are determined, where the candidate video segments are segments that have not been screened using segment identifiers. Among the multiple candidate video segments, the candidate video segments that include segment identifiers are determined as video segments associated with the plot script. Since the segment identifiers are used to indicate the time periods in the sub-videos where the emotional parameters meet the preset high-energy segment conditions, i.e., to locate the high-energy segments, the multiple candidate video segments can be further screened using the segment identifiers as screening conditions to obtain video segments associated with the plot script.

[0127] Therefore, by analyzing the emotional parameters, a plot script including segment identification is generated, so as to accurately locate the high-energy segments in the video to be edited and edit them into the target video, realizing the accurate identification and automatic positioning of high-energy segments in video editing, and improving editing efficiency.

[0128] The plot content of the video to be edited can include multiple storylines, such as a main storyline and branch storylines. The main storyline is the core plot thread throughout the entire video and is the most critical narrative thread in the video. The branch storylines are auxiliary plots that revolve around the main storyline. Based on this, the text description information of all sub-videos can be processed through a text generation model to generate a plot script for the main storyline of the video to be edited. For example, a first prompt word can be input into the text generation model so that the model clearly defines the main storyline as the core goal of editing and generates a corresponding plot script. The first prompt word is used to indicate the generation of a plot script for the main storyline of the video to be edited.

[0129] Then, according to the plot script of the main story line of the video to be edited, multiple video clips associated with the plot script of the main story line are determined, and a target video including the multiple video clips is generated.

[0130] Therefore, by editing video clips that are only related to the main storyline into the target video, editors do not need to analyze the main plot and secondary plots, and can exclude video clips with auxiliary plots and generate a target video that only narrates the main storyline, which makes it easier for viewers to quickly understand the core content of the video to be edited and improves editing efficiency.

[0131] Furthermore, the text description information of all sub-videos can be processed through a text generation model to generate plot scripts for multiple story lines of the video to be edited.

[0132] The multiple storylines may include a main storyline. Specifically, the text generation model can generate multiple plot scripts, each corresponding to a storyline, to subsequently generate target videos corresponding to different storylines. For example, a second prompt word can be input into the text generation model to cause the model to clearly identify the core objective of editing the sub-storylines and generate multiple corresponding plot scripts. The second prompt word is used to indicate the generation of plot scripts for the multiple storylines of the video to be edited.

[0133] Then, according to the plot scripts of each story line of the video to be edited, multiple video clips respectively associated with each story line are determined, and multiple target videos are generated.

[0134] The multiple target videos correspond to different storylines, and the target videos include multiple video clips associated with the corresponding storylines. For example, if a plot script 1 for storyline A and a plot script 2 for storyline B are generated using a text generation model, the multiple video clips associated with storyline A can be determined based on plot script 1, and target video 1 corresponding to storyline A can be generated. The multiple video clips associated with storyline B can be determined based on plot script 2, and target video 2 corresponding to storyline B can be generated.

[0135] As a possible implementation method, target videos corresponding to multiple story lines can be merged to obtain one target video.

[0136] Therefore, for videos to be edited with multiple narrative lines, corresponding target videos can be generated for each story line, so that editors do not need to manually organize the plots of different story lines and match the corresponding video clips, and can achieve automatic editing of multiple story lines, thereby improving editing efficiency.

[0137] See also Figure 3 , Figure 3 The embodiment of the present application provides a video editing device, the device 300 including:

[0138] An acquiring unit 301 is configured to acquire a video to be edited and feature information of the video to be edited, wherein the video to be edited includes a plurality of sub-videos;

[0139] A processing unit 302 is configured to process the target sub-video frame and feature information of the target sub-video using a video description model to generate text description information of the target sub-video, wherein the feature information is temporally associated with the video frame, and the text description information is used to describe the plot content at different moments in the corresponding sub-video;

[0140] The processing unit 302 is further configured to use each of the sub-videos as the target sub-video and obtain text description information of each of the sub-videos;

[0141] The processing unit 302 is further configured to process the text description information of all the sub-videos using a text generation model to generate a plot script of the video to be edited, wherein the plot script is used to describe the plot content of the video to be edited at different moments;

[0142] The editing unit 303 is used to determine multiple video clips associated with the plot script of the video to be edited, and generate a target video including each of the video clips, where the duration of the target video is shorter than that of the video to be edited.

[0143] As a possible implementation, the multiple sub-videos correspond to different scenes respectively, and the sub-videos include multiple shot segments. Then, the acquiring unit 301 is specifically configured to:

[0144] Acquire feature information of each shot segment in the video to be edited and each sub-video;

[0145] The processing unit 302 is specifically configured to:

[0146] For the target shot segment of the target sub-video, processing the target shot segment by the video description model according to the video frame of the target shot segment and the feature information of the target shot segment to generate text description information of the target shot segment;

[0147] Taking each shot segment of the target sub-video as the target shot segment, and obtaining text description information of each shot segment in the target sub-video;

[0148] The text description information of all the shot segments in the target sub-video is processed by the text generation model to generate the text description information of the target sub-video.

[0149] As a possible implementation, if the feature information includes subtitle information, the processing unit 302 is specifically configured to:

[0150] For the target shot segment of the target sub-video, processing the target shot segment through the video description model according to the video frame of the target shot segment, the feature information of the target shot segment, and the subtitle information of the target sub-video to generate text description information of the target shot segment;

[0151] The text description information of all the sub-videos and the subtitle information of the video to be edited are processed by the text generation model to generate a plot script of the video to be edited.

[0152] As a possible implementation, if the to-be-edited videos include multiple videos, and the contents of the to-be-edited videos are associated with each other, the processing unit 302 is specifically configured to:

[0153] Processing the text description information corresponding to all the videos to be edited by the text generation model to generate storyline information corresponding to each of the videos to be edited, wherein the storyline information is used to describe the global plot of each of the videos to be edited and the relationship between each of the videos to be edited;

[0154] According to the story line information and the text description information corresponding to each of the videos to be edited, the text generation model is used to process and generate a plot script for each of the videos to be edited.

[0155] As a possible implementation, the object information includes an emotional parameter of an object in a video frame, where the emotional parameter is used to identify the emotional state or intensity of the object in the video frame. The plot script includes a segment identifier, where the segment identifier is used to indicate a time period in the sub-video when the emotional parameter meets a preset high-energy segment condition. The editing unit 303 is specifically configured to:

[0156] Determining, according to the plot script of the video to be edited, a plurality of candidate video segments associated with the plot script;

[0157] A candidate video segment including the segment identifier among the multiple candidate video segments is determined as the video segment associated with the plot script.

[0158] As a possible implementation, the processing unit 302 is specifically configured to:

[0159] Generate a plot script for the main storyline of the video to be edited by processing the text description information of all the sub-videos and the first prompt word through the text generation model, wherein the first prompt word is used to indicate the generation of the plot script for the main storyline of the video to be edited;

[0160] The editing unit 303 is specifically configured to:

[0161] According to the plot script of the main story line of the video to be edited, a plurality of video clips associated with the plot script of the main story line are determined.

[0162] As a possible implementation, the processing unit 302 is specifically configured to:

[0163] Generate a plot script for multiple story lines of the video to be edited based on the text description information of all the sub-videos and the second prompt word through the text generation model, wherein the second prompt word is used to indicate the generation of the plot script for multiple story lines of the video to be edited;

[0164] The editing unit 303 is specifically configured to:

[0165] According to the plot scripts of each story line of the video to be edited, multiple video clips associated with each story line are determined, and multiple target videos are generated. The multiple target videos correspond to different story lines respectively, and the target videos include multiple video clips associated with the corresponding story lines.

[0166] As a possible implementation, the editing unit 303 is specifically configured to:

[0167] According to the plot script of the video to be edited and the text description information of all the sub-videos, the text generation model is used to process and generate a time information set of the video to be edited, and a target video including each of the video clips is generated based on the time information set, wherein the time information set includes multiple time information for indicating the start and end time of the video clips.

[0168] See also Figure 4, an embodiment of the present application further provides a computer device, the computer device comprising a memory 401 and a processor 402:

[0169] The memory is used to store a computer program and transmit the computer program to the processor;

[0170] The processor is configured to execute the method of the above method embodiment according to the computer program.

[0171] An embodiment of the present application further provides a computer-readable storage medium, characterized in that the computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method of the above method embodiment.

[0172] An embodiment of the present application further provides a computer program product including a computer program, which, when executed on a computer device, enables the computer device to execute the method of the above method embodiment.

[0173] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the systems or devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.

[0174] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0175] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0176] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0177] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0178] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A video editing method, characterized in that: The method comprises: Acquire a video to be edited and feature information of the video to be edited, wherein the video to be edited includes multiple sub-videos; Generate text description information of the target sub-video based on the video frame and feature information of the target sub-video through a video description model, wherein the feature information and the video frame have a temporal correlation relationship, and the text description information is used to describe the plot content at different moments in the corresponding sub-video; Taking each of the sub-videos as the target sub-video, and obtaining text description information of each of the sub-videos; Based on the text description information of all the sub-videos, a text generation model is used to process the text description information to generate a plot script of the video to be edited, wherein the plot script is used to describe the plot content of the video to be edited at different moments; According to the plot script of the video to be edited, a plurality of video segments associated with the plot script are determined, and a target video including each of the video segments is generated, where the duration of the target video is shorter than that of the video to be edited.

2. The method according to claim 1, characterized in that The multiple sub-videos correspond to different scenes respectively, and the sub-videos include multiple shot segments. Then, obtaining the video to be edited and the feature information of the video to be edited includes: Acquire feature information of each shot segment in the video to be edited and each sub-video; The step of processing the video frame of the target sub-video and the feature information of the target sub-video by a video description model to generate text description information of the target sub-video includes: For the target shot segment of the target sub-video, processing the target shot segment by the video description model according to the video frame of the target shot segment and the feature information of the target shot segment to generate text description information of the target shot segment; Taking each shot segment of the target sub-video as the target shot segment, and obtaining text description information of each shot segment in the target sub-video; The text description information of all the shot segments in the target sub-video is processed by the text generation model to generate the text description information of the target sub-video.

3. The method according to claim 2, characterized in that If the characteristic information includes subtitle information, the target shot segment of the target sub-video is processed by the video description model based on the video frame of the target shot segment and the characteristic information of the target shot segment to generate text description information of the target shot segment, including: For the target shot segment of the target sub-video, processing the target shot segment through the video description model according to the video frame of the target shot segment, the feature information of the target shot segment, and the subtitle information of the target sub-video to generate text description information of the target shot segment; The step of processing the text description information of all the sub-videos by a text generation model to generate a plot script of the video to be edited includes: The text description information of all the sub-videos and the subtitle information of the video to be edited are processed by the text generation model to generate a plot script of the video to be edited.

4. The method according to claim 1, wherein If the videos to be edited include multiple videos, and the contents of the videos to be edited are related, then the text description information of all the sub-videos is processed by a text generation model to generate a plot script of the videos to be edited, including: Processing the text description information corresponding to all the videos to be edited by the text generation model to generate storyline information corresponding to each of the videos to be edited, wherein the storyline information is used to describe the global plot of each of the videos to be edited and the relationship between each of the videos to be edited; According to the story line information and the text description information corresponding to each of the videos to be edited, the text generation model is used to process and generate a plot script for each of the videos to be edited.

5. The method according to claim 1, wherein The object information includes an emotional parameter of an object in a video frame, the emotional parameter being used to identify an emotional state or intensity of the object in the video frame; the plot script includes a segment identifier, the segment identifier being used to indicate a time period in the sub-video during which the emotional parameter meets a preset high-energy segment condition; and determining, based on the plot script of the video to be edited, a plurality of video segments associated with the plot script, including: Determining, according to the plot script of the video to be edited, a plurality of candidate video segments associated with the plot script; A candidate video segment including the segment identifier among the multiple candidate video segments is determined as the video segment associated with the plot script.

6. The method according to claim 1, characterized in that The step of processing the text description information of all the sub-videos by a text generation model to generate a plot script of the video to be edited includes: Generate a plot script for the main storyline of the video to be edited by processing the text description information of all the sub-videos and the first prompt word through the text generation model, wherein the first prompt word is used to indicate the generation of the plot script for the main storyline of the video to be edited; The step of determining a plurality of video clips associated with the plot script of the video to be edited comprises: According to the plot script of the main story line of the video to be edited, a plurality of video clips associated with the plot script of the main story line are determined.

7. The method according to claim 1, characterized in that The step of processing the text description information of all the sub-videos by the text generation model to generate a plot script of the main story line of the video to be edited includes: Generate a plot script for multiple story lines of the video to be edited based on the text description information of all the sub-videos and the second prompt word through the text generation model, wherein the second prompt word is used to indicate the generation of the plot script for multiple story lines of the video to be edited; The step of determining, based on the plot script of the video to be edited, a plurality of video segments associated with the plot script, and generating the target video including the respective video segments, comprises: According to the plot scripts of each story line of the video to be edited, multiple video clips associated with each story line are determined, and multiple target videos are generated. The multiple target videos correspond to different story lines respectively, and the target videos include multiple video clips associated with the corresponding story lines.

8. The method according to any one of claims 1 to 7, characterized in that The step of determining a plurality of video segments associated with the plot script of the video to be edited, and generating a target video including each of the video segments, comprises: According to the plot script of the video to be edited and the text description information of all the sub-videos, the text generation model is used to process and generate a time information set of the video to be edited, and a target video including each of the video segments is generated based on the time information set, wherein the time information set includes multiple time information for indicating the start and end time of the video segments.

9. A video editing device, characterized in that: The device comprises: An acquisition unit, configured to acquire a video to be edited and feature information of the video to be edited, wherein the video to be edited includes a plurality of sub-videos; a processing unit configured to generate text description information of a target sub-video based on a video frame and feature information of the target sub-video by processing the target sub-video using a video description model, wherein the feature information is temporally associated with the video frame, and the text description information is used to describe the plot content at different moments in the corresponding sub-video; The processing unit is further configured to use each of the sub-videos as the target sub-video and obtain text description information of each of the sub-videos; The processing unit is further configured to process the text description information of all the sub-videos using a text generation model to generate a plot script of the video to be edited, wherein the plot script is used to describe the plot content of the video to be edited at different moments; The editing unit is used to determine multiple video clips associated with the plot script of the video to be edited according to the plot script of the video to be edited, and generate a target video including each of the video clips, wherein the duration of the target video is shorter than that of the video to be edited.

10. A computer device, characterized in that: The computer device includes a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1 to 8 according to the computer program.

Citation Information

Cited By

  • Media material processing method and device, equipment, storage medium and program product

    CN120935405A