Method, device, equipment, storage medium and program product for editing video
By segmenting the initial video into scenes and associating them with story units, edited sub-videos are generated, solving the problem of people's difficulty in efficiently generating short videos and achieving efficient editing and editing effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
- Filing Date
- 2025-06-23
- Publication Date
- 2026-07-21
Smart Images

Figure CN120658913B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, specifically to artificial intelligence technologies such as computer vision, video processing, and intelligent editing, and particularly to methods, apparatuses, electronic devices, computer-readable storage media, and computer program products for video editing. Background Technology
[0002] With the development of society and technology, people have gradually entered the information age. In the information age, long-form video content provides a large portion of internet traffic, and it has become an important content format in streaming media platforms, social media, and online education.
[0003] However, with changing lifestyles and considering the different paces of life, many people may not always be willing to invest a lot of time and energy in watching these "long videos." Or rather, people may not always be able to spare enough time and energy to watch "long videos" like movies and TV series in their entirety.
[0004] In this context, to improve the efficiency of people's access to video content and enable them to conveniently and efficiently consume video content during fragmented time, "short videos" created from "long videos" are becoming increasingly popular. Therefore, how to generate related "short videos" more efficiently and with higher quality from "long videos" is a noteworthy and urgent need. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for editing video.
[0006] In a first aspect, embodiments of this disclosure propose a method for editing a video, comprising: segmenting an initial video into scenes to obtain scene videos, wherein the scene videos correspond to scenes in the initial video; associating the scene videos with plot units of the initial video; extracting video frames from the initial video based on subtitles included in the scene videos associated with the plot units, and generating edited sub-videos corresponding to the plot units; and generating an edited video of the initial video based on the edited sub-videos.
[0007] Secondly, embodiments of this disclosure provide an apparatus for editing videos, comprising: a scene video segmentation unit configured to segment an initial video into scenes to obtain scene videos, wherein the scene videos correspond to scenes in the initial video; a scene video association unit configured to associate the scene videos with plot units of the initial video; an edited sub-video generation unit configured to extract video frames from the initial video based on subtitles included in the scene videos associated with the plot units, and generate edited sub-videos corresponding to the plot units; and an edited video generation unit configured to generate an edited video of the initial video based on the edited sub-videos.
[0008] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement a method for editing video as described in any implementation of the first aspect.
[0009] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer to perform a method for editing video as described in any implementation of the first aspect.
[0010] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement the method of editing video as described in any implementation of the first aspect.
[0011] The video editing method, apparatus, electronic device, computer-readable storage medium, and computer program product provided in this disclosure firstly segment an initial video into scenes to obtain scene videos, wherein the scene videos correspond to scenes in the initial video; then, the scene videos are associated with plot units of the initial video; next, based on the subtitles included in the scene videos associated with the plot units, video frames are extracted from the initial video to generate edited sub-videos corresponding to the plot units; finally, based on the edited sub-videos, an edited video of the initial video is generated.
[0012] This disclosure can improve the performance and efficiency of video editing and material preparation, and reduce the difficulty of video editing.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is an exemplary system architecture to which this disclosure can be applied; Figure 2 A flowchart illustrating a method for editing video provided in this disclosure embodiment; Figure 3 A flowchart illustrating another method for editing video provided in this disclosure embodiment; Figure 4 A schematic diagram of the architecture for implementing the video editing process in an application scenario provided by an embodiment of this disclosure; Figure 5 A structural block diagram of a video editing apparatus provided in an embodiment of this disclosure; Figure 6 This is a schematic diagram of the structure of an electronic device suitable for performing a method of video editing, provided as an embodiment of the present disclosure. Detailed Implementation
[0015] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0016] Furthermore, the acquisition, storage, use, processing, transportation, provision, and disclosure of user personal information (such as videos, subtitles, etc. that may be provided by users or include user personal information, as discussed later in this disclosure) in the technical solutions disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0017] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of methods, apparatuses, electronic devices, and computer-readable storage media for video editing can be applied.
[0018] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0019] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include video editing applications, short video generation applications, and instant messaging applications.
[0020] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.
[0021] Server 105 can provide various services through its built-in applications. For example, a video editing application that can edit the initial video can achieve the following effects when running this application: First, it obtains the initial video (which needs to be edited as "material") from terminal devices 101, 102, and 103 via network 104; then, server 105 segments the initial video into scene videos, where each scene corresponds to a scene in the initial video; next, server 105 associates the scene videos with the story units of the initial video; then, based on the subtitles included in the scene videos associated with the story units, server 105 extracts video frames from the initial video to generate edited sub-videos corresponding to the story units; finally, server 105 generates an edited video of the initial video based on the edited sub-videos.
[0022] It should be noted that, in addition to being obtained from terminal devices 101, 102, and 103 via network 104, the initial video can also be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (for example, when starting to process previously reserved video editing tasks), it can choose to directly obtain this data from locally. In this case, the exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.
[0023] Since scene segmentation and associating scene videos with plot units require significant computing resources and capabilities, the video editing methods provided in the subsequent embodiments of this disclosure are generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the video editing device is also generally located within the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also perform the aforementioned calculations performed by the server 105 through video editing applications installed on them, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, if the video editing application determines that the terminal device it is using has strong computing power and abundant remaining computing resources, it can allow the terminal device to perform the aforementioned calculations, thereby appropriately reducing the computing pressure on the server 105. Accordingly, the video editing device can also be located within terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.
[0024] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0025] Please refer to Figure 2 , Figure 2 A flowchart of a video editing process provided for an embodiment of this disclosure includes process 200.
[0026] Process 200 specifically includes the following steps: Step 201: Perform scene segmentation on the initial video to obtain scene videos; In embodiments of this disclosure, this step is intended to be performed by the entity executing the video editing method (e.g., Figure 1 The server 105 shown segments the initial video (e.g., a video uploaded by a user via terminal devices 101, 102, 103 that is intended to be edited) according to scenes to obtain scene videos. Accordingly, the obtained scene videos (i.e., segments of the initial video) can correspond to scenes in the initial video.
[0027] For example, the executing entity can use a scene-consistency representation learning scheme, such as the SceneConsistency Representation Learning (SCRL) model, to segment the initial video into scenes and treat the video portions and content of the initial video that correspond to or belong to the same "scene" as scene videos corresponding to that "same scene." Accordingly, the scene video can also be understood as a video segment composed of video frames belonging to the same "scene" in the initial video.
[0028] It should be noted that the initial video can be obtained directly from the local storage device by the aforementioned executing entity, or it can be obtained from a non-local storage device (e.g., Figure 1 The initial video can be obtained from the terminal devices 101, 102, and 103 shown. The local storage device can be a data storage module set within the aforementioned execution entity, such as a server hard drive. In this case, the initial video can be quickly read locally. The non-local storage device can also be any other electronic device set up to store data, such as some user terminals. In this case, the aforementioned execution entity can obtain the required initial video by sending an acquisition command to the electronic device.
[0029] Step 202: Associate the scene video with the story unit of the initial video; In the embodiments of this disclosure, based on step 201, this step aims to have the execution entity, after completing the scene segmentation of the initial video and obtaining the scene videos, associate each obtained scene video with a plot unit of the initial video. That is, after obtaining the scene videos, the execution entity determines the corresponding plot unit associated with each scene video.
[0030] In some optional implementations of this embodiment, the story unit can be obtained by the executing entity through splitting the complete story of the initial video. For example, after obtaining the complete story of the initial video, the executing entity can obtain the individual story units that make up the "complete story" by splitting the "complete story".
[0031] In some embodiments, the complete storyline may be provided by the provider or producer of the initial video, or obtained by the executing entity from the "user" through interaction with the user requesting video editing.
[0032] In general, a "complete storyline" can be the video and plot content included in the "initial video" as described in text form. For example, if the initial video is a movie, the "complete storyline" can be "textual information" obtained by combining the textual descriptions and information of the various events included in the movie according to the order in which the events are shown and performed in the movie.
[0033] Accordingly, the implementing entity can, based on its understanding of the "complete plot," divide the complete plot into multiple plot units according to unit division criteria such as "whether it includes a complete event." For example, under such criteria, each plot unit can completely include an "event."
[0034] In some optional implementations of this embodiment, after determining the scene video based on this step, the executing entity may further select to acquire and read the subtitles included in the scene video (i.e., the subtitles in the initial video that belong to the scene video portion).
[0035] Then, the executing entity compares the subtitles included in the scene video with the text description information of each plot unit to determine the semantic similarity between the subtitles and each text description information, and obtains the semantic matching results between the subtitles included in the scene video and the text description information of each plot unit in the initial video.
[0036] For example, the executing entity can perform semantic matching on the subtitle text {s1,....sn} of each scene video and the text description information {r1, ....rk} of each plot unit in chronological order to determine the semantic similarity between the two and obtain the "semantic matching result".
[0037] It should be understood that the scene videos corresponding to and associated with the same plot unit do not necessarily have to be sequential in time, or in other words, they do not necessarily have to be without any time gaps. This is to avoid loss of context.
[0038] Accordingly, after obtaining the semantic matching results, the executing entity can associate the scene video with the target plot unit with the highest semantic similarity in the semantic matching results to complete the "association" between the scene video and the plot unit. Thus, the executing entity can complete the "coarse correspondence" between the scene video and the plot (unit) (i.e., the correspondence between the plot unit and the scene video) through semantic matching between "subtitles" and "plot".
[0039] Accordingly, after completing the above "coarse correspondence", the subsequent implementing entity can further associate and determine the "video frames" of the corresponding initial video in each plot unit through the "fine correspondence" (i.e., the correspondence between plot units and video frames) discussed below.
[0040] Therefore, by using this coarse-to-fine two-level correspondence method, not only can the correspondence quality be improved, but also the excessive demand and use of computing resources in a single operation can be avoided, thus reducing the configuration requirements for computing resources.
[0041] In some embodiments, after obtaining the complete storyline of the initial video, the executing entity may choose to invoke a large language model to split the complete storyline into multiple storyline units.
[0042] A Large Language Model (LLM) is an artificial intelligence model that can be used to understand and generate human language, and based on its understanding, perform corresponding processing operations to obtain corresponding results. For example, after obtaining a complete storyline, pre-configured prompts can be used to instruct the LLM to break down the complete storyline into multiple storyline units according to certain criteria (e.g., the "event" criterion mentioned above).
[0043] LLMs are typically trained on large amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more. LLMs are characterized by their large scale, often including a large number of parameters to help them learn complex patterns in language data. These models are often based on deep learning architectures, such as transformers, which helps them provide better processing performance on various NLP tasks. In embodiments of this disclosure, the executing agent can utilize a generative large language model (e.g., an LLM) as a "story splitting model" to process and split the complete story into story units. This allows for faster and higher-quality splitting of the "complete story" into "story units."
[0044] In some embodiments, to further improve the speed of LLM calls, the "guide word" can be omitted through default configuration. For example, after obtaining the complete storyline, the LLM can naturally understand that the input is the "complete storyline" and the operations to be performed on the input "complete storyline" based on the default configuration. Thus, by using the default configuration, the LLM can stably and purposefully process the "complete storyline" to break it down into "storyline units," thereby improving the speed of LLM calls.
[0045] In some embodiments, as discussed above, the complete storyline may be obtained by the executing entity through the producer of the initial video or other users (e.g., the other users may be users who provide evaluations or comments on the initial video). In such cases, the user may also send an acquisition instruction for the complete storyline to the executing entity through the target device they are using (e.g., terminal devices 101, 102, 103) to acquire the complete storyline.
[0046] Accordingly, if the executing entity receives a request from the target device for the complete storyline, the executing entity can respond by providing the complete storyline to the target device.
[0047] Then, users can personalize the complete storyline according to their individual needs (e.g., adjusting the content, length, etc. of the complete storyline) and return updated information about the complete storyline to the executing entity.
[0048] Accordingly, if the executing entity receives updated information regarding the complete storyline, it can respond by updating the complete storyline based on the updated information. Then, the executing entity selects the updated complete storyline as input to the LLM to invoke the LLM to split the updated complete storyline into multiple storyline units.
[0049] Therefore, by interacting with users, the implementing entity can dynamically and personally meet the different users' understanding and usage needs of the plot, so as to provide users with "edited videos" that better meet their plot needs in the future.
[0050] When utilizing LLM, leveraging its processing capabilities, the executing entity can similarly delegate the step of "associating the scene video with the target plot unit with the highest semantic similarity in the semantic matching results" to LLM. For example, the executing entity can similarly use other content-based prompts (e.g., generating semantic matching results between subtitles included in the scene video and the textual descriptions of each plot unit in the initial video, and then associating the scene video with the target plot unit with the highest semantic similarity in the semantic matching results) to complete the step of "associating the scene video with the plot unit of the initial video" by calling LLM. This improves the accuracy of determining "semantic similarity."
[0051] In some embodiments, the LLM used to complete the step of “associating the scene video with the target narrative unit with the highest semantic similarity in the semantic matching result based on the semantic matching result” can be the same LLM used in the step of “splitting the complete narrative” (i.e., an LLM trained based on the same initial LLM that has both of these capabilities).
[0052] This allows the executing entity to perform the above two steps by calling a complete LLM, simplifying the calling logic. At the same time, it allows the LLM to process the results of different steps, thereby reducing the possibility of model illusion caused by training differences between different LLMs.
[0053] Step 203: Based on the subtitles included in the scene video associated with the plot unit, extract video frames from the initial video and generate the edited sub-video corresponding to the plot unit; In the embodiments of this disclosure, based on step 202, this step aims to have the execution entity, after determining the association between the plot unit and the scene video, extract video frames from the initial video using subtitles of scene videos belonging to the same plot unit and associated with the same plot unit. Accordingly, the execution entity can generate edited sub-videos corresponding to the plot unit based on these video frames to actually complete the subsequent "detailed correspondence" (i.e., using subtitles to more specifically "correspond" the plot unit to the video frame).
[0054] For example, the executing entity can arrange the scene videos in the order they appear in the initial video, and then determine a time interval based on the time when the subtitles first appear in the first scene video and the time when the subtitles last appear in the last scene video.
[0055] Then, the executing entity extracts the video frames of the initial video within that time interval and assembles and generates the edited sub-videos corresponding to the story unit.
[0056] In some optional implementations of this embodiment, the executing entity may, as an alternative or alternative, first determine the combination result of the subtitles of the scene videos associated with the plot unit (i.e., as discussed above, arrange the subtitles of each scene video in the order of their appearance to achieve "combination" and obtain the combination result of the subtitles).
[0057] Then, based on the combined result, the executing entity more specifically selects and determines the timestamps of each subtitle included in the combined result.
[0058] Then, based on the timestamps of each subtitle in the combined result, the executing entity extracts the corresponding video frames from the initial video, that is, only extracts the video frames containing "subtitles".
[0059] Finally, the executing entity combines the extracted video frames in chronological order to generate edited sub-videos corresponding to the story unit.
[0060] In this way, the implementing entity can further "remove" plot elements that cannot be obtained from the "subtitles," such as actions and atmosphere, so that the plot and subtitles can be matched and correspond more strictly, resulting in a more accurate and granular plot matching result.
[0061] Similarly, in some embodiments, the executing entity can also achieve the "fine correspondence" in this step by calling the LLM (e.g., changing the prompt word, training the LLM in this capability, or configuring other LLMs with this capability, etc.), which will not be repeated here.
[0062] Step 204: Generate a cut video of the initial video based on the edited sub-videos.
[0063] In the embodiments of this disclosure, based on step 203, the executing entity can use the various edited sub-videos generated in step 203 to generate an edited video of the initial video.
[0064] In this step, "Cut Video" can be generated using all the clip sub-videos or only some of the clip sub-videos.
[0065] In some optional implementations of this embodiment, the executing entity can combine the edited sub-videos corresponding to the selected story units based on the user's selection of story units to generate an edited video of the initial video.
[0066] For example, the executing entity can provide the split story units to the user (e.g., to the target device mentioned above, or to present the story units to the user using local presentation components). Accordingly, after receiving these story units, the user can select the desired story units based on their needs, as well as the order in which the story units are arranged (e.g., for certain story units that occur later, the user can instruct that the content of that story unit be presented first in the edited video based on different user needs), generating selection instructions for the story units.
[0067] The user can then return the selection instruction to the executing entity, instructing the entity to generate the edited video.
[0068] Accordingly, if the executing entity receives a selection instruction for a story unit, it can respond by combining the sub-video clips corresponding to the story unit indicated by the selection instruction in the order indicated by the selection instruction to generate the initial video clip (e.g., arranging these sub-video clips sequentially for combination).
[0069] Therefore, it is possible to provide users with personalized editing services for the initial video based on the "storyline" dimension, according to their different needs.
[0070] In some optional implementations of this embodiment, the user can also instruct the execution entity to generate a "scaled-down version" of the initial video by sending a global generation command for the initial video to the execution entity.
[0071] Accordingly, if the executing entity receives a global generation instruction for the initial video, it can respond by combining the corresponding edited sub-videos of each plot unit based on the order of the plot units indicated by the complete storyline of the initial video, thus generating an edited video of the initial video. This provides the user with an "edited video" that is abridged based on the plot content, enabling the user to conveniently and efficiently complete the editing of the extracted plot content from the initial video.
[0072] The video editing method provided in this disclosure involves segmenting an initial video into scenes to obtain scene videos, where each scene video corresponds to a scene in the initial video; associating the scene videos with plot units of the initial video; extracting video frames from the initial video based on subtitles included in the scene videos associated with the plot units to generate edited sub-videos corresponding to the plot units; and generating an edited video of the initial video based on the edited sub-videos. This improves the performance and efficiency of video editing and material preparation, while reducing the difficulty of video editing.
[0073] In some embodiments, during the process of associating scene videos and scene units, such as in step 202 above, the executing entity may also improve the quality of the scene videos used by removing those scene videos that are out of place for the plot unit and may cause errors.
[0074] In some embodiments, if at least two scene videos are associated with the same narrative unit (described as a target narrative unit for ease of understanding), the executing entity may respond by determining a time reference point based on the start times of the individual scene videos associated with the target narrative unit. For example, the time reference point may be the “average” of the individual start times (i.e., the sum of the individual start times divided by the number of individual scene videos associated with the target narrative unit).
[0075] Then, if the time distance between the start time and the time reference point of the target scene video associated with the target plot unit is greater than or equal to the first distance threshold (usually, the size of the "first distance threshold" can be preset based on the size of the time distance that is considered to be out of the group), the executing entity can consider the target scene video to be "out of the group" and then cancel the association between the target video and the target plot unit in order to avoid interference from the "out of the group" target scene video.
[0076] In some embodiments, when there are two or more scene videos associated with the same plot unit, the executing entity may also choose to divide the subtitles involved in achieving "detailed correspondence" according to the scene boundaries to avoid the situation where the subtitles (corresponding plot content) span too large a range due to scene differences.
[0077] Please refer to the following for details. Figure 3 , Figure 3 This is a flowchart illustrating a process for generating a sub-video clip corresponding to a story unit, as provided in an embodiment of this disclosure, including process 300. In some embodiments, process 300 may be an alternative or replacement for step 203 described above.
[0078] Process 300 specifically includes the following steps: Step 301: Combine the subtitles included in the scene videos associated with the plot unit to obtain the initial combination result; Specifically, in this step, the executing entity can, similar to what was discussed above, first combine the subtitles in the various scene videos that belong to and are related to the same plot unit to obtain an initial combination result (for example, arrange the subtitles according to the order of the scene videos in the initial video to achieve "combination").
[0079] Step 302: If the initial combination result includes a first subtitle and a second subtitle, determine whether the time distance between the start times of the first subtitle and the second subtitle is less than a second distance threshold; Specifically, if the executing entity detects and determines that the initial combination result includes a first subtitle and a second subtitle from different scene videos (for example, the first subtitle comes from the first scene video, while the second subtitle comes from the second scene video), the executing entity can further determine whether the time distance between the start time of the first subtitle and the second subtitle is less than a second distance threshold (in some embodiments, the size of the "second distance threshold" can be determined in advance based on the standard that the distance between the first subtitle and the second subtitle is close and that there will be no significant difference in content due to the span between the two).
[0080] If it is less than, the executing entity can continue to execute step 303 to further determine whether the first subtitle and / or the second subtitle is available, using the "plot unit" as a reference.
[0081] Step 303: Obtain the first time center of the first story unit and the second time center of the second story unit adjacent to the story unit; Specifically, the first time center of the first plot unit and the second time center of the second plot unit adjacent to the plot unit are obtained (for example, the first plot unit may be the previous plot unit adjacent to the plot unit, and the second plot unit may be the next plot unit adjacent to the plot unit).
[0082] Regarding the time center, taking the first time center as an example (the determination method of the second time center is similar to that of the first time center, and will not be repeated), the first time center can be the time point in the initial video based on the subtitles corresponding to the "key plot" of the first plot unit (that is, the time point at which the subtitles corresponding to the core plot and key plot in the plot unit are located).
[0083] Step 304: Sum the distance between the first subtitle and the first time center and the distance between the first subtitle and the second time center to obtain the first distance; sum the distance between the second subtitle and the first time center and the distance between the second subtitle and the second time center to obtain the second distance; Specifically, the executing entity can use the time centers of the first subtitle and the second subtitle to represent the first subtitle and the second subtitle, and then determine the distance between the "time center" representing the first subtitle and the first time center, and the distance between the second time center and the first time center, respectively. Then, these two distances are added together to obtain the first distance corresponding to the first subtitle and the second distance corresponding to the second subtitle.
[0084] Step 305: Compare the first distance and the second distance; If the first distance is less than the second distance, it can be considered that the first subtitle is closer to the plot center of the plot unit. In this case, the executing entity can choose to execute step 306, retain the first subtitle in the initial combination result, and delete the second subtitle.
[0085] Similarly, if the first distance is greater than the second distance, it can be considered that the second subtitle is closer to the plot center of the plot unit. In this case, the executing entity can choose to execute step 307, retain the second subtitle in the initial combination result, and delete the second subtitle.
[0086] Step 306: In the initial combined result, keep the first subtitle and delete the second subtitle to obtain the combined result; Step 307: Delete the first subtitle from the initial combined result to obtain the combined result; Step 308: Based on the subtitles included in the combined result, extract video frames from the initial video to generate the edited sub-videos corresponding to the story unit.
[0087] This step is similar to the process of generating edited sub-videos using the "combined results" discussed in step 203 above, and will not be repeated here.
[0088] Furthermore, in some embodiments, if the first distance and the second distance are equal, it can be considered that the two are likely to be equally related to the "plot center". In such cases, in order to avoid content loss, it is possible to retain the edited sub-videos corresponding to the plot units generated later by these two, which will not be repeated here.
[0089] In some optional implementations of this embodiment, as discussed above, if the distance between the first subtitle and the second subtitle is already large, and the difference in content may be obvious due to the span between the two, the executor may choose to delete at least one of them to avoid errors.
[0090] For example, the implementing entity can choose to have the first or second subtitle appear, or delete both (for example, if there are other subtitles that can be used).
[0091] For example, an embodiment involving the deletion of later-time subtitles will be discussed. In such a case, step 309 may also be included in process 300. For instance, in step 302 described above, if it is determined that the time distance between the start times of the first subtitle and the second subtitle is greater than or equal to a second distance threshold, the executing entity may choose to perform step 309.
[0092] Step 309: Compare the start times of the first and second subtitles; If the first subtitle starts later than the second subtitle, the executing entity can choose to continue with step 310.
[0093] If the first subtitle starts earlier than the second subtitle, the executing entity can choose to continue executing step 311.
[0094] Step 310: Delete the first subtitle from the initial combined result to obtain the combined result; Step 311: Delete the second subtitle in the initial combined result to obtain the combined result.
[0095] Subsequently, after obtaining the combined result, the executing entity can similarly jump to step 308 above to extract video frames from the initial video based on the subtitles included in the combined result, and generate the edited sub-videos corresponding to the plot unit. This will not be repeated here.
[0096] To enhance understanding, this disclosure also provides a specific implementation scheme based on a particular application scenario. Please refer to [link to relevant documentation]. Figure 4 . Figure 4This is a schematic diagram of an architecture 400 for implementing a video editing process in an application scenario, as provided in an embodiment of this disclosure. For example, architecture 400 can be implemented by the aforementioned server 105 as the execution entity.
[0097] In architecture 400, server 105, acting as the execution entity, can first execute S401 after obtaining the initial video 410 to perform scene segmentation on the initial video 410, thereby obtaining its corresponding scene videos 421, 422...42N (where N is a positive integer). Each of the scene videos 421, 422...42N can correspond to a "scene" in the initial video 410.
[0098] Then, server 105 can execute S402 to extract the "subtitles" included in each of the scene videos 421, 422...42N.
[0099] For ease of understanding, let's take scene video 421 as an example. Server 105 can extract subtitle 431 from scene video 421 by executing S402.
[0100] For ease of understanding, the following will only show the processing procedures related to scene video 421. The processing procedures for scene videos 422...42N can all be implemented in a similar way to the processing procedures related to scene video 421, and will not be repeated here.
[0101] Furthermore, server 105 can execute S403 based on subtitle 431 to associate scene video 421 with one of plot units 441, 442...44N (where N is a positive integer, and the "N" involved here may be the same as or different from the "N" in the scene video; this disclosure is not intended to limit this). Plot units 441, 442...44N can be obtained by splitting the complete plot (not shown in the figure) of initial video 410.
[0102] For example, server 105 can associate scene video 421 with plot unit 441 based on "semantic matching results".
[0103] Next, server 105 can continue to execute S404 to extract video frames from the initial video 410 based on subtitle 431 (for example, extracting "video frames" from the initial video 410 according to the "timestamp" corresponding to subtitle 431) and generate a sub-video 451 for the plot unit 441 (for example, arranging and combining these video frames in the order of "video frames" to obtain sub-video 451).
[0104] It should be understood that the video frame extraction here is only based on the "subtitles" of scene video 421 as an example. However, as discussed above, if two or more scene videos are simultaneously associated with plot unit 441, the server 105 can use the "subtitles" of two or more scene videos to extract video frames at the same time. This will not be repeated here.
[0105] Next, server 105 can similarly determine the corresponding sub-video clips (e.g., sub-video clips 451, 452, ..., 45N) for each of the story units 441, 442, ..., 44N through the above process.
[0106] After determining the sub-video clips 451, 452...45N, server 105 can generate a clipped video 460 of the initial video 410 based on the sub-video clips 451, 452...45N by executing S405. For example, upon receiving a "global generation instruction" (not shown in the figure), server 105 can choose to utilize all of the sub-video clips 451, 452...45N (for example, combining the sub-video clips 451, 452...45N in the order of story units 441, 442...44N to obtain clipped video 460).
[0107] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a video editing apparatus, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0108] like Figure 5 As shown, the video editing apparatus 500 of this embodiment may include: a scene video segmentation unit 501, a scene video association unit 502, an edited sub-video generation unit 503, and an edited video generation unit 504. The scene video segmentation unit 501 is configured to segment the initial video into scenes to obtain scene videos, wherein the scene videos correspond to scenes in the initial video; the scene video association unit 502 is configured to associate the scene videos with plot units of the initial video; the edited sub-video generation unit 503 is configured to extract video frames from the initial video based on subtitles included in the scene videos associated with the plot units, and generate edited sub-videos corresponding to the plot units; the edited video generation unit 504 is configured to generate an edited video of the initial video based on the edited sub-videos.
[0109] In this embodiment, the specific processing and technical effects of the scene video segmentation unit 501, scene video association unit 502, sub-video editing generation unit 503, and video editing generation unit 504 in the video editing device 500 can be referred to respectively. Figure 2The relevant descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.
[0110] In some optional implementations of this embodiment, the edited sub-video generation unit 503 includes: a video frame extraction sub-unit, configured to extract video frames from the initial video based on the timestamps of each subtitle in the combination result of the subtitles of the scene video associated with the plot unit; and a video frame combination sub-unit, configured to combine the video frames in the order of playback time to generate an edited sub-video corresponding to the plot unit.
[0111] In some optional implementations of this embodiment, the apparatus 500 further includes: a time reference point determination unit, configured to determine a time reference point based on the start time of each scene video associated with the target plot unit in response to the existence of at least two scene videos associated with the target plot unit; and an association update unit, configured to cancel the association between the target video and the target plot unit in response to the time distance between the start time of the target scene video associated with the target plot unit and the time reference point being greater than or equal to a first distance threshold.
[0112] In some optional implementations of this embodiment, the sub-video editing unit 503 includes: a subtitle combination subunit, configured to combine subtitles included in the scene video associated with the plot unit to obtain an initial combination result; a time center determination subunit, configured to, in response to the initial combination result including a first subtitle and a second subtitle, and the time distance between the start times of the first subtitle and the second subtitle being less than a second distance threshold, obtain a first time center of the first plot unit adjacent to the plot unit and a second time center of the second plot unit, wherein the first subtitle and the second subtitle come from different scene videos; and a time center summing subunit, configured to sum the first subtitle and the second subtitle in the first plot unit. The distance between the time centers is summed with the distance between the first subtitle and the second time center to obtain the first distance. The distance between the second subtitle and the first time center is summed with the distance between the second subtitle and the second time center to obtain the second distance. The first initial combination update subunit is configured to delete the second subtitle in the initial combination result in response to the first distance being less than the second distance, and obtain the combination result; or to delete the first subtitle in the initial combination result in response to the first distance being greater than the second distance, and obtain the combination result. The edited sub-video generation subunit is configured to extract video frames from the initial video based on the subtitles included in the combination result and generate the edited sub-video corresponding to the plot unit.
[0113] In some optional implementations of this embodiment, the sub-video editing unit 503 further includes: a subtitle start time comparison subunit, configured to compare the start times of the first subtitle and the second subtitle in response to the initial combination result including the first subtitle and the second subtitle, and the time distance between the start times of the first subtitle and the second subtitle being greater than or equal to a second distance threshold; and a second initial combination update subunit, configured to delete the first subtitle in the initial combination result in response to the start time of the first subtitle being later than that of the second subtitle, and obtain a combination result; or to delete the second subtitle in the initial combination result in response to the start time of the first subtitle being earlier than that of the second subtitle, and obtain a combination result.
[0114] In some optional implementations of this embodiment, the scene video association unit 502 is further configured to associate the scene video with the target plot unit with the highest semantic similarity in the semantic matching results based on the semantic matching results between the subtitles included in the scene video and the text description information of each plot unit in the initial video.
[0115] In some optional implementations of this embodiment, the scene video association unit 502 is further configured to generate semantic matching results between the subtitles included in the scene video and the text description information of each plot unit in the initial video by calling a large language model, and associate the scene video with the target plot unit with the highest semantic similarity in the semantic matching results.
[0116] In some optional implementations of this embodiment, the device 500 further includes: a complete plot acquisition unit, configured to acquire the complete plot of the initial video; and a plot unit splitting unit, configured to call a large language model to split the complete plot into multiple plot units.
[0117] In some optional implementations of this embodiment, the apparatus 500 further includes: a complete storyline providing unit, configured to provide the complete storyline to the target device in response to receiving an acquisition instruction sent by the target device for the complete storyline; a complete storyline updating unit, configured to update the complete storyline based on the update information received for the complete storyline; and a storyline unit splitting unit, further configured to call a large language model to split the updated complete storyline into multiple storyline units.
[0118] In some optional implementations of this embodiment, the video editing unit 504 is further configured to, in response to receiving a selection instruction for a story unit, combine the edited sub-videos corresponding to the story units indicated by the selection instruction in the order indicated by the selection instruction to generate an edited video of the initial video.
[0119] In some optional implementations of this embodiment, the video editing unit 504 is further configured to, in response to receiving a global generation instruction for the initial video, combine the edited sub-videos corresponding to each plot unit according to the order of the plot units indicated by the complete plot of the initial video, and generate an edited video of the initial video.
[0120] This embodiment exists as a device embodiment corresponding to the method embodiment described above. The video editing device provided in this embodiment performs scene segmentation on an initial video to obtain scene videos, wherein the scene videos correspond to scenes in the initial video; the scene videos are associated with plot units of the initial video; based on the subtitles included in the scene videos associated with the plot units, video frames are extracted from the initial video to generate edited sub-videos corresponding to the plot units; based on the edited sub-videos, an edited video of the initial video is generated. Therefore, it can improve the performance and efficiency of video editing and material editing, and reduce the difficulty of video editing.
[0121] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0122] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0123] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded into random access memory (RAM) 603 from storage unit 608. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0124] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0125] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the method of editing video. For example, in some embodiments, the method of editing video may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the method of editing video described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the method of editing video by any other suitable means (e.g., by means of firmware).
[0126] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0127] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0128] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0129] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0130] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0131] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are hosting products within the cloud computing service ecosystem to address the management difficulties and weak business scalability inherent in traditional physical hosts and Virtual Private Servers (VPS) services. Servers can also be categorized as distributed system servers or servers incorporating blockchain technology.
[0132] According to the technical solution of this disclosure, the initial video is segmented into scenes to obtain scene videos, wherein the scene videos correspond to scenes in the initial video; the scene videos are associated with the plot units of the initial video; based on the subtitles included in the scene videos associated with the plot units, video frames are extracted from the initial video to generate edited sub-videos corresponding to the plot units; based on the edited sub-videos, an edited video of the initial video is generated. This improves the performance and efficiency of video editing and material editing, while reducing the difficulty of video editing.
[0133] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.
[0134] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for editing video, comprising: The initial video is segmented into scenes to obtain scene videos, wherein the scene videos correspond to the scenes in the initial video; Associate the scene video with the story unit of the initial video; Based on the subtitles included in the scene videos associated with the plot unit, video frames are extracted from the initial video to generate a sub-video clip corresponding to the plot unit. This includes: combining the subtitles included in the scene videos associated with the plot unit to obtain an initial combination result, which includes: arranging the subtitles in each scene video that belongs to and is simultaneously associated with the same plot unit according to the chronological order of the scene videos in the initial video to obtain the initial combination result; in response to the initial combination result including a first subtitle and a second subtitle, and the time distance between the start times of the first subtitle and the second subtitle being less than a second distance threshold, obtaining the first time center of the first plot unit and the second time center of the second plot unit adjacent to the plot unit. The first subtitle and the second subtitle are from different scene videos. A first distance is obtained by summing the distance between the first subtitle and the first time center and the distance between the first subtitle and the second time center; a second distance is obtained by summing the distance between the second subtitle and the first time center and the distance between the second subtitle and the second time center. In response to the first distance being less than the second distance, the second subtitle is deleted from the initial combination result to obtain a combination result; or in response to the first distance being greater than the second distance, the first subtitle is deleted from the initial combination result to obtain a combination result. Based on the subtitles included in the combination result, video frames are extracted from the initial video to generate the edited sub-video corresponding to the plot unit. Based on the edited sub-videos, an edited video of the initial video is generated.
2. The method according to claim 1, wherein, Based on the subtitles included in the scene video associated with the plot unit, video frames are extracted from the initial video to generate a sub-video clip corresponding to the plot unit, including: Based on the timestamps of each subtitle in the combined result of the subtitles of the scene video associated with the plot unit, video frames are extracted from the initial video accordingly. The video frames are combined in chronological order to generate a sub-video clip corresponding to the story unit.
3. The method according to claim 1, further comprising: In response to the existence of at least two scene videos associated with the target story unit, a time reference point is determined based on the start time of each scene video associated with the target story unit; In response to the fact that the time distance between the start time of the target scene video associated with the target plot unit and the time reference point is greater than or equal to a first distance threshold, the association between the target scene video and the target plot unit is cancelled.
4. The method according to claim 1, further comprising: In response to the initial combination result including a first subtitle and a second subtitle, and the time distance between the start times of the first subtitle and the second subtitle being greater than or equal to the second distance threshold, the start times of the first subtitle and the second subtitle are compared; In response to the first subtitle starting later than the second subtitle, the first subtitle is deleted from the initial combined result to obtain the combined result; or Since the start time of the first subtitle is earlier than that of the second subtitle, the second subtitle is deleted from the initial combined result to obtain the combined result.
5. The method according to claim 1, wherein, The step of associating the scene video with the story unit of the initial video includes: Based on the semantic matching results between the subtitles included in the scene video and the text description information of each plot unit in the initial video, the scene video is associated with the target plot unit with the highest semantic similarity in the semantic matching results.
6. The method according to claim 5, wherein, The step of associating the scene video with the target plot unit with the highest semantic similarity in the semantic matching results based on the semantic matching results between the subtitles included in the scene video and the text description information of each plot unit in the initial video includes: By calling a large language model to generate semantic matching results between the subtitles included in the scene video and the textual description information of each plot unit in the initial video, the scene video is associated with the target plot unit with the highest semantic similarity in the semantic matching results.
7. The method according to claim 1, further comprising: Obtain the complete storyline of the initial video; The complete storyline is broken down into multiple storyline units by invoking a large language model.
8. The method according to claim 7, further comprising: In response to receiving an acquisition instruction from the target device for the complete storyline, the complete storyline is provided to the target device; In response to receiving update information regarding the complete storyline, the complete storyline is updated based on the update information; as well as The invocation of the large language model breaks down the complete storyline into multiple storyline units, including: The updated complete storyline is broken down into multiple storyline units by invoking a large language model.
9. The method according to any one of claims 1-8, wherein, The process of generating the edited video based on the edited sub-video includes: In response to receiving the selection instruction for the story unit, the edited sub-videos corresponding to the story units indicated by the selection instruction are combined in the order indicated by the selection instruction to generate the edited video of the initial video.
10. The method of claim 9, further comprising: In response to receiving a global generation instruction for the initial video, based on the order of the story units indicated by the complete storyline of the initial video, the corresponding edited sub-videos of each story unit are combined to generate an edited video of the initial video.
11. An apparatus for editing video, comprising: A scene video segmentation unit is configured to segment an initial video into scenes to obtain scene videos, wherein the scene videos correspond to scenes in the initial video; A scene video association unit is configured to associate the scene video with a story unit of the initial video; The sub-video editing unit is configured to extract video frames from the initial video based on subtitles included in the scene videos associated with the plot unit, and generate sub-videos corresponding to the plot unit. This includes: combining subtitles included in the scene videos associated with the plot unit to obtain an initial combination result; arranging subtitles in each scene video belonging to and associated with the same plot unit according to the order of the scene videos in the initial video to obtain the initial combination result; and, in response to the initial combination result including a first subtitle and a second subtitle, and the time distance between the start times of the first subtitle and the second subtitle being less than a second distance threshold, obtaining the first time center of the first plot unit adjacent to the plot unit and the time center of the second plot unit. A second time center is established, wherein the first subtitle and the second subtitle originate from different scene videos; a first distance is obtained by summing the distance between the first subtitle and the first time center and the distance between the first subtitle and the second time center, and a second distance is obtained by summing the distance between the second subtitle and the first time center and the distance between the second subtitle and the second time center; in response to the first distance being less than the second distance, the second subtitle is deleted from the initial combination result to obtain a combination result; or in response to the first distance being greater than the second distance, the first subtitle is deleted from the initial combination result to obtain a combination result; based on the subtitles included in the combination result, video frames are extracted from the initial video to generate the edited sub-video corresponding to the plot unit; The video editing unit is configured to generate an edited video of the initial video based on the edited sub-video.
12. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method of editing video according to any one of claims 1-10.
13. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method of editing video according to any one of claims 1-10.
14. A computer program product comprising a computer program that, when executed by a processor, implements the method of editing video according to any one of claims 1-10.