Video editing method and device

By performing image group segmentation and parallel editing tasks on the video, the problem of low video editing efficiency is solved, and a more efficient video editing workflow is achieved.

CN122053884APending Publication Date: 2026-05-15BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-02-26
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing video editing technologies are inefficient, especially when dealing with stuttering or high latency, which prolongs the entire process and affects editing efficiency.

Method used

By segmenting the video to be edited into image groups, multiple independent video segments are obtained. The segments to be edited are then selected based on time information. Parallel editing tasks are created, and editing is performed using at least two task nodes to ultimately obtain the target video.

Benefits of technology

It improves the efficiency of video editing by reducing task execution time through parallel processing, thereby enhancing the overall efficiency of video editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053884A_ABST
    Figure CN122053884A_ABST
Patent Text Reader

Abstract

The invention provides a video editing method and device, and relates to the technical field of video processing. The method comprises the following steps: after acquiring a to-be-edited video and time information of to-be-edited video frames in the to-be-edited video, firstly, performing video cutting on the to-be-edited video according to image groups of the to-be-edited video to obtain a plurality of video clips, and then according to the time information of the to-be-edited video frames in the to-be-edited video and time intervals corresponding to the video clips, performing video editing on the to-be-edited video. Obtaining at least one to-be-edited video clip from the plurality of video clips, creating an editing task corresponding to the to-be-edited video clip, executing the editing task through the at least two task nodes to obtain a target video clip corresponding to the to-be-edited video clip, and finally, editing the target video clip. And obtaining a target video corresponding to the to-be-edited video according to the target video clip corresponding to the to-be-edited video clip. The method is used for solving the problem of low video editing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This article relates to the field of video processing technology, and in particular to a video editing method and apparatus. Background Technology

[0002] With the rapid development of the Internet, more and more data is maintained and disseminated through the Internet, such as the dissemination of videos or multimodal data containing video and audio. In this case, in order to improve the video viewers' awareness of the video-related content, relevant content can be displayed in the video through video editing. For example, the author's identity can be added to the video as a watermark, so that the video viewers can be aware of the video's author and the author's ownership of the video can be protected. Currently, video editing processes rely on linear operations to edit each frame of the video, with each subsequent frame depending on the result of the previous frame, resulting in low efficiency. If there are stutters or high latency during the operation, the execution time of the entire process will be further extended, affecting the efficiency of video editing. Summary of the Invention

[0003] In view of this, this paper provides a video editing method and apparatus to solve the problem of low video editing efficiency.

[0004] To achieve the above objectives, this paper provides the following technical solution: Firstly, this paper provides a video editing method, including: Obtain the time information of the video to be edited and the video frames to be edited in the video to be edited; cut the video to be edited according to the image group of the video to be edited to obtain multiple video segments; Based on the time information and the time interval corresponding to the video segment, at least one video segment to be edited is obtained from the plurality of video segments; Create an editing task corresponding to the video segment to be edited, the editing task being used to edit the video frames to be edited in the corresponding video segment; By executing the editing task corresponding to the video segment to be edited through at least two task nodes, the target video segment corresponding to the video segment to be edited is obtained. Obtain the target video corresponding to the video to be edited based on the target video segment corresponding to the video segment to be edited.

[0005] Secondly, this paper provides a video editing device, including: The video to be edited acquisition unit is used to acquire the video to be edited and the time information of the video frames to be edited in the video to be edited, and to cut the video to be edited according to the image group of the video to be edited to obtain multiple video segments; The video segment acquisition unit is used to acquire at least one video segment to be edited from the plurality of video segments based on the time information and the time interval corresponding to the video segment. The task creation unit is used to create an editing task corresponding to the video segment to be edited, and the editing task is used to edit the video frame to be edited in the corresponding video segment to be edited. The task execution unit is used to execute the editing task corresponding to the video segment to be edited through at least two task nodes to obtain the target video segment corresponding to the video segment to be edited. The target video acquisition unit is used to acquire the target video corresponding to the video to be edited based on the target video segment corresponding to the video segment to be edited.

[0006] Thirdly, this article provides an electronic device, including: a memory and a processor, wherein the memory is used to store a computer program and the processor is used to cause the electronic device to implement the video editing method described in any of the above embodiments when executing the computer program.

[0007] Fourthly, this document provides a computer-readable storage medium that, when executed by a computing device, causes the computing device to implement the video editing method described in any of the above embodiments.

[0008] Fifthly, this document provides a computer program product that, when run on a computer, enables the computer to implement the video editing method described in any of the above embodiments.

[0009] The video editing method presented in this paper first segments the video to be edited according to image groups, obtaining multiple video segments. This image group segmentation allows for independent processing of the resulting video segments. Next, based on the time information of the video frames to be edited and the time intervals corresponding to the video segments, at least one video segment to be edited is obtained from the multiple video segments. This eliminates the need to process all video segments. After obtaining at least one video segment to be edited, an editing task is created to edit the corresponding video frames within that segment. This task is executed through at least two task nodes to obtain the target video segment corresponding to the video frame to be edited. This improves task execution efficiency. Finally, based on the target video segment corresponding to the video frame to be edited, the target video corresponding to the video frame to be edited is obtained, thus improving both task execution efficiency and the overall video editing efficiency. Attached Figure Description

[0010] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this description and, together with the description, serve to explain the principles herein.

[0011] To more clearly illustrate the technical solutions in this document or the prior art, the accompanying drawings that need to be called in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0012] Figure 1 This article provides a flowchart of the video editing method steps. Figure 2 This is an illustration of the video clip provided in this article; Figure 3 This document provides a diagram illustrating the editing tasks required for this article. Figure 4 The flowchart for the editing task provided for this article; Figure 5 This is a schematic diagram of the video editing device provided in this article; Figure 6 This is a schematic diagram of the electronic device provided in this article. Detailed Implementation

[0013] To better understand the objectives, features, and advantages outlined above, the solutions described herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0014] Numerous specific details are set forth in the following description to provide a full understanding of this document, but other practices may also be employed in ways different from those described herein. Obviously, the embodiments in the specification are only a part of the embodiments described herein, and not all of them.

[0015] In this document, the words "exemplary" or "for example" are used to indicate that something is an example, illustration, or illustration. Any embodiment or design described herein as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner. Furthermore, in the description herein, unless otherwise stated, "a plurality of" means two or more.

[0016] Reference Figure 1 As shown, the video editing method provided in this article includes the following steps S11 to S15: S11. Obtain the time information of the video to be edited and the video frames to be edited in the video to be edited, and cut the video to be edited according to the image group of the video to be edited to obtain multiple video segments.

[0017] The aforementioned editing includes editing videos; for example, adding a watermark or markings to a video. The video to be edited includes the video that needs to be edited; for example, the video to be edited includes the video to be watermarked.

[0018] In practice, editing each video frame would consume a significant amount of time and resources. To achieve the desired editing effect while saving time and resources, some embodiments may edit only a portion of the video frames to be edited. During execution, the video frames to be edited can be determined based on time. Therefore, in addition to acquiring the video to be edited, the time information of the video frames to be edited within that video can also be obtained. This time information can be regular, with examples of 1s, 2s, ..., ns; alternatively, the time information can be irregular, with examples of 1s, 1.3s, 2s, 2.5s, ...

[0019] Furthermore, the aforementioned acquisition of the time information of the video to be edited and the video frames to be edited within the video to be edited can be replaced by acquiring the video to be edited, thereby enabling editing of each video frame within the video to be edited.

[0020] In practical applications, video is usually played in conjunction with audio to enhance the user's perception, that is, to enable the user to perceive through multimodal data including audio and video; however, video and audio are data of different modalities, and in order to improve processing efficiency and ensure processing quality, data of different modalities can be processed separately in specific processing.

[0021] During video editing, data from other modalities can be left unprocessed to improve processing efficiency. In some embodiments, when acquiring the video to be edited, multimodal data can be acquired first, and then modality separation processing can be performed on the multimodal data to obtain the video to be edited. The multimodal data may also include data from other modalities; for example, the data from other modalities may be audio. Therefore, in addition to obtaining the video to be edited, modality separation processing can also obtain reference modal data; for example, the reference modal data may be reference audio. After obtaining the reference modal data, the reference modal data can be stored so that after obtaining the target video, the target video and the reference modal data can be encapsulated to obtain the target multimodal data.

[0022] For example, multimodal data includes audio and video. After acquiring the multimodal data, modal separation processing is performed on the multimodal data to obtain the video to be edited and the reference audio, and the reference audio is stored.

[0023] Multimodal data includes audio and video. After acquiring the multimodal data, modality separation processing is performed to obtain the video to be edited and the reference audio, which is then stored. Furthermore, the reference modal data can also be text or other modal data; no limitation is made here.

[0024] In practice, after obtaining the video to be edited and the time information of the video frames within it, the video is segmented according to the image groups of the video to be edited, resulting in multiple video segments. Optionally, each video segment obtained by segmenting the video segments according to the image groups can be decoded independently. Thus, by segmenting the video to be edited into multiple independently decodeable video segments, multiple video segments can be processed independently in subsequent processing, without waiting for the previous video segment to be processed before processing the next, thereby improving video processing efficiency. The aforementioned image groups include GOPs (Group of Pictures).

[0025] In the specific execution process, multiple video segments can be video container formats that support independent playback, such as Fragment MP4 (Fragmented MP4) or MPEG-TS (MPEG2-TS, MPEG-2 transport stream) slices.

[0026] S12. Based on the time information and the time interval corresponding to the video segment, obtain at least one video segment to be edited from multiple video segments.

[0027] In practice, after obtaining the time information of the video frames to be edited from the video to be edited, and after segmenting the video into multiple video segments from the image groups to be edited, at least one video segment to be edited is selected from these multiple video segments based on the time information and the corresponding time intervals. This filters out video segments that do not require frame editing, eliminating the need for subsequent processing of all video segments. This reduces the number of video segments to be edited while still achieving effective editing, saving processing time and improving efficiency. Specifically, in the process of selecting at least one video segment to be edited from multiple video segments based on the time information and the corresponding time intervals, at least one video segment to be edited can be selected from multiple video segments based on the time information and the corresponding time intervals.

[0028] In the specific execution process, when obtaining at least one video segment to be edited from multiple video segments based on time information and the time intervals corresponding to each video segment, the video segment whose time interval contains the time information can be determined as the video segment to be edited. Alternatively, the video frame to be edited can be determined based on the time information, and the video segment corresponding to the video frame to be edited can be determined as the video segment to be edited.

[0029] For example, such as Figure 2 As shown, TS1 includes video frames 1 to 5; TS2 includes video frames 6 to 10; TS3 includes video frames 11 to 15; the video frames to be edited, determined based on the time information, are video frames 2, 4, 12, and 14. Among them, video frames 2 and 4 belong to TS1, and video frames 12 and 14 belong to TS3. Therefore, TS1 and TS3 are determined as the video segments to be edited.

[0030] In practice, after obtaining at least one video segment to be edited, video segments other than the video segment to be edited can be used as associated video segments corresponding to the video segment to be edited, so as to obtain the target video corresponding to the video segment to be edited in subsequent collaboration with the target video segment.

[0031] S13. Create an editing task corresponding to the video clip to be edited.

[0032] In practice, after acquiring at least one video segment to be edited, an editing task corresponding to that video segment is created; that is, an editing task is created for each video segment to be edited. The editing task can be used to edit the video frames to be edited within the corresponding video segment. Optionally, there is a one-to-one correspondence between an editing task and a video segment to be edited.

[0033] For example, such as Figure 3 As shown, create TASK1 corresponding to TS1 and TASK2 corresponding to TS2.

[0034] S14. Execute the editing task corresponding to the video segment to be edited through at least two task nodes to obtain the target video segment corresponding to the video segment to be edited.

[0035] In practice, after creating the editing task corresponding to the video segment to be edited, the editing task is executed through at least two task nodes to obtain the target video segment corresponding to the video segment to be edited. This parallel execution of the editing task through at least two task nodes improves task execution efficiency. The specific number of task nodes can be equal to the number of editing tasks, thus improving task execution efficiency through fully parallel execution. It should be noted that obtaining the target video segment corresponding to the video segment to be edited through at least two task nodes can mean executing the editing task corresponding to each video segment to be edited through at least two task nodes to obtain the target video segment corresponding to each video segment to be edited.

[0036] In the specific execution process, if it is a single server, the total available resources can be evenly distributed among the task nodes so that each task node can perform the editing task. In some embodiments, in the process of obtaining the target video segment corresponding to the video segment to be edited by executing the editing task corresponding to the video segment to be edited through at least two task nodes, the average available resources can first be calculated based on the total available resources and the number of video segments of at least one video segment to be edited. Then, the average available resources can be allocated to the task nodes corresponding to the number of video segments, so that the editing task can be executed by the task nodes corresponding to the number of video segments to obtain the target video segment. Specifically, in the process of allocating the average available resources to the task nodes corresponding to the number of video segments, the average available resources can be allocated to each task node among the task nodes corresponding to the number of video segments.

[0037] The aforementioned task nodes can be task threads or task processes.

[0038] For example, the number of editing tasks is 2, the total available resources of the server are m, task thread 1 and task thread 2 are created, and the available resources of task thread 1 and task thread 2 are m / 2 respectively. Task thread 1 is used to execute TASK1, and task thread 2 is used to execute TASK2.

[0039] Furthermore, in a distributed cluster, each editing task can be sent to a single server within the cluster, utilizing the total available resources of that single server for task execution. In other embodiments, during the process of obtaining the target video segment corresponding to the video segment to be edited by executing the editing task corresponding to the video segment to be edited through at least two task nodes, the available servers in the distributed cluster can be first obtained, and the target server corresponding to each editing task can be determined from among the available servers. Then, each editing task can be sent to the corresponding target server for task execution based on the target server.

[0040] Specifically, in determining the target server for each editing task from the available servers, the target server for each editing task can be randomly determined from the available servers. However, to improve task execution efficiency, an one-to-one correspondence can be established between editing tasks and target servers.

[0041] The above describes in detail two methods for obtaining the target video segment corresponding to the video segment to be edited by executing the editing task corresponding to the video segment to be edited through at least two task nodes.

[0042] In practice, since the video segment to be edited includes multiple video frames, during task execution, video frame decoding is required first. After decoding, the video frame to be edited is edited, and then the edited video frame is encoded. During this process, video frame editing can be performed according to a sequential execution strategy during the execution of any editing task. In some embodiments, during the execution of the editing task corresponding to the video segment to be edited, the current video frame in the video segment to be edited can be decoded first to obtain a decoded video frame. If the decoded video frame is the video frame to be edited, the decoded video frame is edited to obtain a target decoded video frame. Then, the target decoded video frame is encoded, and after the encoding is completed, the next video frame is decoded until the encoding of the last video frame in the video segment to be edited is completed.

[0043] In the specific execution process, the editing task corresponding to any video segment to be edited can decode the current video frame in that segment to obtain a decoded video frame. If the decoded video frame is the video frame to be edited, it is edited to obtain the target decoded video frame. The target decoded video frame is then encoded, and after encoding, the next video frame is decoded, and so on, until the encoding of the last video frame in any video segment is completed, thus obtaining the target video segment corresponding to that segment. If the decoded video frame is not the video frame to be edited, it can be directly encoded, and after encoding, the next video frame is decoded. In other words, the task is executed according to the order of the video frames in any video segment to be edited.

[0044] Furthermore, to further reduce task execution time and improve efficiency, an asynchronous execution strategy can be adopted for video frame editing tasks. Editing tasks can be divided into subtasks, and parallel processing of these subtasks can further improve execution efficiency. Optionally, editing tasks include: decoding subtasks, editing subtasks, and / or encoding subtasks; parallel execution of decoding, editing, and / or encoding subtasks is achieved by executing these subtasks through at least two subtask nodes.

[0045] The decoding subtask is used to decode video frames in the video segment to be edited, without waiting for the previous video frame to be decoded. The editing subtask is used to edit the video frames in the video segment to be edited. The encoding subtask is used to encode the target decoded video frame. Optionally, the target decoded video frame here includes the target decoded video frame obtained after editing the decoded video frame and / or the decoded video frame that does not need to be edited.

[0046] In the specific execution process, the decoding subtask can decode the current video frame in the video segment to be edited. After obtaining the decoded video frame, it checks whether the current video frame is the video frame to be edited. If it is, the decoded video frame is passed to the editing subtask; if not, the decoded video frame is passed to the encoding subtask. Then, the next video frame in the video segment to be edited is decoded, and the above process is repeated. The editing subtask can edit the decoded video frame passed in from the decoding subtask to obtain the target decoded video frame and pass it in to the encoding subtask; The encoding subtask can encode the decoded video frames passed in from the decoding subtask and the target decoded video frames passed in from the editing subtask.

[0047] Furthermore, the aforementioned decoding, editing, and encoding can be executed in parallel through two subtasks. For example, the editing task includes a decoding subtask and an encoding subtask. The decoding subtask can decode the current video frame in the video segment to be edited, obtain the decoded video frame, and then pass it to the encoding subtask. The encoding subtask can then decode the next video frame in the video segment to be edited and repeat the above process. The encoding subtask can first detect whether the decoded video frame is the video frame to be edited. If it is, it can edit the decoded video frame to obtain the target decoded video frame and encode the target decoded video frame. If not, it can encode the decoded video frame.

[0048] It should be noted that subtasks can be executed on a single server or in a distributed server cluster during parallel execution. The specific execution process is similar to the relevant content of the editing task mentioned above, and will not be repeated here.

[0049] In practice, during the execution of a specific task (video frame editing), duration detection can determine whether to use a sequential or asynchronous execution strategy. During execution, it can be checked whether the video frame processing duration meets the asynchronous execution conditions. If so, the asynchronous execution strategy is used; otherwise, a sequential execution strategy is used. The video frame processing duration includes the video frame editing duration and / or the video frame decoding duration. Correspondingly, the asynchronous execution conditions include the video frame editing duration being greater than the video frame decoding duration.

[0050] Specifically, the video frame editing duration and video frame decoding duration can be obtained. If the video frame editing duration is greater than the video frame decoding duration, an asynchronous execution strategy is used to execute the task. If the video frame editing duration is less than or equal to the decoding duration, a sequential execution strategy is used to execute the task.

[0051] In addition, the editing time percentage can be calculated based on the video frame editing duration and the baseline duration. If the editing time percentage is greater than the percentage threshold, an asynchronous execution strategy is used for task execution; otherwise, a sequential execution strategy is used. The baseline duration can include the time between the completion of decoding the current video frame and the completion of decoding the next video frame; the baseline duration is calculated based on the current video frame editing duration, the current video frame encoding duration, and the next video frame decoding duration.

[0052] The above provides two methods for duration detection. In addition, the two methods can be combined. In some embodiments, during the duration detection process, the video editing duration and video decoding duration are obtained. The editing duration ratio is calculated based on the video editing duration and the baseline duration. If the video frame editing duration is greater than the video frame decoding duration and the editing duration ratio is greater than the ratio threshold, then a decoding subtask, an editing subtask, and / or an encoding subtask are created.

[0053] It should be noted that during the encoding process of video frames, the encoding and decoding format (Codec, Coder-Decoder) of the video to be edited can be used to ensure that the target video segment is consistent with the video to be edited in terms of format compatibility and playback quality, thus avoiding anomalies or quality loss caused by changes in the encoding and decoding format.

[0054] S15. Obtain the target video corresponding to the video to be edited based on the target video segment corresponding to the video segment to be edited.

[0055] In practice, after obtaining the target video segment corresponding to the video segment to be edited, the target video corresponding to the video segment to be edited is obtained based on the target video segment corresponding to the video segment to be edited. That is, after obtaining the target video segment corresponding to each video segment to be edited, the target video corresponding to the video segment to be edited is obtained based on the target video segment corresponding to each video segment to be edited.

[0056] To ensure that the target video and the video to be edited are identical except for the frames to be edited, thus improving the effectiveness of the obtained target video, in some embodiments, during the process of obtaining the target video corresponding to the video to be edited based on the target video segment corresponding to the video segment to be edited, the associated video segment corresponding to the video segment to be edited is first obtained. Based on the time interval corresponding to the target video segment and the time interval corresponding to the associated video segment, the target video segment and the associated video segment are encapsulated to obtain the target video. Optionally, the associated video segment includes video segments other than the video to be edited from multiple video segments.

[0057] Encapsulation processing includes remux (Remultiplexing). Remux encapsulation allows independent media streams to be recombined into a new media file container without changing the encoded data of the video and / or audio streams themselves.

[0058] In the process of encapsulating the target video segment and related video segments according to the time intervals corresponding to the target video segment and the time intervals corresponding to the related video segments, the target video segment and related video segments can be sorted according to the time intervals to obtain the sorting order. The target video segment and related video segments are then encapsulated according to the sorting order to obtain the target video. In this way, the obtained target video and the video to be edited are consistent in time sequence.

[0059] If the target video is available, and the video to be edited also has corresponding baseline modal data, the baseline modal data corresponding to the video to be edited can be obtained. The target video and the baseline modal data are then encapsulated to obtain the target multimodal data.

[0060] For example, if the video to be edited is obtained from multimodal data containing audio and video, then the reference audio corresponding to the video to be edited is obtained, and the target video and the reference audio are encapsulated to obtain the target multimodal data.

[0061] The following describes the processing procedure starting from any editing task. The processing procedure for any editing task includes steps S41 and S48.

[0062] S41. Obtain the editing task for the video clip to be edited.

[0063] S42. Obtain the video frame processing duration of the video segment to be edited.

[0064] Optionally, the video frame processing duration can be obtained by decoding, editing, and / or encoding one or two frames of the video segment to be edited.

[0065] S43. Detect whether the asynchronous execution conditions are met based on the video frame processing time; If not, proceed with steps S44 and S47 to S48 as follows; If so, proceed with steps S45 to S48.

[0066] S44. Decode, edit, and / or encode the video frames according to the order of the video frames in the video segment to be edited to obtain the target video segment.

[0067] S45. Create decoding subtasks, editing subtasks, and encoding subtasks.

[0068] S46. The decoding subtask, editing subtask, and encoding subtask are executed in parallel to obtain the target video segment.

[0069] S47. Obtain the associated video segments of the video segment to be edited, encapsulate the target video segment and the associated video segments, and obtain the target video.

[0070] S48. Obtain the baseline modal data corresponding to the video to be edited, and encapsulate the target video and the baseline modal data to obtain the target multimodal data.

[0071] Based on the same inventive concept, as an implementation of the above method, this document also provides a video editing device. This embodiment corresponds to the aforementioned method embodiment. For ease of reading, this document will not repeat the details of the aforementioned method embodiment one by one, but it should be clear that the video editing device in this document can implement all the contents of the aforementioned method embodiment.

[0072] This article provides a video editing device. Figure 5 This is a schematic diagram of the video editing device, as shown below. Figure 5 As shown, the video editing device 500 includes: The video to be edited acquisition unit 51 is used to acquire the video to be edited and the time information of the video frames to be edited in the video to be edited, and to cut the video to be edited according to the image group of the video to be edited to obtain multiple video segments; The video segment acquisition unit 52 is used to acquire at least one video segment to be edited from the plurality of video segments based on the time information and the time interval corresponding to the video segment. The task creation unit 53 is used to create an editing task corresponding to the video segment to be edited, and the editing task is used to edit the video frame to be edited in the corresponding video segment to be edited. Task execution unit 54 is used to execute the editing task corresponding to the video segment to be edited through at least two task nodes to obtain the target video segment corresponding to the video segment to be edited; The target video acquisition unit 55 is used to acquire the target video corresponding to the video to be edited based on the target video segment corresponding to the video segment to be edited.

[0073] As an optional implementation, the task execution unit 54 is specifically configured to: calculate the average available resources based on the total available resources and the number of video segments to be edited; allocate the average available resources to the task nodes of the number of video segments, so as to execute the editing task through the task nodes of the number of video segments to obtain the target video segment.

[0074] As an optional implementation, the task execution unit 54, in the process of executing the editing task corresponding to the video segment to be edited, is specifically used to: decode the current video frame in the video segment to be edited to obtain a decoded video frame; if the decoded video frame is the video frame to be edited, edit the decoded video frame to obtain a target decoded video frame; encode the target decoded video frame, and decode the next video frame after the encoding process is completed.

[0075] As an optional implementation, the editing task includes: a decoding subtask, an editing subtask, and / or an encoding subtask; the decoding subtask, the editing subtask, and / or the encoding subtask are executed through at least two subtask nodes.

[0076] As an optional implementation, the device further includes a duration detection unit, which is specifically used for: acquiring the video frame editing duration and the video frame decoding duration; calculating the editing duration ratio based on the video frame editing duration and a reference duration; the reference duration includes the duration between the completion of decoding the current video frame and the completion of decoding the next video frame; if the video frame editing duration is greater than the video frame decoding duration, and the editing duration ratio is greater than a ratio threshold, then creating the decoding subtask, the editing subtask, and / or the encoding subtask.

[0077] As an optional implementation, the target video acquisition unit 55 is specifically used to: acquire associated video segments corresponding to the video segment to be edited; the associated video segments include video segments other than the video to be edited among the plurality of video segments; and encapsulate the target video segment and the associated video segments according to the time interval corresponding to the target video segment and the time interval corresponding to the associated video segments to obtain the target video.

[0078] As an optional implementation, the device further includes a modality separation unit, which is specifically used to: acquire multimodal data; the multimodal data includes audio and video; perform modality separation processing on the multimodal data to obtain the video to be edited and the reference audio, and store the reference audio.

[0079] As an optional implementation, the device further includes a modal encapsulation unit, which is specifically used for: acquiring the reference audio corresponding to the video to be edited; and encapsulating the target video and the reference audio to obtain target multimodal data.

[0080] The video editing device provided in this article can execute the video editing method provided in any of the above embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0081] Based on the same inventive concept, this article also provides an electronic device. Figure 6 The schematic diagram of the electronic device provided in this article is as follows: Figure 6 As shown, the electronic device provided herein includes a memory 601 and a processor 602, wherein the memory 601 is used to store a computer program and the processor 602 is used to execute the video editing method provided in the above embodiments when executing the computer program.

[0082] Based on the same inventive concept, this document also provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the computing device to implement the video editing method provided in the above embodiments.

[0083] Based on the same inventive concept, this document also provides a computer program product that, when run on a computer, enables the computing device to implement the video editing method provided in the above embodiments.

[0084] Those skilled in the art will understand that this document may be provided as a method, system, or computer program product. Therefore, this document may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this document may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0085] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0086] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0087] Computer-readable media include both permanent and non-permanent, removable and non-removable storage media. Storage media can store information using any method or technology; the information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this article, and are not intended to limit them. Although this article has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments in this article.

Claims

1. A video editing method, comprising: Obtain the time information of the video to be edited and the video frames to be edited in the video to be edited; cut the video to be edited according to the image group of the video to be edited to obtain multiple video segments; Based on the time information and the time interval corresponding to the video segment, at least one video segment to be edited is obtained from the plurality of video segments; Create an editing task corresponding to the video segment to be edited, the editing task being used to edit the video frames to be edited in the corresponding video segment; By executing the editing task corresponding to the video segment to be edited through at least two task nodes, the target video segment corresponding to the video segment to be edited is obtained. Obtain the target video corresponding to the video to be edited based on the target video segment corresponding to the video segment to be edited.

2. The method according to claim 1, wherein obtaining the target video segment corresponding to the video segment to be edited by executing the editing task corresponding to the video segment to be edited through at least two task nodes includes: Calculate the average available resources based on the total available resources and the number of video segments to be edited; The average available resources are allocated to the number of task nodes corresponding to the number of video segments, so that the editing task can be executed by the number of task nodes corresponding to the number of video segments to obtain the target video segment.

3. The method according to claim 1, wherein executing the editing task corresponding to the video segment to be edited includes: The current video frame in the video segment to be edited is decoded to obtain a decoded video frame; If the decoded video frame is a video frame to be edited, the decoded video frame is edited to obtain the target decoded video frame; The target video frame is encoded, and the next video frame is decoded after the encoding process is completed.

4. The method according to claim 1, wherein the editing task includes: Decoding subtasks, editing subtasks, and / or encoding subtasks; The decoding subtask, the editing subtask, and / or the encoding subtask are executed through at least two subtask nodes.

5. The method according to claim 4, further comprising: Get the video frame editing duration and video frame decoding duration; The percentage of editing time is calculated based on the video frame editing duration and the baseline duration. The reference duration includes the duration between the completion of decoding the current video frame and the completion of decoding the next video frame. If the video frame editing duration is greater than the video frame decoding duration, and the editing duration percentage is greater than the percentage threshold, then the decoding subtask, the editing subtask, and / or the encoding subtask are created.

6. The method according to claim 1, wherein obtaining the target video corresponding to the video to be edited based on the target video segment corresponding to the video segment to be edited comprises: Obtain the associated video segment corresponding to the video segment to be edited; The associated video segments include video segments other than the video to be edited among the plurality of video segments; Based on the time interval corresponding to the target video segment and the time interval corresponding to the associated video segment, the target video segment and the associated video segment are encapsulated to obtain the target video.

7. The method according to claim 1, further comprising: Acquire multimodal data; The multimodal data includes audio and video; The multimodal data is subjected to modal separation processing to obtain the video to be edited and the reference audio, and the reference audio is stored.

8. The method according to claim 7, further comprising: Obtain the reference audio corresponding to the video to be edited; The target video and the reference audio are encapsulated to obtain target multimodal data.

9. A video editing device, comprising: The video to be edited acquisition unit is used to acquire the video to be edited and the time information of the video frames to be edited in the video to be edited, and to cut the video to be edited according to the image group of the video to be edited to obtain multiple video segments; The video segment acquisition unit is used to acquire at least one video segment to be edited from the plurality of video segments based on the time information and the time interval corresponding to the video segment. The task creation unit is used to create an editing task corresponding to the video segment to be edited, and the editing task is used to edit the video frame to be edited in the corresponding video segment to be edited. The task execution unit is used to execute the editing task corresponding to the video segment to be edited through at least two task nodes to obtain the target video segment corresponding to the video segment to be edited. The target video acquisition unit is used to acquire the target video corresponding to the video to be edited based on the target video segment corresponding to the video segment to be edited.

10. An electronic device, comprising: A memory and a processor, the memory being used to store a computer program and the processor being used to cause the electronic device to implement the video editing method according to any one of claims 1-8 when executing the computer program.