Video editing method and terminal device
By generating a mask image and identifying the image region of the object to be processed in the video editing interface, the problem of not being able to eliminate or redraw video footage in multi-track video editing is solved, enabling direct editing and display of video footage.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2025-10-16
- Publication Date
- 2026-06-04
AI Technical Summary
Existing technologies cannot directly eliminate or redraw video footage within a track in a multi-track video editing scenario.
A video editing method is provided, which generates a mask image in response to the user's editing operation in the video editing interface, identifies the image region of the object to be processed, edits it as the target object, replaces the video material segment, and displays the editing effect of the target video frame.
It enables the direct deletion or redrawing of video footage within a multi-track video editing interface, overcoming the shortcomings of existing technologies and improving the flexibility and efficiency of video editing.
Smart Images

Figure CN2025128201_04062026_PF_FP_ABST
Abstract
Description
A video editing method and terminal device
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411733064.2, filed on November 28, 2024, entitled "A Video Editing Method and Terminal Device", the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of video editing technology, and in particular to a video editing method and terminal device. Background Technology
[0004] A common function in image editing is that users can use a brush to paint over image content, and the image editing device automatically identifies the objects in the painted area and replaces them with corresponding background content or a specified object. The function that replaces the identified objects with corresponding background content is generally called the erase function, while the function that replaces the identified objects with a specified object is generally called the redraw function. Summary of the Invention
[0005] This application provides a video editing method and a terminal device. The technical solution provided by this application is as follows:
[0006] In a first aspect, embodiments of this application provide a video editing method, including:
[0007] In response to an editing operation on the initial video preview image in the video editing interface, the object to be processed indicated by the editing operation is determined. The video editing interface includes a video editing track and the initial video preview image. The video editing track is used to display the material clips corresponding to the video to be edited. The initial video preview image is used to display the editing effect of a specified video frame of the video to be edited. The object to be processed is at least one image content object presented in the specified video frame.
[0008] Obtain the image region corresponding to the object to be processed in each video frame of the video to be edited;
[0009] According to the image region corresponding to the object to be processed in each video frame of the video to be edited, the object to be processed in each video frame of the video to be edited is edited into the target object corresponding to each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited;
[0010] The material clips of the video to be edited are replaced with material clips of the target video on the video track, and a preview image of the target video frame is displayed on the video editing interface. The target video frame is used to demonstrate the editing effect of a specified video frame of the target video.
[0011] As an optional implementation of this application, the step of determining the object to be processed indicated by the editing operation in response to the editing operation on the initial video preview image in the video editing interface includes:
[0012] A mask image is generated based on the sliding trajectory of the editing operation;
[0013] The mask region of the target video frame is obtained based on the mask image, and the target video frame is the video frame displayed in the initial video preview image;
[0014] The image content within the mask area of the target video frame is identified to obtain the object to be processed.
[0015] As an optional implementation of this application, the step of editing the objects to be processed in each video frame of the video to be edited into target objects corresponding to each video frame of the video to be edited, according to the image regions corresponding to the objects to be processed in each video frame of the video to be edited, to obtain the target video corresponding to the video to be edited, includes:
[0016] The image content within the image area corresponding to the object to be processed in each video frame of the video to be edited is edited to the background content corresponding to the object to be processed, so as to obtain the target video corresponding to the video to be edited.
[0017] As an optional implementation of this application, before editing the objects to be processed in each video frame of the video to be edited into target objects corresponding to each video frame of the video to be edited according to the image regions corresponding to the objects to be processed in each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited, the method further includes:
[0018] Receive object description information input by the user;
[0019] The target objects corresponding to each video frame of the video to be edited are generated based on the object description information.
[0020] As an optional implementation of this application, the step of editing the objects to be processed in each video frame of the video to be edited into target objects corresponding to each video frame of the video to be edited, according to the image regions corresponding to the objects to be processed in each video frame of the video to be edited, to obtain the target video corresponding to the video to be edited, includes:
[0021] The video editing task is sent to the video editing server so that the video editing server edits the objects to be processed in each video frame of the video to be edited into the target objects corresponding to each video frame of the video to be edited, according to the image regions corresponding to the objects to be processed in each video frame of the video to be edited, so as to obtain the target video;
[0022] Receive the target video sent by the video editing server.
[0023] As an optional implementation of this application, after sending the video editing task to the video editing server, the method further includes:
[0024] Receive background operation input from the user for the video editing task;
[0025] In response to the background operation, the video editing task is switched to background operation, a task record corresponding to the video editing task is generated, the task execution information is polled, and the polled task information is updated in the task record.
[0026] As an optional implementation of this application, the method further includes:
[0027] The execution progress of the video editing task is obtained based on the task record;
[0028] The execution progress of the video editing task is displayed in the video editing interface.
[0029] As an optional implementation of this application, before generating a video editing task based on the target video frame, the mask image, and the video to be edited, the method further includes:
[0030] Determine whether the resolution of the video to be edited is greater than a threshold resolution, and determine whether the frame rate of the video to be edited is greater than a threshold frame rate;
[0031] If the resolution of the video to be edited is greater than the threshold resolution, then the video to be edited is downsampled so that the resolution of the video to be edited is less than or equal to the threshold resolution;
[0032] If the frame rate of the video to be edited is greater than the threshold frame rate, then the frame rate of the video to be edited is reduced to be less than or equal to the threshold frame rate.
[0033] As an optional implementation of this application, before editing the objects to be processed in each video frame of the video to be edited into target objects corresponding to each video frame of the video to be edited according to the image regions corresponding to the objects to be processed in each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited, the method further includes:
[0034] Determine if the length of the video to be edited is greater than the threshold length;
[0035] If the length of the video to be edited is greater than the threshold length, a prompt message is output, which prompts the user to cut out a video segment from the video to be edited that is less than or equal to the threshold length.
[0036] Receive cropping operations from the user on the video to be edited;
[0037] The video to be edited is cropped according to the cropping operation.
[0038] As an optional implementation of this application, when the length of the video to be edited is greater than the threshold length, after obtaining the target video corresponding to the video to be edited, the method further includes:
[0039] The target video and the video to be edited are spliced together, excluding the cropped video frequency bands.
[0040] Secondly, embodiments of this application provide a terminal device, including:
[0041] A processing unit is configured to, in response to an editing operation on an initial video preview image in a video editing interface, determine the object to be processed indicated by the editing operation. The video editing interface includes a video editing track and the initial video preview image. The video editing track is used to display material segments corresponding to the video to be edited. The initial video preview image is used to display the editing effect of a specified video frame of the video to be edited. The object to be processed is at least one image content object presented in the specified video frame.
[0042] The acquisition unit is used to acquire the image region corresponding to the object to be processed in each video frame of the video to be edited;
[0043] The editing unit is used to edit the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited, according to the image region corresponding to the object to be processed in each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited;
[0044] The display unit is used to replace the material segment of the video to be edited with the material segment of the target video on the video track, and to display the target video frame preview image on the video editing interface. The target video frame is used to display the editing effect of a specified video frame of the target video.
[0045] As an optional implementation of this application, the processing unit is specifically configured to generate a mask image based on the sliding trajectory of the editing operation; obtain a mask region of a target video frame based on the mask image, wherein the target video frame is the video frame displayed in the initial video preview image; and identify the image content within the mask region of the target video frame to obtain the object to be processed.
[0046] As an optional implementation of this application, the editing unit is specifically used to edit the image content within the image area corresponding to the object to be processed in each video frame of the video to be edited into the background content corresponding to the object to be processed, so as to obtain the target video corresponding to the video to be edited.
[0047] As an optional implementation of this application, the editing unit is further configured to, before editing the object to be processed in each video frame of the video to be edited into a target object corresponding to each video frame of the video to be edited according to the image region corresponding to the object to be processed in each video frame of the video to be edited, and obtaining the target video corresponding to the video to be edited, receive object description information input by the user; and generate the target object corresponding to each video frame of the video to be edited according to the object description information.
[0048] As an optional implementation of this application, the editing unit is specifically used to send the video editing task to the video editing server, so that the video editing server edits the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited according to the image region corresponding to the object to be processed in each video frame of the video to be edited, so as to obtain the target video; and receives the target video sent by the video editing server.
[0049] As an optional implementation of this application, the processing unit is further configured to, after sending the video editing task to the video editing server, receive a background operation input by the user for the video editing task; in response to the background operation, switch the video editing task to background operation, generate a task record corresponding to the video editing task, poll the task execution information, and update the task information obtained by polling to the task record.
[0050] As an optional implementation of this application, the display unit is further configured to obtain the execution progress of the video editing task based on the task record; and display the execution progress of the video editing task in the video editing interface.
[0051] As an optional implementation of this application, the editing unit is further configured to, before generating a video editing task based on the target video frame, the mask image, and the video to be edited, determine whether the resolution of the video to be edited is greater than a threshold resolution and whether the frame rate of the video to be edited is greater than a threshold frame rate; if the resolution of the video to be edited is greater than the threshold resolution, then the video to be edited is downsampled to make the resolution of the video to be edited less than or equal to the threshold resolution; if the frame rate of the video to be edited is greater than the threshold frame rate, then the video to be edited is subjected to a frame rate reduction operation to make the frame rate of the video to be edited less than or equal to the threshold frame rate.
[0052] As an optional implementation of this application, the editing unit is further configured to, before editing the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited according to the image region corresponding to the object to be processed in each video frame of the video to be edited, to obtain the target video corresponding to the video to be edited, determine whether the length of the video to be edited is greater than a threshold length; if the length of the video to be edited is greater than the threshold length, output a prompt message, the prompt message being used to prompt the user to cut out video segments with a length less than or equal to the threshold length from the video to be edited; receive the user's cropping operation input for the video to be edited; and crop the video to be edited according to the cropping operation.
[0053] As an optional implementation of this application, the editing unit is further configured to, after obtaining the target video corresponding to the video to be edited, splice the target video and the video to be edited, excluding the cropped video frequency bands, when the length of the video to be edited is greater than the threshold length.
[0054] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor, wherein the memory is used to store a computer program and the processor is used to cause the electronic device to implement the video editing method described in any of the above embodiments when executing the computer program.
[0055] Fourthly, embodiments of this application provide a computer-readable storage medium that, when executed by a computing device, enables the computing device to implement any of the video editing methods described above.
[0056] Fifthly, embodiments of this application provide a computer program product that, when run on a computer, enables the computer to implement any of the aforementioned video editing methods. Attached Figure Description
[0057] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0058] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings that need to be called in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0059] Figure 1 is a flowchart of the video editing method provided in an embodiment of this application;
[0060] Figure 2 is one of the interface diagrams of the video editing method provided in the embodiments of this application;
[0061] Figure 3 is one of the interactive flowcharts of the video editing method provided in the embodiments of this application;
[0062] Figure 4 is the second interactive flowchart of the video editing method provided in the embodiments of this application;
[0063] Figure 5 is the third interactive flowchart of the video editing method provided in the embodiments of this application;
[0064] Figure 6 is the fourth interactive flowchart of the video editing method provided in the embodiments of this application;
[0065] Figure 7 is a second schematic diagram of the interface of the video editing method provided in the embodiments of this application;
[0066] Figure 8 is a schematic diagram of the structure of the terminal device provided in an embodiment of this application;
[0067] Figure 9 is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of this application. Detailed Implementation
[0068] To better understand the above-mentioned objectives, features, and advantages of this application, the solution of this application will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.
[0069] Many specific details are set forth in the following description in order to provide a full understanding of this application, but this application may also be implemented in other ways different from those described herein. Obviously, the embodiments in the specification are only some embodiments of this application, and not all embodiments.
[0070] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner. Furthermore, in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.
[0071] The terminal device provided in this application embodiment can be a mobile phone, tablet computer, laptop computer, personal computer (PC), netbook, personal digital assistant (PDA), smartwatch, smart bracelet, or other terminal device, or the terminal device can be other types of terminal devices, which is not limited in this application embodiment.
[0072] With the booming development of the digital media industry, the application scope of erase and redraw functions is becoming increasingly widespread. However, the current erase and redraw functions only support the erasure and redrawing of images, not videos, and even less so the direct elimination or redrawing of mid-track video footage within a multi-track video editing scenario.
[0073] This application provides a video editing method and terminal device to solve the problem in related technologies that it is impossible to directly eliminate or redraw video materials in a multi-track video editing scenario. When the video editing method provided in this application receives an editing operation from a user on the initial video preview image in the video editing interface, it first responds to the editing operation by determining the object to be processed indicated by the editing operation. Then, it obtains the image region corresponding to the object to be processed in each video frame of the video to be edited. Next, according to the image region corresponding to the object to be processed in each video frame of the video to be edited, it edits the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited, thereby obtaining the target video corresponding to the video to be edited. It also replaces the material segments of the video to be edited with the material segments of the target video on the video track, and displays a preview image of the target video frame on the video editing interface. The target video frame is used to demonstrate the editing effect of a specified video frame of the target video. Since the video editing method provided in this application embodiment allows users to directly edit the initial video preview image in the multi-track video editing interface, the object to be processed in each video frame of the video to be edited can be edited into the target object corresponding to each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited. Therefore, this application embodiment can solve the problem in the related technology that it is impossible to directly eliminate or redraw the video material in the track in the multi-track video editing scenario.
[0074] This application provides a video editing method. Referring to FIG1, the video editing method includes the following steps:
[0075] S11. In response to the editing operation on the initial video preview image in the video editing interface, determine the object to be processed indicated by the editing operation.
[0076] The video editing interface includes a video editing track and an initial video preview image. The video editing track is used to display the material clips corresponding to the video to be edited, and the initial video preview image is used to display the editing effect of a specified video frame of the video to be edited. The object to be processed is at least one image content object presented in the specified video frame.
[0077] In some embodiments, the video editing interface is a multi-track video editing interface, and the video editing tracks of the video editing interface may include one or more of the following: editing tracks corresponding to video materials, editing tracks corresponding to audio materials, editing tracks corresponding to subtitle materials, and editing tracks corresponding to special effects materials, and the number of editing tracks for the same type of material can be any number.
[0078] Multitrack video editing is a video editing method that allows users to independently edit video footage, audio footage, special effects footage, subtitle footage, and other elements on multiple parallel tracks. In multitrack video editing, the footage on each track can be considered an independent layer, and users can freely add, modify, and delete footage on different tracks, greatly improving the flexibility of video editing. For example, a user can place the main video content on one track and add special effects, filters, or overlay graphics on another track. As another example, a user can arrange background music, narration, and sound effects on multiple audio tracks to achieve the desired audio effect.
[0079] In some embodiments, determining the object to be processed by the editing operation instruction includes the following steps a to c:
[0080] Step a: Generate a mask image based on the sliding trajectory of the editing operation.
[0081] In some embodiments, generating a mask image based on the sliding trajectory of the editing operation includes: initializing the handwriting path data of the editing operation when the business layer receives a press event; recording the handwriting path point data of the editing operation when the business layer receives a move event; ending the recording of the handwriting path point data of the editing operation when the business layer receives a lift event; generating the sliding trajectory of the editing operation based on the recorded handwriting path point data; and generating a mask image based on the sliding trajectory of the editing operation.
[0082] Referring to FIG2, in some embodiments, the video editing method provided in this application further includes: when receiving an editing operation input by a user on a target video frame in a multi-track video editing interface, displaying a sliding trajectory 20 corresponding to the editing operation in the video editing interface.
[0083] In some embodiments, displaying the sliding trajectory corresponding to the editing operation in the video editing interface includes: when the business layer receives a lift-off event, calling the drawing interface of the editor layer, the editor layer performs texture drawing based on the preset brush resources (including brush material and other information), stroke parameters (brush size, transparency and other information) and stroke path point data passed in by the drawing interface, and finally merges them into the rendering node of the target video frame to display the sliding trajectory corresponding to the editing operation in the video editing interface.
[0084] In some embodiments, generating a mask image based on the sliding trajectory of the editing operation includes: acquiring an initial image including the sliding trajectory of the editing operation; performing binarization processing on the initial image; acquiring a preprocessed image corresponding to the editing operation; and drawing the preprocessed image onto a background image with the same size as the target video frame based on the scaling information of the target video frame corresponding to the preview image, so as to generate a mask image.
[0085] In some embodiments, generating a mask image based on the sliding trajectory of the editing operation includes: acquiring the initial image output by the texture output interface provided by the editor side.
[0086] The initial image corresponding to the editing operation output by the texture output interface provided by the editor is an image containing only the sliding trajectory. This image is not adapted according to the size of the canvas and the size of the target video frame. However, when determining the object to be processed based on the target video frame and the mask image, the mask image must be a binary image with the same size as the target video frame (the mask area is pure white with 0 transparency, and the rest of the area is colorless). Therefore, after obtaining the initial image corresponding to the editing operation, the above implementation first performs binarization processing on the initial image to obtain the preprocessed image corresponding to the editing operation. Then, based on the scaling information of the target video frame, the preprocessed image is drawn on a background image with the same size as the original size of the target video frame, so that the mask image corresponding to the editing operation meets the requirements of subsequent processing.
[0087] Step b: Obtain the mask region of the target video frame based on the mask image.
[0088] The target video frame is the video frame displayed in the initial video preview image.
[0089] In some embodiments, obtaining the mask region of a target video frame based on the mask image includes: determining the orthographic projection of a first region of the mask image onto the target video frame as the mask region of the target video frame, wherein the first region of the mask image is a region in the mask image composed of pixels whose pixel values identify a sliding trajectory with an editing operation.
[0090] Step c: Identify the image content within the mask area of the target video frame to obtain the object to be processed.
[0091] For example, the image content within the mask area can be identified by an image recognition algorithm or image recognition model deployed on a terminal device to obtain the object to be processed.
[0092] S12. Obtain the image region corresponding to the object to be processed in each video frame of the video to be edited.
[0093] In some embodiments, an image recognition algorithm or image recognition model deployed on a terminal device can be used to identify the objects to be processed in each video frame of the video to be edited, and the area occupied by the objects to be processed in each video frame can be determined as the image region corresponding to the objects to be processed in each video frame.
[0094] In other embodiments, the object to be processed can be used as a tracking target, and the object to be processed in each video frame of the video to be edited can be obtained through a target tracking algorithm.
[0095] S13. According to the image region corresponding to the object to be processed in each video frame of the video to be edited, edit the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited.
[0096] In some embodiments, the step of editing the object to be processed in each video frame of the video to be edited into a target object corresponding to each video frame of the video to be edited, according to the image region corresponding to the object to be processed in each video frame of the video to be edited, to obtain a target video corresponding to the video to be edited, includes:
[0097] The image content within the image area corresponding to the object to be processed in each video frame of the video to be edited is edited to the background content corresponding to the object to be processed, so as to obtain the target video corresponding to the video to be edited.
[0098] That is, the video editing task is to erase the objects to be processed in each video frame of the video to be edited.
[0099] In some embodiments, the image content within the image region corresponding to the object to be processed in each video frame of the video to be edited is edited to the background content corresponding to the object to be processed, so as to obtain the target video corresponding to the video to be edited, including: obtaining the background content corresponding to the object to be processed in each video frame of the video to be edited according to the image inpainting algorithm, and replacing the object to be processed in each video frame of the video to be edited with the corresponding background content.
[0100] In some embodiments, before editing the objects to be processed in each video frame of the video to be edited into target objects corresponding to each video frame of the video to be edited according to the image regions corresponding to the objects to be processed in each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited, the method further includes: receiving object description information input by a user; and generating target objects corresponding to each video frame of the video to be edited based on the object description information.
[0101] That is, the video editing task is to redraw the objects to be processed in each video frame of the video to be edited.
[0102] In some embodiments, the terminal and device may execute steps S11 to S14 above using a video editing model or video editing algorithm to obtain the target video. The video editing model may be a network model trained on machine learning models such as convolutional neural networks, recurrent neural networks, and deep learning neural networks based on sample data.
[0103] S14. Replace the material segment of the video to be edited with the material segment of the target video on the video track, and display the target video frame preview image on the video editing interface. The target video frame is used to display the editing effect of a specified video frame of the target video.
[0104] The video editing method provided in this application, upon receiving an editing operation from a user on the initial video preview image in the video editing interface, first responds to the editing operation by determining the object to be processed indicated by the editing operation, then obtains the image region corresponding to the object to be processed in each video frame of the video to be edited, and then edits the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited according to the image region corresponding to the object to be processed in each video frame of the video to be edited, thereby obtaining the target video corresponding to the video to be edited, and replacing the material segment of the video to be edited with the material segment of the target video on the video track, and displays the target video frame preview image on the video editing interface, wherein the target video frame is used to display the editing effect of the specified video frame of the target video. Since the video editing method provided in this application embodiment allows users to directly edit the initial video preview image in the multi-track video editing interface, the object to be processed in each video frame of the video to be edited can be edited into the target object corresponding to each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited. Therefore, this application embodiment can solve the problem in the related technology that it is impossible to directly eliminate or redraw the video material in the track in the multi-track video editing scenario.
[0105] This application embodiment also provides another video editing method. Referring to FIG3, the video editing method includes the following steps:
[0106] S31. The terminal device receives the user's editing operation on the initial video preview image in the video editing interface.
[0107] The video editing interface includes a video editing track and an initial video preview image. The video editing track is used to display the material clips corresponding to the video to be edited, and the initial video preview image is used to display the editing effect of a specified video frame of the video to be edited. The object to be processed is at least one image content object presented in the specified video frame.
[0108] S32. The terminal device generates a mask image based on the sliding trajectory of the editing operation.
[0109] The method for generating the mask image can be referred to step a above, and will not be repeated here to avoid redundancy.
[0110] S33. The terminal device generates a video editing task based on the target video frame, the mask image, and the video to be edited.
[0111] In some embodiments, the terminal device generates a video editing task based on the target video frame, the mask image, and the video to be edited, including:
[0112] Based on the target video frame, the mask image, and the video to be edited, a video editing task is generated to instruct the image content within the image area corresponding to the object to be processed in each video frame of the video to be edited to be edited as the background content corresponding to the object to be processed.
[0113] That is, the video editing task is to erase the objects to be processed in each video frame of the video to be edited.
[0114] In some embodiments, the terminal device generates a video editing task based on the target video frame, the mask image, and the video to be edited, including:
[0115] Based on the target video frame, the mask image, the video to be edited, and object description information, a video editing task is generated to instruct the replacement of the object to be processed in each video frame of the video to be edited with an object generated based on the object description information.
[0116] That is, the video editing task is to redraw the objects to be processed in each video frame of the video to be edited.
[0117] S34. The terminal device sends the video editing task to the video editing server.
[0118] Correspondingly, the video editing server receives the video editing task sent by the terminal device.
[0119] This application embodiment does not limit the implementation method of the terminal device sending the video editing task to the video editing server. The terminal device can directly communicate with the video editing server to send the video editing task to the video editing server, or it can send the video editing task to the video editing server through other devices.
[0120] S35. The video editing server identifies the image content within the mask area of the target video frame to obtain the object to be processed.
[0121] S36. The video editing server obtains the image region corresponding to the object to be processed in each video frame of the video to be edited.
[0122] In some embodiments, the video editing server may use the object to be processed as a tracking target and obtain the object to be processed in each video frame of the video to be edited through a target tracking algorithm.
[0123] In some embodiments, the video editing server may use the object to be processed as a recognition target and obtain the object to be processed in each video frame of the video to be edited through an image recognition algorithm.
[0124] S37. The video editing server edits the objects to be processed in each video frame of the video to be edited into target objects corresponding to each video frame of the video to be edited, according to the image regions corresponding to the objects to be processed in each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited.
[0125] In some embodiments, the video editing server edits the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited, including: the video editing server obtaining the background content corresponding to the object to be processed in each video frame of the video to be edited according to an image filling algorithm, and replacing or overwriting the object to be processed in each video frame of the video to be edited with the corresponding background content.
[0126] In some embodiments, the video editing server edits the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited, including: the video editing server generating at least one object based on object description information, and replacing or overwriting the object to be processed in each video frame of the video to be edited with at least one object generated based on the object description information.
[0127] In some embodiments, the video editing server may perform the above steps S35 to S37 through a video editing model or video editing algorithm to obtain the target video.
[0128] S38. The video editing server sends the target video to the terminal device.
[0129] Accordingly, the terminal device receives the target video sent by the video editing server.
[0130] This application embodiment does not limit the implementation method of the video editing server sending the target video to the terminal device. The video editing server can directly communicate with the terminal device to send the target video to the terminal device, or it can send the target video to the terminal device through other devices or methods.
[0131] S39. The terminal device replaces the material segment of the video to be edited with the material segment of the target video on the video track, and displays a preview image of the target video frame on the video editing interface. The target video frame is used to display the editing effect of a specified video frame of the target video.
[0132] As an extension and refinement of the above embodiments, this application provides another video editing method. Referring to FIG4, the video editing method includes the following steps:
[0133] S401, The terminal device receives the user's editing operation on the initial video preview image in the video editing interface.
[0134] The video editing interface includes a video editing track and an initial video preview image. The video editing track is used to display the material clips corresponding to the video to be edited, and the initial video preview image is used to display the editing effect of a specified video frame of the video to be edited. The object to be processed is at least one image content object presented in the specified video frame.
[0135] S402. The terminal device generates a mask image based on the sliding trajectory of the editing operation.
[0136] S403. The terminal device determines whether the length of the video to be edited is greater than the threshold length.
[0137] In some embodiments, the threshold length can be dynamically sent to the terminal device through configuration information.
[0138] For example, the threshold length can be 10 seconds.
[0139] In step S403 above, if the terminal device determines that the length of the video to be edited is greater than the threshold length, then the following steps S404 to S406 are executed:
[0140] S404, The terminal device outputs a prompt message.
[0141] The prompt information is used to prompt the user to cut out video segments from the video to be edited that are shorter than or equal to the threshold length.
[0142] S405. The terminal device receives the user's cropping operation input for the video to be edited.
[0143] S406. The terminal device performs cropping on the video to be edited according to the cropping operation.
[0144] It should be noted that cropping the video to be edited according to the cropping operation can be to crop out a continuous video segment or to crop out multiple non-contiguous video segments. This application embodiment does not limit this, and the length of the video to be edited after cropping is less than or equal to the threshold length.
[0145] In step S403 above, if the terminal device determines that the length of the video to be edited is less than or equal to the threshold length, then the following step S407 is executed:
[0146] S407. The terminal device determines whether the resolution of the video to be edited is greater than the threshold resolution.
[0147] Similarly, the threshold resolution can also be dynamically sent to the terminal device through configuration information.
[0148] For example, the threshold resolution can be 2048×1080.
[0149] In step S407 above, if the terminal device determines that the resolution of the video to be edited is greater than the threshold resolution, then step S408 is executed; if the terminal device determines that the resolution of the video to be edited is less than or equal to the threshold resolution, then step S408 is skipped and step S409 is executed.
[0150] S408. The terminal device downsamples the video to be edited so that the resolution of the video to be edited is less than or equal to the threshold resolution.
[0151] S409. The terminal device determines whether the frame rate of the video to be edited is greater than the threshold frame rate.
[0152] Similarly, the threshold frame rate can also be dynamically sent to the terminal device through configuration information.
[0153] For example, the threshold resolution can be 30 frames per second (FPS).
[0154] In step S409 above, if the terminal device determines that the frame rate of the video to be edited is greater than the threshold frame rate, then step S410 is executed. If the terminal device determines that the frame rate of the video to be edited is less than or equal to the threshold resolution, then step S410 is skipped and step S411 is executed.
[0155] S410. The terminal device performs a frame rate reduction operation on the video to be edited, so that the frame rate of the video to be edited is less than or equal to the threshold frame rate.
[0156] Erasing objects from an image only requires processing a single image, while erasing objects from a video requires processing all video frames. Therefore, the processing time and performance overhead for erasing objects from a video are significantly greater than for image processing. The above embodiments limit the length, resolution, and frame rate of the video to be edited, thereby ensuring that the video editing process is not lengthy.
[0157] S411. The terminal device sends the video to be edited to the video storage server.
[0158] Correspondingly, the video storage server receives the video to be edited sent by the terminal device.
[0159] It should be noted that when the length of the video to be edited is less than or equal to the threshold length, the resolution of the video to be edited is less than or equal to the threshold resolution, and the frame rate of the video to be edited is less than or equal to the threshold frame rate, the encoded data of the video to be edited can be directly sent to the video storage server as the encoded data of the video to be edited. Otherwise, it is necessary to perform cropping and / or transcoding operations on the video to be edited to obtain the encoded data of the video to be edited.
[0160] S412, The video storage server sends the first access address to the terminal device.
[0161] Accordingly, the terminal device receives the first access address sent by the video storage server.
[0162] Wherein, the first access address is the access address of the video to be edited.
[0163] S413. The terminal device generates a video editing task based on the first access address, the mask image, and the position information of the target video frame in the video to be edited, which instructs the image content within the image area corresponding to the object to be processed in each video frame of the video to be edited to be edited as the background content corresponding to the object to be processed.
[0164] S414. The terminal device sends the video editing task to the application server.
[0165] Accordingly, the application server receives the video editing task sent by the terminal device.
[0166] S415. The application server sends the video editing task to the video editing server.
[0167] Accordingly, the video editing server receives the video editing task sent by the application server.
[0168] S416. The video editing server obtains the video to be edited from the video storage server according to the first access address.
[0169] S417. The video editing server determines the target video frame based on the position information of the target video frame in the video to be edited.
[0170] S418. The video editing server determines the object to be processed based on the target video frame and the mask image.
[0171] S419. The video editing server obtains the image region corresponding to the object to be processed in each video frame of the video to be edited.
[0172] S420. The video editing server edits the image content within the image area corresponding to the object to be processed in each video frame of the video to be edited into the background content corresponding to the object to be processed, so as to obtain the target video corresponding to the video to be edited.
[0173] S421. The video editing server sends the target video to the video storage server.
[0174] Accordingly, the video storage server receives the target video.
[0175] S422. The video storage server sends the second access address to the video editing server.
[0176] Accordingly, the video editing server receives the second access address sent by the video storage server.
[0177] The second access address is the access address of the video frequency band to be edited.
[0178] S423. The video editing server sends the second access address to the application server.
[0179] Accordingly, the application server receives the second access address sent by the video editing server.
[0180] S424. The application server sends the second access address to the terminal device.
[0181] Accordingly, the terminal device receives the second access address sent by the application server.
[0182] S425. The terminal device obtains the target video from the video storage server according to the second access address.
[0183] In step S403 above, if the terminal device determines that the length of the video to be edited is greater than the threshold length, then after executing step S425 above, the video editing method further includes:
[0184] S426. The terminal device splices the target video and the video to be edited, excluding the video frequency bands that have been cropped.
[0185] If the terminal device determines that the length of the video to be edited is greater than the threshold length, it will execute steps S404 to S406 to trim the video to be edited so that the length of the video to be edited is less than or equal to the threshold length. Since the duration of the trimmed video to be edited has changed compared to the original video, directly replacing the video to be edited with the target video would prevent the user from restoring the video to its original duration. For example, if the duration of the video to be edited is 1 minute, while the duration of both the trimmed video to be edited and the target video is 10 seconds, the user cannot restore the target video to 1 minute. Therefore, in the above embodiment, when the length of the video to be edited is determined to be greater than the threshold length, after obtaining the target video, it will also splice the target video and the video to be edited (excluding the trimmed video segments) to achieve the effect of processing the content within a specific segment of the video to be edited.
[0186] In some embodiments, after the terminal device sends the video editing task to the video editing server, the video editing method provided in this application further includes: receiving a background operation input by the user for the video editing task; in response to the background operation, switching the video editing task to background operation, generating a task record corresponding to the video editing task, polling the task execution information, and updating the task information obtained by polling to the task record.
[0187] For example, referring to FIG5, after the terminal device sends the video editing task to the video editing server, it can display a control 500 for triggering the background operation of the video editing task. The user can operate the control 500 for triggering the background operation of the video editing task to trigger the switching of the video editing task to run in the background.
[0188] In some embodiments, generating a task record corresponding to the video editing task includes: persistently recording the task parameters of the video editing task as a local JSON file.
[0189] The task parameters of the video editing task may include: the target video frame, the mask image, and the video to be edited.
[0190] The above embodiments can switch the video editing task to run in the background during the execution of the video editing task. Therefore, the above embodiments can perform other editing operations during the execution of the editing task, avoiding the video editing task blocking the video editing process.
[0191] In some embodiments, the video editing method further includes: a terminal device obtaining the execution progress of the video editing task based on the task record; and displaying the execution progress of the video editing task in the multi-track video editing interface.
[0192] Displaying the execution progress of the video editing task in the multi-track video editing interface allows users to easily obtain the execution progress of the video editing task.
[0193] As an extension and refinement of the above embodiments, this application provides another video editing method. Referring to FIG6, the video editing method includes the following steps:
[0194] S601, The terminal device receives the user's editing operation on the initial video preview image in the video editing interface.
[0195] The video editing interface includes a video editing track and an initial video preview image. The video editing track is used to display the material clips corresponding to the video to be edited, and the initial video preview image is used to display the editing effect of a specified video frame of the video to be edited. The object to be processed is at least one image content object presented in the specified video frame.
[0196] S602. The terminal device generates a mask image based on the sliding trajectory of the editing operation.
[0197] S603, Terminal device displays description information input interface.
[0198] S604. The terminal device obtains object description information based on the user's input in the description information input interface.
[0199] For example, referring to FIG7, when the terminal device receives the user's editing operation on the target video frame, the terminal device displays a description information input interface, which includes a text input box 70. When the user enters the text "a small boat" in the text input box 70, the object description information obtained by the terminal device is "a small boat".
[0200] S605. The terminal device determines whether the length of the video to be edited is greater than the threshold length.
[0201] The video to be edited is the video to which the target video frame belongs.
[0202] In step S605 above, if the terminal device determines that the length of the video to be edited is greater than the threshold length, then the following steps S606 to S608 are executed:
[0203] S606, The terminal device outputs a prompt message.
[0204] The prompt information is used to prompt the user to cut out a portion of the video to be edited that is less than or equal to the threshold length.
[0205] S607. The terminal device receives the user's cropping operation input for the video to be edited.
[0206] S608. The terminal device performs cropping on the video to be edited according to the cropping operation.
[0207] In step S605 above, if the terminal device determines that the length of the video to be edited is less than or equal to the threshold length, then the following step S609 is executed:
[0208] S609. The terminal device determines whether the resolution of the video to be edited is greater than the threshold resolution.
[0209] In step S609 above, if the terminal device determines that the resolution of the video to be edited is greater than the threshold resolution, then step S610 is executed; if the terminal device determines that the resolution of the video to be edited is less than or equal to the threshold resolution, then step S610 is skipped and step S611 is executed.
[0210] S610. The terminal device downsamples the video to be edited so that the resolution of the video to be edited is less than or equal to the threshold resolution.
[0211] S611. The terminal device determines whether the frame rate of the video to be edited is greater than the threshold frame rate.
[0212] In step S611 above, if the terminal device determines that the frame rate of the video to be edited is greater than the threshold frame rate, then step S612 is executed. If the terminal device determines that the frame rate of the video to be edited is less than or equal to the threshold resolution, then step S612 is skipped and step S613 is executed.
[0213] S612. The terminal device performs a frame rate reduction operation on the video to be edited, so that the frame rate of the video to be edited is less than or equal to the threshold frame rate.
[0214] Erasing objects from an image only requires processing a single image, while erasing objects from a video requires processing all video frames. Therefore, the processing time and performance overhead for erasing objects from a video are significantly greater than for image processing. The above embodiments limit the length, resolution, and frame rate of the video to be edited, thereby ensuring that the video editing process is not lengthy.
[0215] S613. The terminal device sends the video to be edited to the video storage server.
[0216] Correspondingly, the video storage server receives the video to be edited sent by the terminal device.
[0217] S614, The video storage server sends the first access address to the terminal device.
[0218] Accordingly, the terminal device receives the first access address sent by the video storage server.
[0219] Wherein, the first access address is the access address of the video to be edited.
[0220] S615. The terminal device generates a video editing task based on the target video frame, the mask image, the video to be edited, and the object description information, which instructs the processing object in each video frame of the video to be edited to be edited into an object generated according to the object description information.
[0221] S616. The terminal device sends the video editing task to the application server.
[0222] Accordingly, the application server receives the video editing task sent by the terminal device.
[0223] S617. The application server sends the video editing task to the video editing server.
[0224] Accordingly, the video editing server receives the video editing task sent by the application server.
[0225] S618. The video editing server generates the target object based on the object description information.
[0226] S619. The video editing server obtains the video to be edited from the video storage server according to the first access address.
[0227] S620. The video editing server determines the target video frame based on the position information of the target video frame in the video to be edited.
[0228] S621. The video editing server determines the object to be processed based on the target video frame and the mask image.
[0229] S622. The video editing server obtains the image region corresponding to the object to be processed in each video frame of the video to be edited.
[0230] S623. The video editing server edits the objects to be processed in each video frame of the video to be edited into objects generated according to the object description information, so as to obtain the target video corresponding to the video to be edited.
[0231] It should be noted that if the video editing server generates multiple objects based on the object description information in step S618, then in step S623, the object to be processed in each video frame of the video to be edited can be replaced with the multiple objects to obtain multiple target videos.
[0232] S624. The video editing server sends the target video to the video storage server.
[0233] Accordingly, the video storage server receives the target video.
[0234] S625. The video storage server sends the second access address to the video editing server.
[0235] Accordingly, the video editing server receives the second access address sent by the video storage server.
[0236] The second access address is the access address of the video band to be edited.
[0237] S626. The video editing server sends the second access address to the application server.
[0238] Accordingly, the application server receives the second access address sent by the video editing server.
[0239] S627. The application server sends the second access address to the terminal device.
[0240] Accordingly, the terminal device receives the second access address sent by the application server.
[0241] S628. The terminal device obtains the target video from the video storage server according to the second access address.
[0242] It should be noted that if the video editing server generates multiple objects based on the object description information in step S618, and the video editing server replaces the object to be processed in each video frame of the video to be edited with the multiple objects in step S623, then the target video obtained by the terminal device from the video storage server based on the second access address includes multiple objects. At this time, one of the target videos can be determined as the target video for subsequent processing based on the user's selection operation.
[0243] In step S605 above, if the terminal device determines that the length of the video to be edited is greater than the threshold length, then after executing step S628 above, the video editing method further includes:
[0244] S629. The terminal device splices together the target video and other video segments in the video to be edited, excluding the video to be edited.
[0245] Based on the same inventive concept, as an implementation of the above method, this application embodiment also provides a terminal device, which corresponds to the aforementioned method embodiment. For ease of reading, this embodiment will not repeat the details of the aforementioned method embodiment one by one, but it should be clear that the terminal device in this embodiment can implement all the contents of the aforementioned method embodiment.
[0246] This application provides a terminal device, and Figure 8 is a schematic diagram of the structure of the terminal device. As shown in Figure 8, the terminal device 800 includes:
[0247] The processing unit 81 is configured to, in response to an editing operation on an initial video preview image in a video editing interface, determine the object to be processed indicated by the editing operation. The video editing interface includes a video editing track and the initial video preview image. The video editing track is used to display material segments corresponding to the video to be edited. The initial video preview image is used to display the editing effect of a specified video frame of the video to be edited. The object to be processed is at least one image content object presented in the specified video frame.
[0248] The acquisition unit 82 is used to acquire the image region corresponding to the object to be processed in each video frame of the video to be edited;
[0249] Editing unit 83 is used to edit the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited according to the image region corresponding to the object to be processed in each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited;
[0250] Display unit 84 is used to replace the material segment of the video to be edited with the material segment of the target video on the video track, and to display a preview image of the target video frame on the video editing interface. The target video frame is used to display the editing effect of a specified video frame of the target video.
[0251] As an optional implementation of this application, the processing unit 81 is specifically configured to generate a mask image based on the sliding trajectory of the editing operation; obtain the mask region of the target video frame based on the mask image, wherein the target video frame is the video frame displayed in the initial video preview image; and identify the image content within the mask region of the target video frame to obtain the object to be processed.
[0252] As an optional implementation of this application, the editing unit 83 is specifically used to edit the image content within the image area corresponding to the object to be processed in each video frame of the video to be edited into the background content corresponding to the object to be processed, so as to obtain the target video corresponding to the video to be edited.
[0253] As an optional implementation of this application, the editing unit 83 is further configured to receive object description information input by the user before editing the object to be processed in each video frame of the video to be edited into a target object corresponding to each video frame of the video to be edited, according to the image region corresponding to the object to be processed in each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited; and generate the target object corresponding to each video frame of the video to be edited according to the object description information.
[0254] As an optional implementation of this application, the editing unit 83 is specifically used to send the video editing task to the video editing server, so that the video editing server edits the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited according to the image region corresponding to the object to be processed in each video frame of the video to be edited, so as to obtain the target video; and receives the target video sent by the video editing server.
[0255] As an optional implementation of this application, the processing unit 81 is further configured to, after sending the video editing task to the video editing server, receive a background operation input by the user for the video editing task; in response to the background operation, switch the video editing task to background operation, generate a task record corresponding to the video editing task, poll the task execution information, and update the task information obtained by polling to the task record.
[0256] As an optional implementation of this application, the display unit 84 is further configured to obtain the execution progress of the video editing task based on the task record; and display the execution progress of the video editing task in the video editing interface.
[0257] As an optional implementation of this application, the editing unit 83 is further configured to, before generating a video editing task based on the target video frame, the mask image, and the video to be edited, determine whether the resolution of the video to be edited is greater than a threshold resolution and whether the frame rate of the video to be edited is greater than a threshold frame rate; if the resolution of the video to be edited is greater than the threshold resolution, then the video to be edited is downsampled so that the resolution of the video to be edited is less than or equal to the threshold resolution; if the frame rate of the video to be edited is greater than the threshold frame rate, then the video to be edited is subjected to a frame rate reduction operation so that the frame rate of the video to be edited is less than or equal to the threshold frame rate.
[0258] As an optional implementation of this application, the editing unit 83 is further configured to, before editing the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited according to the image region corresponding to the object to be processed in each video frame of the video to be edited, to obtain the target video corresponding to the video to be edited, determine whether the length of the video to be edited is greater than a threshold length; if the length of the video to be edited is greater than the threshold length, output a prompt message, the prompt message being used to prompt the user to cut out video segments with a length less than or equal to the threshold length from the video to be edited; receive the user's cropping operation input for the video to be edited; and crop the video to be edited according to the cropping operation.
[0259] As an optional implementation of this application, the editing unit 83 is further configured to, after obtaining the target video corresponding to the video to be edited, splice the target video and the video to be edited, excluding the cropped video frequency bands, when the length of the video to be edited is greater than the threshold length.
[0260] The terminal device provided in this application embodiment can execute the video editing method provided in any of the above embodiments. Its implementation principle and technical effect are similar, and will not be described again here.
[0261] Based on the same inventive concept, this application also provides an electronic device. Figure 9 is a schematic diagram of the structure of the electronic device provided in this application embodiment. As shown in Figure 8, the electronic device provided in this embodiment includes: a memory 901 and a processor 902. The memory 901 is used to store a computer program, and the processor 902 is used to execute any of the video editing methods provided in the above embodiments when executing the computer program.
[0262] Based on the same inventive concept, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the computing device to implement any of the video editing methods provided in the above embodiments.
[0263] Based on the same inventive concept, this application also provides a computer program product that, when run on a computer, enables the computing device to implement any of the video editing methods provided in the above embodiments.
[0264] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0265] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0266] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0267] Computer-readable media include both permanent and non-permanent, removable and non-removable storage media. Storage media can store information using any method or technology; the information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0268] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A video editing method, comprising: In response to an editing operation on the initial video preview image in the video editing interface, the object to be processed indicated by the editing operation is determined. The video editing interface includes a video editing track and the initial video preview image. The video editing track is used to display the material clips corresponding to the video to be edited. The initial video preview image is used to display the editing effect of a specified video frame of the video to be edited. The object to be processed is at least one image content object presented in the specified video frame. Obtain the image region corresponding to the object to be processed in each video frame of the video to be edited; According to the image region corresponding to the object to be processed in each video frame of the video to be edited, the object to be processed in each video frame of the video to be edited is edited into the target object corresponding to each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited; The material clips of the video to be edited are replaced with material clips of the target video on the video track, and a preview image of the target video frame is displayed on the video editing interface. The target video frame is used to demonstrate the editing effect of a specified video frame of the target video.
2. The method according to claim 1, wherein, The step of determining the object to be processed indicated by the editing operation in response to an editing operation on an initial video preview image in the video editing interface includes: A mask image is generated based on the sliding trajectory of the editing operation; The mask region of the target video frame is obtained based on the mask image, and the target video frame is the video frame displayed in the initial video preview image; The image content within the mask area of the target video frame is identified to obtain the object to be processed.
3. The method according to claim 1, wherein, The step of editing the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited, according to the image region corresponding to the object to be processed in each video frame of the video to be edited, to obtain the target video corresponding to the video to be edited, includes: The image content within the image area corresponding to the object to be processed in each video frame of the video to be edited is edited to the background content corresponding to the object to be processed, so as to obtain the target video corresponding to the video to be edited.
4. The method according to claim 1, wherein, Before editing the objects to be processed in each video frame of the video to be edited into target objects corresponding to each video frame of the video to be edited, according to the image regions corresponding to the objects to be processed in each video frame of the video to be edited, to obtain the target video corresponding to the video to be edited, the method further includes: Receive object description information input by the user; The target objects corresponding to each video frame of the video to be edited are generated based on the object description information.
5. The method according to claim 1, wherein, The step of editing the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited, according to the image region corresponding to the object to be processed in each video frame of the video to be edited, to obtain the target video corresponding to the video to be edited, includes: The video editing task is sent to the video editing server so that the video editing server edits the objects to be processed in each video frame of the video to be edited into the target objects corresponding to each video frame of the video to be edited, according to the image regions corresponding to the objects to be processed in each video frame of the video to be edited, so as to obtain the target video; Receive the target video sent by the video editing server.
6. The method according to claim 5, wherein, After sending the video editing task to the video editing server, the method further includes: Receive background operation input from the user for the video editing task; In response to the background operation, the video editing task is switched to background operation, a task record corresponding to the video editing task is generated, the task execution information is polled, and the polled task information is updated in the task record.
7. The method according to claim 6, wherein, The method further includes: The execution progress of the video editing task is obtained based on the task record; The execution progress of the video editing task is displayed in the video editing interface.
8. The method according to claim 1, wherein, Before generating a video editing task based on the target video frame, the mask image, and the video to be edited, the method further includes: Determine whether the resolution of the video to be edited is greater than a threshold resolution, and determine whether the frame rate of the video to be edited is greater than a threshold frame rate; If the resolution of the video to be edited is greater than the threshold resolution, then the video to be edited is downsampled so that the resolution of the video to be edited is less than or equal to the threshold resolution; If the frame rate of the video to be edited is greater than the threshold frame rate, then the frame rate of the video to be edited is reduced to be less than or equal to the threshold frame rate.
9. The method according to claim 1, wherein, Before editing the objects to be processed in each video frame of the video to be edited into target objects corresponding to each video frame of the video to be edited, according to the image regions corresponding to the objects to be processed in each video frame of the video to be edited, to obtain the target video corresponding to the video to be edited, the method further includes: Determine if the length of the video to be edited is greater than the threshold length; If the length of the video to be edited is greater than the threshold length, a prompt message is output, which prompts the user to cut out a video segment from the video to be edited that is less than or equal to the threshold length. Receive cropping operations from the user on the video to be edited; The video to be edited is cropped according to the cropping operation.
10. The method according to claim 9, wherein, If the length of the video to be edited is greater than the threshold length, after obtaining the target video corresponding to the video to be edited, the method further includes: The target video and the video to be edited are spliced together, excluding the cropped video frequency bands.
11. A terminal device, comprising: A processing unit is configured to, in response to an editing operation on an initial video preview image in a video editing interface, determine the object to be processed indicated by the editing operation. The video editing interface includes a video editing track and the initial video preview image. The video editing track is used to display material segments corresponding to the video to be edited. The initial video preview image is used to display the editing effect of a specified video frame of the video to be edited. The object to be processed is at least one image content object presented in the specified video frame. The acquisition unit is used to acquire the image region corresponding to the object to be processed in each video frame of the video to be edited; The editing unit is used to edit the object to be processed in each video frame of the video to be edited into the target object corresponding to each video frame of the video to be edited, according to the image region corresponding to the object to be processed in each video frame of the video to be edited, so as to obtain the target video corresponding to the video to be edited; The display unit is used to replace the material segment of the video to be edited with the material segment of the target video on the video track, and to display the target video frame preview image on the video editing interface. The target video frame is used to display the editing effect of a specified video frame of the target video.
12. An electronic device, comprising: A memory and a processor, the memory being used to store a computer program and the processor being used to cause the electronic device to implement the video editing method according to any one of claims 1-10 when executing the computer program.
13. A computer-readable storage medium storing a computer program that, when executed by a computing device, causes the computing device to implement the video editing method according to any one of claims 1-10.
14. A computer program product, when run on a computer, causes the computer to implement the video editing method according to any one of claims 1-10.