Video editing method and apparatus, and terminal device
Automatically analyze and crop video frame images through terminal devices to generate videos with the lens moving with the target object, solving the problem of low manual editing efficiency and achieving efficient video editing.
Patent Information
- Application Number
- PCT/CN2024/139569
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2024-12-16
- Publication Date
- 2025-07-31
AI Technical Summary
In the prior art, the user generates videos that follow the characters through manual editing and generates a video that follows the camera and moves with low efficiency and high complexity.
The terminal device obtains the video frame image to be edited and the preset crop size, uses the plug-in to analyze the target position and perform cropping processing, generates the target video frame image, and finally generates a video where the lens moves with the target object.
Reduces the complexity of video cropping, improves the efficiency of video cropping, avoids jitter in the clipped video window, and improves the video display effect.
Smart Images

Figure CN2024139569_31072025_PF_FP_ABST
Abstract
Description
Video editing method, device and terminal equipment
[0001] This application claims priority to Chinese Patent Application No. 202410094711.3 filed on January 23, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The embodiments of the present disclosure relate to a video editing method, apparatus, and terminal device. Background Art
[0003] Users can edit videos with a fixed camera lens to create a video with the effect that the camera lens moves along with the characters in the video.
[0004] Currently, users can create videos that follow a person through manual editing. For example, users can identify the person image in each frame, crop that person image, and then stitch together multiple cropped images to create a video that gives the effect of the camera following the person. However, this manual process is complex, resulting in low video cropping efficiency. Summary of the Invention
[0005] The present disclosure provides a video editing method, apparatus, and terminal device.
[0006] In a first aspect, the present disclosure provides a video editing method, the video editing method comprising:
[0007] Obtaining the first video frame image and a preset cropping size in the video to be edited;
[0008] Analyzing the first video frame image according to the plug-in to determine the target position in the first video frame image, and cropping the first video frame image according to the cropping size and the target position to obtain a target video frame image;
[0009] A cropped video is generated according to the target video frame image.
[0010] In a second aspect, the present disclosure provides a video editing device, which includes a first acquisition module, a determination module, a cropping module, and a generation module, wherein:
[0011] The first acquisition module is used to acquire a first video frame image and a preset cropping size in the video to be edited;
[0012] The determination module is configured to analyze the first video frame image according to the plug-in to determine the target position in the first video frame image;
[0013] The cropping module is configured to crop the first video frame image according to the cropping size and the target position to obtain a target video frame image;
[0014] The generating module is used to generate a cropped video according to the target video frame image.
[0015] In a third aspect, an embodiment of the present disclosure provides a terminal device including: a processor and a memory;
[0016] The memory stores computer-executable instructions;
[0017] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video editing method as described in the first aspect and various possible aspects of the first aspect.
[0018] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video editing method as described in the first aspect and various possible aspects of the first aspect are implemented.
[0019] The present disclosure provides a video editing method, apparatus, and terminal device. The terminal device can obtain a first video frame image and a preset cropping size in a video to be edited, analyze the first video frame image according to a plug-in, and determine a target position in the first video frame image, wherein the target position is determined according to the position of a target object to be cropped in the first video frame image. The terminal device can crop the first video frame image according to the cropping size and the target position to obtain a target video frame image, and generate a cropped video based on the target video frame image. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, a brief introduction to the drawings required for use in the embodiments will be given below. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] FIG1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure;
[0022] FIG2 is a flow chart of a video editing method provided by an embodiment of the present disclosure;
[0023] FIG3 is a schematic diagram of a process for determining a target object according to an embodiment of the present disclosure;
[0024] FIG4 is a schematic diagram of a process for determining the position of a target object provided by an embodiment of the present disclosure;
[0025] FIG5A is a schematic diagram of a cropping area provided by an embodiment of the present disclosure;
[0026] FIG5B is a schematic diagram of another cropping area provided by an embodiment of the present disclosure;
[0027] FIG6 is a schematic diagram of a process for determining a target video frame image provided by an embodiment of the present disclosure;
[0028] FIG7 is a processing process of a plug-in provided by an embodiment of the present disclosure;
[0029] FIG8 is a schematic diagram of a method for cropping a second video frame image provided by an embodiment of the present disclosure;
[0030] FIG9 is a schematic diagram of a second video frame image provided by an embodiment of the present disclosure;
[0031] FIG10 is a schematic diagram of a video editing method according to an embodiment of the present disclosure;
[0032] FIG11 is a schematic structural diagram of a video editing device provided by an embodiment of the present disclosure;
[0033] FIG12 is a schematic structural diagram of another video editing device provided by an embodiment of the present disclosure;
[0034] FIG13 is a schematic structural diagram of a terminal device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0036] To facilitate understanding, the concepts involved in the embodiments of the present disclosure are explained below.
[0037] Terminal device: is a device with wireless transceiver function. The terminal device can be deployed on land, including indoors or outdoors, handheld, wearable or vehicle-mounted. The terminal device can be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a vehicle-mounted terminal device, a wireless terminal in self-driving, a wireless terminal device in remote medical, a wireless terminal device in smart grid, a wireless terminal device in transportation safety, a wireless terminal device in smart city, a wireless terminal device in smart home, a wearable terminal device, etc. The terminal device involved in the embodiments of the present disclosure can also be called a terminal, user equipment (UE), an access terminal device, a vehicle-mounted terminal, an industrial control terminal, a UE unit, a UE station, a mobile station, a mobile station, a remote station, a remote terminal device, a mobile device, a UE terminal device, a wireless communication device, a UE agent or a UE device, etc. Terminal devices can also be fixed or mobile.
[0038] In the related art, users can edit videos with a fixed camera lens to produce a video with the effect that the camera lens follows the movement of the characters in the video. For example, if the position of the camera lens is fixed, and the character in the video shot by the camera lens moves from the left side of the video to the right side of the video, the user can edit the video to produce a video with the effect that the camera lens moves from the left side to the right side following the user. Currently, users can edit videos based on manual editing methods. For example, users can manually determine the character image in each frame of the video, crop the character image, and splice multiple cropped images to produce a video with the effect that the camera lens follows the movement of the character. However, the complexity of manual processing is relatively high, resulting in low efficiency in video cropping.
[0039] In order to solve the technical problems in the related art, the embodiment of the present disclosure provides a video editing method, wherein a terminal device can obtain a first video frame image and a preset cropping size in a video to be edited, and determine the position of a target object in the first video frame image, and store the position of the target object in a cache, wherein the cache includes the position of the target object in multiple video frame images, and the position of the target object in the cache and the multiple positions are smoothed to obtain the target position of the first video frame image. The terminal device can crop the first video frame image according to the cropping size and the target position to obtain the target video frame image, and generate a cropped video based on the target video frame image. In this way, since the terminal device can smooth the position of the target object, the problem of window jitter of the cropped video can be avoided, and the video display effect can be improved. Moreover, since the terminal device can perform nonlinear cropping on the video frame image based on the plug-in, the terminal device can automatically generate the cropped video without manual intervention, thereby reducing the complexity of video cropping and improving the efficiency of video cropping.
[0040] The application scenario of the embodiment of the present disclosure is described below with reference to FIG1 .
[0041] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. Please refer to Figure 1, which includes: video A and a terminal device. Among them, video A may include 3 frames of images, in the first frame of the image, the ball is on the left side of the square, in the second frame of the image, the square blocks the ball, and in the third frame of the image, the ball is on the right side of the square, that is, in video A, the ball moves from the left side of the square to the right side of the square (the lens position for shooting video A is fixed). After the terminal device obtains video A, it can display video A. After the user clicks the edit control, the terminal device can generate video B, wherein video B also includes 3 frames of images, the first frame of the image includes the ball and part of the square, the second frame of the image includes the square and the blocked ball, and the third frame of the image includes part of the square and the ball. In this way, the video effect of video B is the effect of the lens following the movement of the ball, and there is no need to manually crop the video, which reduces the complexity of video cropping and improves the efficiency of video cropping.
[0042] It should be noted that FIG1 is only an example of an application scenario of the embodiment of the present disclosure, and does not limit the application scenario of the embodiment of the present disclosure.
[0043] The following detailed description of the technical solution of the present disclosure and how the technical solution of the present disclosure solves the above-mentioned technical problems is provided with specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present disclosure will be described below in conjunction with the accompanying drawings.
[0044] FIG2 is a flow chart of a video editing method provided by an embodiment of the present disclosure. Referring to FIG2 , the method may include:
[0045] S201: Obtain a first video frame image and a preset cropping size in a video to be edited.
[0046] The execution subject of the embodiment of the present disclosure may be a terminal device, or a video editing device provided in the terminal device. The video editing device may be implemented based on software, or based on a combination of software and hardware, which is not limited in the embodiment of the present disclosure.
[0047] The video to be edited may be a video with a fixed viewing angle. For example, the video to be edited may be shot with a fixed viewing angle. For example, the video to be edited may be a video shot with a fixed lens.
[0048] Optionally, the video to be edited may include an object. For example, the video to be edited may include any object such as an animal, a person, or an object. The video to be edited may also include any other type of object, which is not limited in the present embodiment.
[0049] It should be noted that the terminal device can obtain the video to be edited based on any feasible implementation method (for example, the terminal device can obtain the video to be edited from the database, or receive the video to be edited sent by other devices, or shoot the video to be edited in real time), and the embodiments of the present disclosure do not limit this.
[0050] The first video frame image may be an image in the video to be edited. For example, as shown in FIG1 , video A may be the video to be edited, and the first, second, and third frames of image in video A may be the first video frame image.
[0051] Optionally, the cropping size may be a pre-set image size. For example, the cropping size may be 3:4, or 16:9, etc., which is not limited in the present embodiment. For example, if the cropping size is 3:4, the terminal device may crop the first video frame image to a 3:4 size; if the cropping size is 16:9, the terminal device may crop the first video frame image to a 16:9 size.
[0052] It should be noted that the terminal device can determine the cropping size based on business needs (for example, when a video is published on the platform, the required size of the video is 3:4, then the corresponding preset cropping size of the platform can be 3:4). The terminal device can also determine the preset cropping size based on any feasible implementation method, and the embodiments of the present disclosure are not limited to this.
[0053] S202: Analyze the first video frame image according to the plug-in to determine the target position in the first video frame image.
[0054] Optionally, the user interface (UI) of the terminal device can send the video to be edited to the video editing side. The video editing side can read the video stream and extract frames of the video to be edited. The extracted video frames can be sent to the plug-in based on the link of the video editing side, so that the plug-in can obtain the video frame image and process the video frame image.
[0055] Among them, the target position is determined according to the position of the target object to be cropped in the first video frame image. The terminal device can determine the target position in the first video frame image according to the following feasible implementation method: determine the position of the target object in the first video frame image, smooth the position of the target object, and obtain the target position in the first video frame image.
[0056] Optionally, the target object may be an object to be cropped. For example, the target object may be a moving object in the video to be edited. For example, the video to be edited may include object 1 and object 2. If object 1 is moving and object 2 is stationary, the target object in the first video frame may be object 1. If object 1 is stationary and object 2 is moving, the target object in the first video frame may be object 2.
[0057] Optionally, the terminal device may determine the target object in the first video frame image based on the following feasible implementation manner: processing the first video frame image according to an object detection algorithm to obtain the target object in the first video frame image.
[0058] The target object may include a human portrait object. For example, the target object in the first video frame image may be a human face, a human limb, a human torso, etc., which is not limited in the present embodiment.
[0059] It should be noted that the above embodiments are only examples of target objects and are not limitations on the target objects. For example, the target object may also be a moving football, a flying bird, a swimming fish, etc., and the embodiments of the present disclosure do not limit this.
[0060] Optionally, an object detection algorithm can be used to detect the target object in the first video frame image. For example, the object detection algorithm can be an artificial intelligence (AI) algorithm. For example, if the target object is a face, the object detection algorithm can be a face detection algorithm. Based on the face detection algorithm, the first video frame image is processed, and the terminal device can determine whether the first video frame image includes a face image. For example, if the target object is a limb, the object detection algorithm can be a limb detection algorithm. Based on the limb detection algorithm, the first video frame image is processed, and the terminal device can determine whether the first video frame image includes a limb image.
[0061] It should be noted that the terminal device can obtain the object detection algorithm based on any feasible implementation method, and the object detection algorithm can also be implemented based on a model, which is not limited in the embodiments of the present disclosure.
[0062] It should be noted that the terminal device can determine the target object in the first video frame image based on any feasible implementation method (such as based on a trained model implementation), and the embodiments of the present disclosure are not limited to this.
[0063] Next, the determination of the target object in the first video frame image will be described with reference to FIG. 3 .
[0064] FIG3 is a schematic diagram of a process for determining a target object provided by an embodiment of the present disclosure. Please refer to FIG3 , which includes a first video frame image and a face detection module. The first video frame image may include a portrait, and the target object may be a face in the portrait. The terminal device (not shown in FIG3 ) may process the first video frame image based on the face detection module, and further determine that the first video frame image may include a face, and the face detection module may also determine the position or area of the face in the first video frame image, thereby improving the recognition accuracy of the target object and thus improving the accuracy of video editing.
[0065] Among them, the terminal device can determine the position of the target object in the first video frame image according to the following feasible implementation method: determine the target area of the target object in the first video frame image, and determine the position of the target object in the first video frame image based on the target area.
[0066] The target area may be the area occupied by the target object in the first video frame. For example, the target area may be the area enclosed by the outline of the target object in the first video frame, or the target area may be the area enclosed by multiple vertices (e.g., the upper vertex, the lower vertex, the left vertex, and the right vertex) of the target object in the first video frame. This disclosure is not limited to this aspect.
[0067] Optionally, when the terminal device processes the first video frame image based on the object detection algorithm, the target area of the target object in the first video frame image can be obtained. The terminal device can also determine the target area of the target object in the first video frame image based on any feasible implementation method. The embodiments of the present disclosure are not limited to this.
[0068] Optionally, the terminal device determines the position of the target object in the first video frame image based on the target area, and specifically may determine the center position of the target area as the position of the target object.
[0069] The center position may be the geometric center of the target area or the midpoint of the central axis of the target area, which is not limited in the embodiments of the present disclosure.
[0070] The terminal device may determine the coordinates of the center position of the target area, and then determine the coordinates of the center position as the position of the target object. For example, if the center of the target area is the 10th pixel on the long side (i.e., the x-axis of the pixel coordinate system) and the 20th pixel on the high side (i.e., the y-axis of the pixel coordinate system) of the first video frame image, the position of the target object may be (10, 20); if the center of the target area is the 100th pixel on the long side and the 70th pixel on the high side of the first video frame image, the position of the target object may be (100, 70).
[0071] It should be noted that the terminal device can determine the center position of the target area as the position of the target object in the first video frame image, and the terminal device can also determine any point in the target area as the position of the target object in the first video frame image. The embodiment of the present disclosure does not limit this.
[0072] 4 , the process of determining the position of the target object in the first video frame image by the terminal device is described below.
[0073] FIG4 is a schematic diagram of a process for determining the position of a target object provided by an embodiment of the present disclosure. Please refer to FIG4 , which includes: a first video frame image. The first video frame image may include a human portrait. The terminal device (not shown in FIG4 ) can process the first video frame image based on a face detection algorithm to determine the human face image in the first video frame image and the target area corresponding to the human face image. The terminal device can determine the center of the target area as the position of the human face in the first video frame image. In this way, the terminal device can accurately determine the position of the target object in the first video frame image, thereby improving the accuracy of video editing.
[0074] Among them, the terminal device can smooth the position of the target object to obtain the target position in the first video frame image, which can avoid the window jitter of the generated cropped video and thus improve the display effect of the video.
[0075] Among them, the terminal device smoothes the position of the target object to obtain the target position in the first video frame image, which can be specifically: storing the position of the target object in the cache, smoothing the position of the target object in the cache, and multiple positions to obtain the target position of the first video frame image.
[0076] The cache may include multiple positions of the target object in multiple video frame images. For example, the video to be edited may include multiple first video frame images. Therefore, the terminal device can determine the position of the target object in each first video frame image and store multiple positions in the cache. The terminal device can smooth the multiple positions to obtain the target position corresponding to each first video frame image. For example, the terminal device can store position 1 of the target object in the first video frame A and position 2 of the target object in the first video frame B in the cache. The terminal device can smooth position 1 and position 2 to obtain target position a corresponding to position 1 and target position b corresponding to position 2, wherein target position a can be the target position of the first video frame A, and target position b can be the target position of the first video frame B. In this way, since the target position is a smoothed position, the window jitter of the cropped video generated by the terminal device is smaller, thereby improving the video playback effect.
[0077] It should be noted that the terminal device can smooth multiple positions in the cache to obtain the target position, or it can smooth the two positions when the position difference of the target object in two adjacent first video frame images is greater than a preset threshold. The embodiments of the present disclosure do not limit this.
[0078] It should be noted that the terminal device can perform smoothing processing on the position and multiple positions of the target object in the cache based on any feasible implementation method, and the embodiments of the present disclosure are not limited to this.
[0079] S203 : Crop the first video frame image according to the cropping size and the target position to obtain a target video frame image.
[0080] The target video frame image may be a video frame image that is cropped from the first video frame image, and the target video frame image may include the target object. For example, based on the target position and the cropping size, the terminal device may crop the target video frame image that includes the target object from the first video frame image, and the size of the target video frame image is the same as the preset cropping size. In this way, the cropping accuracy of the target video frame image and the efficiency of video editing can be improved.
[0081] The terminal device can obtain the target video frame image based on the following feasible implementation method: obtaining the target position of the first video frame image in the cache, using the target position as the cropping center, and cropping the first video frame image according to the cropping size to obtain the target video frame image. Because the preset cropping size matches the business requirements (e.g., requiring a 16:9 video), not only the efficiency of video cropping can be improved, but also the accuracy of video cropping can be improved.
[0082] Optionally, the terminal device may determine a cropping region in the first video frame image based on the target position and the cropping size, and crop the cropping region to obtain the target video frame image. For example, the cropping region may include the target object, and the image in the cropping region may be the target video frame image.
[0083] The size of the cropping area is the same as the cropping size. For example, if the preset cropping size is 16:9, the size of the cropping area can also be 16:9; if the preset cropping size is 800*600, the size of the cropping area can also be 800*600.
[0084] The cropping area is described below with reference to FIG. 5A and FIG. 5B .
[0085] Figure 5A is a schematic diagram of a cropping area provided by an embodiment of the present disclosure. Please refer to Figure 5A, which includes a first video frame image. The first video frame image may include a portrait. If the target object is a face in the portrait, the terminal device (not shown in Figure 5A) may determine the center of the face as the target position of the target object in the first video frame image, that is, point A is the target position of the first video frame image. If the preset cropping size is 9:16 (aspect ratio, other cropping sizes in the embodiment of the present disclosure may also be aspect ratios, and the embodiment of the present disclosure will not be repeated), the terminal device determines a 16:9 area with point A as the center, and determines the area as the cropping area. The cropping area may include a face image, and the center of the cropping area may be the center of the face.
[0086] FIG5B is a schematic diagram of another cropping area provided by an embodiment of the present disclosure. Please refer to FIG5B , which includes a first video frame image. The first video frame image may include a portrait. If the target object is a face in the portrait, the terminal device (not shown in FIG5B ) may determine the center of the face as the target position of the target object in the first video frame image, that is, point A is the target position of the first video frame image. If the preset cropping size is 600*800 (number of pixels), the terminal device may determine an 800*600 area in the first video frame image based on point A, and determine the area as the cropping area. The cropping area includes a facial image, and the center of the facial image may be located on the central axis of the cropping area (if it is located at the center of the cropping area, it is not possible to obtain enough pixels, and the center of the facial image may also be located at any position in the cropping area, which is not limited in the embodiment of the present disclosure). In this way, the accuracy of cropping can be improved and the effect of video display can be improved.
[0087] Optionally, the electronic device may determine the upper left vertex, the upper right vertex, the lower left vertex, and the lower right vertex in the first video frame image based on the target position and the cropping size, and determine the area enclosed by the upper left vertex, the upper right vertex, the lower left vertex, and the lower right vertex as the cropping area. For example, the coordinates of the upper left vertex, the upper right vertex, the lower left vertex, and the lower right vertex may be the pixel coordinates of the four vertices of the cropping area in the first video frame image. For example, if the cropping area is a rectangle, the terminal device may determine the pixel coordinates of the upper left vertex of the cropping area in the first video frame image as the upper left vertex coordinates, the terminal device may determine the pixel coordinates of the upper right vertex of the cropping area in the first video frame image as the upper right vertex coordinates, the terminal device may determine the pixel coordinates of the lower left vertex of the cropping area in the first video frame image as the lower left vertex coordinates, and the terminal device may determine the pixel coordinates of the lower right vertex of the cropping area in the first video frame image as the lower right vertex coordinates.
[0088] Optionally, after the terminal device determines the upper left vertex, the upper right vertex, the lower left vertex, and the lower right vertex, it can determine the area enclosed by the upper left vertex, the upper right vertex, the lower left vertex, and the lower right vertex as the clipping area. For example, if the coordinates of the upper left vertex are (1, 3), the coordinates of the upper right vertex are (3, 3), the coordinates of the lower left vertex are (1, 0), and the coordinates of the lower right vertex are (3, 0), then the terminal device can determine the area enclosed by (1, 3), (3, 3), (1, 0), and (3, 0) as the clipping area.
[0089] It should be noted that the terminal device can determine the cropping area in the first video frame image based on any other feasible implementation method, and the embodiment of the present disclosure is not limited to this.
[0090] The following describes the process of determining the target video frame image by the terminal device in conjunction with FIG6 .
[0091] FIG6 is a schematic diagram of a process for determining a target video frame image according to an embodiment of the present disclosure. Referring to FIG6 , the process includes: a first video frame image. The first video frame image may include a portrait. If the target object is a face within the portrait, the terminal device (not shown in FIG6 ) may determine the center of the face as the target position of the target object within the first video frame image. That is, point A is the target position of the first video frame image.
[0092] Please refer to Figure 6. If the preset cropping size is 800*800 (number of pixels), the terminal device takes point A as the center and, based on the cropping size of 800*800, determines the coordinates B of the upper left vertex, the coordinates C of the lower left vertex, the coordinates D of the upper right vertex, and the coordinates E of the lower right vertex in the first video frame image. The terminal device can determine the area surrounded by B, C, D, and E as the cropping area, and crop the cropping area to obtain the target video frame image. The target video frame image may include a face image, and the center of the face image may be located at the center of the target video frame image. In this way, the terminal device can automatically crop the first video frame image to obtain a target video frame image including the target object, thereby improving the efficiency of image cropping and the efficiency of video editing.
[0093] The processing of the above plug-in will be described below with reference to FIG7 .
[0094] Figure 7 shows a processing process of a plug-in provided by an embodiment of the present disclosure. Please refer to Figure 7, which includes: a plug-in. Among them, the plug-in can receive a video frame image sent by the video editing side, and the plug-in can process the video frame image to generate cropping data, and store the cropping data in the cache. The cropping data can be the position of the target object in the video frame image, and the terminal device can smooth the cropping data in the cache. When the plug-in uses the cropping data, the video frame image can be cropped according to the preset cropping size and cropping data to obtain the target video frame image. In this way, the terminal device can realize nonlinear intelligent cropping based on the plug-in, improve the cropping efficiency of the video frame image, and then improve the generation efficiency of the cropped video.
[0095] S204: Generate a cropped video according to the target video frame image.
[0096] Optionally, the terminal device may synthesize the target video frame image corresponding to each first video frame image to obtain a cropped video. For example, after the terminal device obtains the target video frame image based on the plug-in, it may send the target video frame image to the video synthesis side. After the video synthesis side obtains the target video frame image corresponding to each first video frame image (the video synthesis side may also process the target video frame image frame by frame, which is not limited in this embodiment of the present disclosure), a cropped video may be generated. The cropped video includes the target object, and the display effect of the cropped video is that the lens moves along with the movement of the target object.
[0097] In other words, after the terminal device obtains the target video frame image based on the plug-in, it can send the target video frame image to the upper screen module. The upper screen module can display the target video frame image in the preview page. The user can preview the cropping result of the terminal device. If an error occurs, the user can correct it in time. In addition, the user can also adjust the cropping result, which can improve the accuracy of video cropping and enhance the user experience.
[0098] The disclosed embodiment provides a video editing method, wherein a terminal device can obtain a first video frame image and a preset cropping size in a video to be edited, determine the position of a target object in the first video frame image, and store the position of the target object in a cache, wherein the cache includes the positions of the target object in multiple video frame images, and smooth the position of the target object and the multiple positions in the cache to obtain the target position of the first video frame image. The terminal device can obtain the target position of the first video frame image in the cache, and use the target position as the cropping center and crop the first video frame image according to the cropping size to obtain the target video frame image. The terminal device can generate a cropped video based on the target video frame image. In this way, since the terminal device can smooth the position of the target object to obtain the target position, the problem of window jitter in the cropped video can be avoided, thereby improving the video display effect. Moreover, since the terminal device can perform nonlinear cropping on the video frame image based on a plug-in, the terminal device can automatically generate the cropped video without manual intervention, thereby reducing the complexity of video cropping and improving the efficiency of video cropping.
[0099] Based on the embodiment shown in Figure 2, in the above-mentioned video editing method, if the first video frame image is an image obtained based on interval frame extraction (the terminal device extracts the first video frame image in the video to be edited based on interval frame extraction), then the target positions of the multiple video frame images that have not been extracted in the video to be edited do not exist in the cache. Therefore, after the terminal device smoothes the position of the target object and obtains the target position in the first video frame image, the above-mentioned video editing method also includes a cropping method for the second video frame image that has not been extracted. Below, in combination with Figure 8, the cropping method for the second video frame image is described in detail.
[0100] FIG8 is a schematic diagram of a method for cropping a second video frame image provided by an embodiment of the present disclosure. Referring to FIG8 , the method process may include:
[0101] S801: Acquire a second video frame image between a first video frame image and a next extracted video frame image.
[0102] The second video frame image is a video frame image between the first video frame image and the next video frame image extracted. For example, the terminal device's frame extraction method for the video to be edited includes frame-by-frame extraction and interval frame extraction. If the terminal device's frame extraction method for the video to be edited is frame-by-frame extraction, then each video frame image in the video to be edited is the first video frame image. Based on the above video editing method, the terminal device can determine the target video frame image corresponding to each video frame image, and then obtain the cropped video. If the frame extraction method for the video to be edited is interval frame extraction, then the video frame image extracted by the terminal device in the video to be edited is the first video frame image, and the video frame image not extracted by the terminal device can be the second video frame image.
[0103] For example, the terminal device uses interval frame extraction as the frame extraction method for the edited video. If the terminal device determines video frame image 1 as the current first video frame image and determines video frame image 2 as the next frame extracted video frame image, then the terminal device can determine the video frame image between video frame image 1 and video frame image 2 as the second video frame image corresponding to video frame image 1 and video frame image 2.
[0104] It should be noted that the terminal device may also determine the second video frame image that the terminal device has not extracted based on any other feasible implementation method, and the embodiment of the present disclosure is not limited to this.
[0105] Optionally, the terminal device may determine the frame extraction method for the video to be edited based on the optical flow information between the video frame images. For example, if the optical flow difference between adjacent video frame images in the video to be edited is less than a preset optical flow value, the terminal device may determine that the frame extraction method for the video to be edited is interval frame extraction. If the optical flow difference between adjacent video frame images in the video to be edited is greater than or equal to the preset optical flow value, the terminal device may determine that the frame extraction method for the video to be edited is frame-by-frame frame extraction. In this way, the accuracy of the frame extraction of the video to be edited can be improved.
[0106] It should be noted that the terminal device can also determine the frame extraction method of the video to be edited based on any other feasible implementation method (for example, based on a trained model implementation), and the embodiments of the present disclosure are not limited to this.
[0107] The second video frame image is described below with reference to FIG9 .
[0108] FIG9 is a schematic diagram of a second video frame image provided by an embodiment of the present disclosure. Please refer to FIG9 , which includes: a video to be edited. The video to be edited includes image 1 (first frame), image 2 (second frame), image 3 (third frame), image 4 (fourth frame), image 5 (fifth frame), ..., image n (nth frame). If the terminal device (not shown in FIG9 ) uses an interval frame extraction method for the video to be edited, the first video frame of the interval frame extraction of the terminal device may include image 1, image 5, ..., image n.
[0109] Referring to Figure 9, since Image 5 is the first video frame following Image 1, the terminal device can determine that the second video frame corresponding to Image 1 and Image 5 may include Image 2, Image 3, and Image 4. In this way, the terminal device can quickly and accurately determine the second video frame, improving the accuracy and efficiency of video editing.
[0110] S802: Obtain a smooth curve corresponding to the target position of the first video frame image and the target position of the next extracted video frame image.
[0111] Optionally, after the terminal device performs smoothing processing on the target position of the first video frame image and the target position of the next extracted video frame image, a smooth curve corresponding to the smoothing processing can be obtained.
[0112] It should be noted that the terminal device can determine the smooth curve according to any feasible implementation method, and the embodiments of the present disclosure are not limited to this.
[0113] S803: Determine the target position of the second video frame image according to the smooth curve.
[0114] Wherein, after the terminal device determines the smooth curve, it can determine the target position of the second video frame image according to the numerical value of the horizontal coordinate and the numerical value of the vertical coordinate in the smooth curve. For example, the smooth curve is a smooth curve corresponding to target position 1 and target position 2, and the terminal device can determine the horizontal coordinate and vertical coordinate corresponding to the point on the smooth curve as the target position of the second video frame image. For example, the coordinates of the target position 1 of the first video frame image A are (1, 1), the coordinates of the target position 2 of the first video frame image B are (3, 3), and there is a second video frame image between the first video frame image A and the first video frame image B (only an example, not a limitation), then the target position of the second video frame image can be the coordinates (2, 2), which can avoid the problem of window jitter of the cropped video and improve the effect of video playback.
[0115] S804 : Crop the second video frame image according to the target position and cropping size of the second video frame image to obtain a target video frame image corresponding to the second video frame image.
[0116] After the terminal device determines the target position of the second video frame image, it can crop the second video frame image based on the cropping method for the first video frame image to obtain the target video frame image corresponding to the second video frame image. For example, after the terminal device determines the target position of the second video frame image, the terminal device crops the second video frame image based on the cropping size with the target position of the second video frame image as the center, thereby obtaining the target video frame image corresponding to the second video frame image. In this way, the terminal device can quickly crop each video frame image in the video to be edited, thereby improving the efficiency of video editing.
[0117] The disclosed embodiment provides a method for cropping a second video frame image. A terminal device can obtain a second video frame image between a first video frame image and a next extracted video frame image, obtain a smooth curve corresponding to the target position of the first video frame image and the target position of the next extracted video frame image, determine the target position of the second video frame image based on the smooth curve, and crop the second video frame image based on the target position and cropping size of the second video frame image to obtain a target video frame image corresponding to the second video frame image. In this way, the terminal device can quickly determine the target position of each second video frame image based on the smooth curve, thereby improving the efficiency of video cropping.
[0118] Based on any of the above embodiments, the process of the above video editing method will be described below in conjunction with FIG. 10 .
[0119] FIG10 is a process diagram of a video editing method provided by an embodiment of the present disclosure. Referring to FIG10 , it includes: Video A. Video A includes a first frame, a second frame, and a third frame. Each frame includes a ball and a square. In the first frame, the ball is on the left side of the square. In the second frame, the square partially obscures the ball. In the third frame, the ball is on the right side of the square. That is, in Video A, the square is stationary, and the ball moves from the left side of the square to the right side.
[0120] Please refer to Figure 10. The terminal device (not shown in Figure 10) can determine the target object in each frame of the image, wherein the target object in each frame of the image is a sphere. The terminal device can determine the center of the sphere as the target position of the sphere in each video frame, that is, point A in Figure 10 can be the target position in each frame of the image. If the cropping size is 600:800 (aspect ratio), the terminal device determines a 600*800 cropping area with point A as the center, and crops the image based on the cropping area to obtain the target video frame image corresponding to each video frame.
[0121] As shown in Figure 10, the terminal device can combine three target video frames into Video B. The first frame of Video B includes the sphere and a portion of the square located to the right of the sphere; the second frame includes the square and a portion of the sphere (the square obscures the sphere); and the third frame includes the sphere and a portion of the square located to the left of the sphere. This allows the camera to follow the movement of the sphere, eliminating the need for manual video editing and improving the accuracy and efficiency of video cropping.
[0122] FIG11 is a schematic diagram of the structure of a video editing device provided by an embodiment of the present disclosure. Referring to FIG11 , the video editing device 110 includes a first acquisition module 111, a determination module 112, a cropping module 113, and a generation module 114, wherein:
[0123] The first acquisition module 111 is used to acquire a first video frame image and a preset cropping size in the video to be edited;
[0124] The determining module 112 is configured to analyze the first video frame image according to the plug-in to determine the target position in the first video frame image;
[0125] The cropping module 113 is configured to crop the first video frame image according to the cropping size and target position to obtain a target video frame image;
[0126] The generating module 114 is configured to generate a cropped video according to the target video frame image.
[0127] According to one or more embodiments of the present disclosure, the determining module 112 is specifically configured to:
[0128] Determining a position of a target object in the first video frame image;
[0129] The position of the target object is smoothed to obtain the target position in the first video frame image.
[0130] According to one or more embodiments of the present disclosure, the determining module 112 is specifically configured to:
[0131] Storing the position of the target object in a cache, wherein the cache includes multiple positions of the target object in multiple video frame images;
[0132] Smoothing is performed on the position of the target object and the multiple positions in the cache to obtain the target position of the first video frame image.
[0133] According to one or more embodiments of the present disclosure, the determining module 112 is specifically configured to:
[0134] Determining a target area of the target object in the first video frame image;
[0135] A position of a target object in the first video frame image is determined according to the target area.
[0136] According to one or more embodiments of the present disclosure, the determining module 112 is specifically configured to:
[0137] The center position of the target area is determined as the position of the target object.
[0138] According to one or more embodiments of the present disclosure, the determining module 112 is specifically configured to:
[0139] Obtaining a target position of the first video frame image in the cache;
[0140] The first video frame image is cropped with the target position as the cropping center according to the cropping size to obtain the target video frame image.
[0141] The video editing device provided in the embodiment of the present disclosure can be used to execute the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0142] FIG12 is a schematic diagram of the structure of another video editing device provided by an embodiment of the present disclosure. Based on the embodiment shown in FIG11 , referring to FIG12 , the video editing device 110 further includes a second acquisition module 115 , wherein the second acquisition module 115 is configured to:
[0143] Acquire a second video frame image between the first video frame image and the next extracted video frame image;
[0144] Obtaining a smooth curve corresponding to the target position of the first video frame image and the target position of the next extracted video frame image;
[0145] A target position of the second video frame image is determined according to the smooth curve.
[0146] The video editing device provided in the embodiment of the present disclosure can be used to execute the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0147] FIG13 is a schematic diagram of the structure of a terminal device provided by an embodiment of the present disclosure. Please refer to FIG13 , which shows a schematic diagram of the structure of a terminal device 1300 suitable for implementing an embodiment of the present disclosure. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The terminal device shown in FIG13 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.
[0148] As shown in FIG13 , terminal device 1300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1301, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1302 or programs loaded from a storage device 1308 into a random access memory (RAM) 1303. RAM 1303 also stores various programs and data required for the operation of terminal device 1300. Processing device 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to bus 1304.
[0149] Typically, the following devices may be connected to the I / O interface 1305: an input device 1306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1309. The communication device 1309 may allow the terminal device 1300 to communicate with other devices wirelessly or by wire to exchange data. Although FIG13 illustrates a terminal device 1300 having various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.
[0150] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 1309, or installed from the storage device 1308, or installed from the ROM 1302. When the computer program is executed by the processing device 1301, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0151] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0152] The computer-readable medium may be included in the terminal device, or may exist independently without being incorporated into the terminal device.
[0153] The computer-readable medium carries one or more programs. When the one or more programs are executed by the terminal device, the terminal device executes the method shown in the above embodiment.
[0154] An embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, various methods that may be involved in the above embodiments are implemented.
[0155] An embodiment of the present disclosure provides a computer program product, including a computer program, which implements various possible methods involved in the above embodiments when executed by a processor.
[0156] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0157] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0158] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0159] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0160] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0161] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0162] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0163] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0164] For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require obtaining and using the user's personal information. Thereby, the user can independently choose whether to provide personal information to the software or hardware such as the terminal device, application, server or storage medium that performs the operation of the technical solution of the present disclosure based on the prompt message. As an optional but non-limiting implementation method, in response to receiving an active request from the user, the method of sending the prompt message to the user can be, for example, a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the terminal device.
[0165] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0166] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws and regulations. Data may include information, parameters and messages, such as flow switching indication information.
[0167] In a first aspect, an embodiment of the present disclosure provides a video editing method, the video editing method comprising:
[0168] Obtaining the first video frame image and a preset cropping size in the video to be edited;
[0169] analyzing the first video frame image according to the plug-in to determine a target position in the first video frame image, and performing a cropping process on the first video frame image according to the cropping size and the target position to obtain a target video frame image, wherein the target position is determined according to a position of a target object to be cropped in the first video frame image;
[0170] A cropped video is generated according to the target video frame image.
[0171] According to one or more embodiments of the present disclosure, analyzing the first video frame image according to the plug-in to determine the target position in the first video frame image includes:
[0172] Determining a position of a target object in the first video frame image;
[0173] The position of the target object is smoothed to obtain the target position in the first video frame image.
[0174] According to one or more embodiments of the present disclosure, the step of smoothing the position of the target object to obtain the target position in the first video frame image includes:
[0175] Storing the position of the target object in a cache, wherein the cache includes multiple positions of the target object in multiple video frame images;
[0176] Smoothing is performed on the position of the target object and the multiple positions in the cache to obtain the target position of the first video frame image.
[0177] According to one or more embodiments of the present disclosure, determining the position of the target object in the first video frame image includes:
[0178] Determining a target area of the target object in the first video frame image;
[0179] A position of a target object in the first video frame image is determined according to the target area.
[0180] According to one or more embodiments of the present disclosure, determining the position of the target object in the first video frame image according to the target area includes:
[0181] The center position of the target area is determined as the position of the target object.
[0182] According to one or more embodiments of the present disclosure, the step of performing cropping processing on the first video frame image according to the cropping size and the target position to obtain the target video frame image includes:
[0183] Obtaining a target position of the first video frame image in the cache;
[0184] The first video frame image is cropped with the target position as the cropping center according to the cropping size to obtain the target video frame image.
[0185] According to one or more embodiments of the present disclosure, the first video frame image is an image obtained based on interval frame extraction, and after smoothing the position of the target object to obtain the target position in the first video frame image, the method further includes:
[0186] Acquire a second video frame image between the first video frame image and the next extracted video frame image;
[0187] Obtaining a smooth curve corresponding to the target position of the first video frame image and the target position of the next extracted video frame image;
[0188] A target position of the second video frame image is determined according to the smooth curve.
[0189] In a second aspect, an embodiment of the present disclosure provides a video editing device, the video editing device including a first acquisition module, a determination module, a cropping module, and a generation module, wherein:
[0190] The first acquisition module is used to acquire a first video frame image and a preset cropping size in the video to be edited;
[0191] The determination module is configured to analyze the first video frame image according to the plug-in to determine the target position in the first video frame image;
[0192] The cropping module is configured to crop the first video frame image according to the cropping size and the target position to obtain a target video frame image;
[0193] The generating module is used to generate a cropped video according to the target video frame image.
[0194] According to one or more embodiments of the present disclosure, the determining module is specifically configured to:
[0195] Determining a position of a target object in the first video frame image;
[0196] The position of the target object is smoothed to obtain the target position in the first video frame image.
[0197] According to one or more embodiments of the present disclosure, the determining module is specifically configured to:
[0198] Storing the position of the target object in a cache, wherein the cache includes multiple positions of the target object in multiple video frame images;
[0199] Smoothing is performed on the position of the target object and the multiple positions in the cache to obtain the target position of the first video frame image.
[0200] According to one or more embodiments of the present disclosure, the determining module is specifically configured to:
[0201] Determining a target area of the target object in the first video frame image;
[0202] A position of a target object in the first video frame image is determined according to the target area.
[0203] According to one or more embodiments of the present disclosure, the determining module is specifically configured to:
[0204] The center position of the target area is determined as the position of the target object.
[0205] According to one or more embodiments of the present disclosure, the determining module is specifically configured to:
[0206] Obtaining a target position of the first video frame image in the cache;
[0207] The first video frame image is cropped with the target position as the cropping center according to the cropping size to obtain the target video frame image.
[0208] According to one or more embodiments of the present disclosure, the video editing device further includes a second acquisition module, wherein the second acquisition module is configured to:
[0209] Acquire a second video frame image between the first video frame image and the next extracted video frame image;
[0210] Obtaining a smooth curve corresponding to the target position of the first video frame image and the target position of the next extracted video frame image;
[0211] A target position of the second video frame image is determined according to the smooth curve.
[0212] In a third aspect, an embodiment of the present disclosure provides a terminal device including: a processor and a memory;
[0213] The memory stores computer-executable instructions;
[0214] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video editing method as described in the first aspect and various possible aspects of the first aspect.
[0215] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video editing method as described in the first aspect and various possible aspects of the first aspect are implemented.
[0216] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0217] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0218] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A video editing method, comprising: Obtaining a first video frame image in a video to be edited and a preset cropping size; Analyzing the first video frame image according to a plug-in to determine a target position in the first video frame image, and cropping the first video frame image according to the cropping size and the target position to obtain a target video frame image, where the target position is determined according to the position of a target object to be cropped in the first video frame image; Generating a cropped video according to the target video frame image.
2. The method according to claim 1, wherein, The analyzing the first video frame image according to the plug-in to determine the target position in the first video frame image includes: Determining the position of a target object in the first video frame image; Smoothing the position of the target object to obtain the target position in the first video frame image.
3. The method according to claim 2, wherein The smoothing the position of the target object to obtain the target position in the first video frame image includes: Storing the position of the target object in a cache, where the cache includes multiple positions of the target object in multiple video frame images; Smoothing the position of the target object and the multiple positions in the cache to obtain the target position of the first video frame image.
4. The method according to claim 2 or 3, wherein The determining the position of the target object in the first video frame image includes: Determining a target region of the target object in the first video frame image; Determining the position of the target object in the first video frame image according to the target region.
5. The method according to claim 4, wherein The determining the position of the target object in the first video frame image according to the target region includes: Determining the center position of the target region as the position of the target object.
6. The method according to any one of claims 3-5, wherein, The cropping the first video frame image according to the cropping size and the target position to obtain the target video frame image includes: Obtaining the target position of the first video frame image in the cache; Taking the target position as a cropping center and cropping the first video frame image according to the cropping size to obtain the target video frame image.
7. The method according to any one of claims 2-6, wherein, The first video frame image is an image obtained by interval frame extraction. After smoothing the position of the target object to obtain the target position in the first video frame image, the method further includes: Obtaining a second video frame image between the first video frame image and the next video frame image obtained by frame extraction; Obtaining a smoothing curve corresponding to the target position of the first video frame image and the target position of the next video frame image obtained by frame extraction; Determining the target position of the second video frame image according to the smoothing curve.
8. A video editing apparatus, including a first obtaining module, a determining module, a cropping module, and a generating module, where The first obtaining module is configured to obtain a first video frame image in a video to be edited and a preset cropping size; The determining module is configured to analyze the first video frame image according to a plug-in to determine the target position in the first video frame image; The cropping module is configured to crop the first video frame image according to the cropping size and the target position to obtain a target video frame image; The generating module is configured to generate a cropped video according to the target video frame image.
9. A terminal device, comprising a processor and a memory, wherein, The memory is configured to store computer-executable instructions; The processor is configured to execute the computer-executable instructions stored in the memory, so that the processor executes the video editing method according to any one of claims 1-7.
10. A computer-readable storage medium stores computer-executable instructions, wherein, When the processor executes the computer-executable instructions, the video editing method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Video clipping method and device, storage medium and electronic equipment
CN112561839A
Video cropping method and device, storage medium and electronic equipment
CN112565890A
Machine vision system capable of automatically identifying, tracking, shooting and editing
CN112702535A
Method for providing a video, transmitting device, and receiving device
US20140245367A1
Video processing method, video processing device, and storage medium
US20210287009A1