Video editing method and device and terminal equipment

The terminal device uses plug-ins to analyze the target position of the video frame image for nonlinear cropping, which solves the problem of low video cropping efficiency, and realizes automatic cropping and efficient video generation.

CN120378763APending Publication Date: 2025-07-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410094711.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In the prior art, video cutting efficiency is low and manual processing complexity is high, resulting in low video cutting efficiency.

Method used

The terminal device obtains the video frame image to be edited and the preset crop size, uses the plug-in to analyze the target position and performs non-linear cropping to generate the cropped video.

Benefits of technology

It realizes video cropping without manual participation, reduces complexity, improves video cropping efficiency, and avoids jitter in the cropped video window and improves display effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378763A_ABST
    Figure CN120378763A_ABST
Patent Text Reader

Abstract

The invention provides a video editing method and apparatus, and a terminal device. The method comprises the steps of obtaining a first video frame image and a preset cutting size in a to-be-edited video; analyzing the first video frame image according to a plug-in, determining a target position in the first video frame image, and cutting the first video frame image according to the cutting size and the target position to obtain a target video frame image, the target position is determined according to the position of a target object to be cut in the first video frame image; and generating a clipped video according to the target video frame image. And the video cutting efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of video processing technologies, and in particular, to a video editing method, apparatus, and terminal device. Background Art

[0002] Users can perform editing processing on a video with a fixed shooting lens to obtain a video with an effect that the shooting lens follows the movement of the characters in the video.

[0003] Currently, users can obtain a video with the lens following the movement of the characters through manual editing. For example, the user can determine the character images in each frame of the image, crop the character images, and splice the multiple cropped images to obtain a video with the effect that the lens follows the movement of the characters. However, the complexity of manual processing is relatively high, resulting in low efficiency of video cropping. Summary of the Invention

[0004] The present disclosure provides a video editing method, apparatus, and terminal device for solving the technical problem of low efficiency of video cropping in the prior art.

[0005] In a first aspect, the present disclosure provides a video editing method, which includes:

[0006] Obtain a first video frame image in a video to be edited and a preset cropping size;

[0007] Analyze the first video frame image according to a plug-in to determine a target position in the first video frame image, and crop the first video frame image according to the cropping size and the target position to obtain a target video frame image;

[0008] Generate a cropped video according to the target video frame image.

[0009] In a second aspect, the present disclosure provides a video editing apparatus, which includes a first acquisition module, a determination module, a cropping module, and a generation module, where:

[0010] The first acquisition module is configured to obtain a first video frame image in a video to be edited and a preset cropping size;

[0011] The determination module is configured to analyze the first video frame image according to a plug-in to determine a target position in the first video frame image;

[0012] The cropping module is configured to crop the first video frame image according to the cropping size and the target position to obtain a target video frame image;

[0013] The generation module is configured to generate a cropped video according to the target video frame image.

[0014] In a third aspect, embodiments of the present disclosure provide a terminal device, including: a processor and a memory;

[0015] The memory stores computer-executable instructions;

[0016] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video editing method as described in the first aspect and various possible aspects related to the first aspect above.

[0017] In a fourth aspect, embodiments of the present disclosure provide a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the video editing method as described in the first aspect and various possible aspects related to the first aspect above is implemented.

[0018] The present disclosure provides a video editing method, apparatus, and terminal device. The terminal device can obtain the first video frame image in the video to be edited and a preset cropping size, analyze the first video frame image according to the plug-in, and determine the target position in the first video frame image, where the target position is determined according to the position of the target object to be cropped in the first video frame image. The terminal device can crop the first video frame image according to the cropping size and the target position to obtain the target video frame image, and generate a cropped video according to the target video frame image. In the above method, since the terminal device can perform non-linear cropping processing on the video frame image based on the plug-in, the terminal device can automatically generate a cropped video without manual participation, reducing the complexity of video cropping and improving the efficiency of video cropping. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic diagram of an application scenario provided by an embodiment of the present disclosure;

[0021] Figure 2 It is a schematic flowchart of a video editing method provided by an embodiment of the present disclosure;

[0022] Figure 3 It is a schematic diagram of a process for determining a target object provided by an embodiment of the present disclosure;

[0023] Figure 4Schematic diagram of a process for determining the position of a target object provided by an embodiment of the present disclosure;

[0024] Figure 5A Schematic diagram of a cropping area provided by an embodiment of the present disclosure;

[0025] Figure 5B Schematic diagram of another cropping area provided by an embodiment of the present disclosure;

[0026] Figure 6 Schematic diagram of a process for determining a target video frame image provided by an embodiment of the present disclosure;

[0027] Figure 7 Processing process of a plug-in provided by an embodiment of the present disclosure;

[0028] Figure 8 Schematic diagram of a method for cropping a second video frame image provided by an embodiment of the present disclosure;

[0029] Figure 9 Schematic diagram of a second video frame image provided by an embodiment of the present disclosure;

[0030] Figure 10 Schematic diagram of a process of a video editing method provided by an embodiment of the present disclosure;

[0031] Figure 11 Schematic diagram of the structure of a video editing device provided by an embodiment of the present disclosure;

[0032] Figure 12 Schematic diagram of the structure of another video editing device provided by an embodiment of the present disclosure;

[0033] Figure 13 Schematic diagram of the structure of a terminal device provided by an embodiment of the present disclosure. Detailed implementation manners

[0034] Here, exemplary embodiments will be described in detail, and examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0035] For ease of understanding, the concepts related to the embodiments of the present disclosure will be described below.

[0036] Terminal device: A device with wireless transceiver capabilities. The terminal device can be deployed on land, including indoor or outdoor, handheld, wearable, or vehicle-mounted. The terminal device can be a mobile phone, a tablet (Pad), a computer with wireless transceiver capabilities, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a vehicle-mounted terminal device, a wireless terminal in self-driving, a wireless terminal device in remote medical, a wireless terminal device in smart grid, a wireless terminal device in transportation safety, a wireless terminal device in smart city, a wireless terminal device in smart home, a wearable terminal device, etc. The terminal device involved in the embodiments of the present disclosure can also be referred to as a terminal, a user equipment (UE), an access terminal device, a vehicle-mounted terminal, an industrial control terminal, a UE unit, a UE station, a mobile station, a mobile unit, a remote station, a remote terminal device, a mobile device, a UE terminal device, a wireless communication device, a UE agent, or a UE device, etc. The terminal device can also be fixed or mobile.

[0037] In the related art, a user can perform clip processing on a video with a fixed shooting lens to obtain a video with an effect that the shooting lens follows the movement of the characters in the video. For example, the position of the lens for shooting the video is fixed, and the characters in the video shot by this lens move from the left side to the right side of the video. The user can perform clip processing on this video to obtain a video with an effect that the lens moves from the left side to the right side following the user. Currently, the user can perform clip processing on the video based on manual clip. For example, the user can determine the character images in each frame of the video based on a manual method, crop the character images, and splice the multiple cropped images to obtain a video with an effect that the lens follows the movement of the characters. However, the complexity of manual processing is relatively high, resulting in low efficiency of video cropping.

[0038] To solve the technical problems in the related art, an embodiment of the present disclosure provides a video editing method. The terminal device can obtain the first video frame image in the video to be edited and a preset cropping size, and determine the position of the target object in the first video frame image, and store the position of the target object in the cache, where the cache includes the positions of the target object in multiple video frame images. Smooth the position of the target object in the cache and multiple positions to obtain the target position of the first video frame image. The terminal device can crop the first video frame image according to the cropping size and the target position to obtain the target video frame image, and generate a cropped video according to the target video frame image. In this way, since the terminal device can smooth the position of the target object, the problem of window jitter in the cropped video can be avoided, and the video display effect can be improved. Moreover, since the terminal device can perform non-linear cropping processing on the video frame image based on the plug-in, the terminal device can automatically generate a cropped video without manual participation, reducing the complexity of video cropping and improving the efficiency of video cropping.

[0039] Next, in combination with Figure 1 , the application scenarios of the embodiments of the present disclosure will be described.

[0040] Figure 1 FIG. is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. Please refer to Figure 1 , including: Video A and a terminal device. Among them, Video A may include 3 frame images. In the first frame image, the sphere is on the left side of the square. In the second frame image, the square blocks the sphere. In the third frame image, the sphere is on the right side of the square. That is, in Video A, the sphere moves from the left side of the square to the right side of the square (the position of the camera shooting this Video A is fixed). After the terminal device obtains Video A, it can display Video A. After the user clicks the editing control, the terminal device can generate Video B, where Video B also includes 3 frame images. The first frame image includes the sphere and part of the square. The second frame image includes the square and the blocked sphere. The third frame image includes part of the square and the sphere. In this way, the video effect of Video B is the effect of the camera following the sphere moving, without manual cropping of the video, reducing the complexity of video cropping and improving the efficiency of video cropping.

[0041] It should be noted 1, Figure 1 This is only an example of the application scenario of the embodiments of the present disclosure, and is not a limitation on the application scenario of the embodiments of the present disclosure.

[0042] Next, the technical solutions of the present disclosure and how the technical solutions of the present disclosure solve the above technical problems will be described in detail with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. Next, the embodiments of the present disclosure will be described with reference to the drawings.

[0043] Figure 2 The flowchart of a video editing method provided by an embodiment of the present disclosure. Please refer to Figure 2 , the method may include:

[0044] S201. Obtain a first video frame image in the video to be edited and a preset cropping size.

[0045] The execution subject of the embodiment of the present disclosure may be a terminal device or a video editing device provided in the terminal device. Among them, the video editing device may be implemented based on software, or may be implemented based on the combination of software and hardware. The embodiment of the present disclosure does not limit this.

[0046] Among them, the video to be edited may be a video with a fixed perspective. For example, the shooting perspective of the video to be edited may be a fixed perspective. For example, the video to be edited may be a video obtained after the camera is fixed.

[0047] Optionally, the video to be edited may include an object. For example, the video to be edited may include any object such as an animal, a person, an object, etc., and the video to be edited may also include any other type of object. The embodiment of the present disclosure does not limit this.

[0048] It should be noted that the terminal device may obtain the video to be edited based on any feasible implementation method (for example, the terminal device may obtain the video to be edited in a database, or may receive the video to be edited sent by other devices, or may capture the video to be edited in real time). The embodiment of the present disclosure does not limit this.

[0049] Among them, the first video frame image may be an image in the video to be edited. For example, as Figure 1 shown, Video A may be the video to be edited, and the first frame image, the second frame image, and the third frame image in Video A may be the first video frame images.

[0050] Optionally, the cropping size may be a pre-set image size. For example, the cropping size may be 3:4, or may be 16:9, etc. The embodiment of the present disclosure does not limit this. For example, if the cropping size is 3:4, the terminal device may crop the first video frame image to a size of 3:4. If the cropping size is 16:9, the terminal device may crop the first video frame image to a size of 16:9.

[0051] It should be noted that the terminal device may determine the cropping size based on service requirements (for example, when publishing a video on a platform, if the required size of the video is 3:4, the preset cropping size corresponding to the platform may be 3:4), and the terminal device may also determine the preset cropping size based on any feasible implementation method. The embodiment of the present disclosure does not limit this.

[0052] S202. Analyze the first video frame image according to the plug-in to determine the target position in the first video frame image.

[0053] Optionally, the user interface (UI) side of the terminal device can send the video to be edited to the video editing side. The video editing side can read the video stream, perform frame extraction on the video to be edited, and the extracted video frames can be sent to the plug-in based on the link of the video editing side. In this way, the plug-in can obtain the video frame image and process the video frame image.

[0054] Among them, the target position is determined according to the position of the target object to be cropped in the first video frame image. The terminal device can determine the target position in the first video frame image according to the following feasible implementation: determine the position of the target object in the first video frame image, and smooth the position of the target object to obtain the target position in the first video frame image.

[0055] Optionally, the target object can be the object to be cropped. For example, the target object can be a moving object in the video to be edited. For example, the video to be edited may include object 1 and object 2. If object 1 moves and object 2 is stationary, the target object in the first video frame image can be object 1. If object 1 is stationary and object 2 moves, the target object in the first video frame image can be object 2.

[0056] Optionally, the terminal device can determine the target object in the first video frame image according to the following feasible implementation: process the first video frame image according to the object detection algorithm to obtain the target object in the first video frame image.

[0057] Among them, the target object can include a portrait object. For example, the target object in the first video frame image can be a human face, portrait limbs, a portrait torso, etc. The embodiments of the present disclosure do not limit this.

[0058] It should be noted that the above embodiments are only examples of the target object, not limitations on the target object. For example, the target object can also be a moving football, a flying bird, a swimming fish, etc. The embodiments of the present disclosure do not limit this.

[0059] Optionally, an object detection algorithm can be used to detect a target object in the first video frame image. For example, the object detection algorithm can be an Artificial Intelligence (AI) algorithm. For example, if the target object is a human face, the object detection algorithm can be a face detection algorithm. Based on this face detection algorithm to process the first video frame image, the terminal device can determine whether the first video frame image includes a face image. For example, if the target object is a limb, the object detection algorithm can be a limb detection algorithm. Based on this limb detection algorithm to process the first video frame image, the terminal device can determine whether the first video frame image includes a limb image

[0060] It should be noted that the terminal device can obtain the object detection algorithm based on any feasible implementation manner, and the object detection algorithm can also be implemented based on a model. The embodiments of the present disclosure do not limit this

[0061] It should be noted that the terminal device can determine the target object in the first video frame image based on any feasible implementation manner (such as, implemented based on a trained model). The embodiments of the present disclosure do not limit this

[0062] Next, in combination with Figure 3 , the determination of the target object in the first video frame image will be described

[0063] Figure 3 is a schematic diagram of a process for determining a target object provided by an embodiment of the present disclosure. Please refer to Figure 3 , including a first video frame image and a face detection module. Among them, the first video frame image may include a portrait, and the target object may be the face in the portrait. The terminal device ( Figure 3 not shown) can process the first video frame image based on the face detection module, and then can determine that the first video frame image may include a face, and the face detection module can also determine the position or area of the face in the first video frame image, which can improve the recognition accuracy of the target object, and then improve the accuracy of video editing

[0064] Among them, the terminal device can determine the position of the target object in the first video frame image according to the following feasible implementation manner: determine the target area of the target object in the first video frame image, and according to the target area, determine the position of the target object in the first video frame image

[0065] Among them, the target area can be the area occupied by the target object in the first video frame image. For example, the target area can be the area enclosed by the contour of the target object in the first video frame image, or the target area can also be the area enclosed by multiple vertices of the target object (such as the upper vertex, lower vertex, left vertex, and right vertex) in the first video frame image. The embodiments of the present disclosure do not limit this.

[0066] Optionally, when the terminal device processes the first video frame image based on the object detection algorithm, it can obtain the target area of the target object in the first video frame image. The terminal device can also determine the target area of the target object in the first video frame image based on any feasible implementation manner. The embodiments of the present disclosure do not limit this.

[0067] Optionally, the terminal device determines the position of the target object in the first video frame image according to the target area. Specifically, it can be: determining the central position of the target area as the position of the target object.

[0068] Among them, the central position can be the geometric center of the target area or the midpoint of the central axis of the target area. The embodiments of the present disclosure do not limit this.

[0069] Among them, the terminal device can determine the coordinates of the central position of the target area, and then determine the coordinates of this central position as the position of the target object. For example, if the center of the target area is the 10th pixel of the long side (i.e., the x-axis of the pixel coordinate system) of the first video frame image and the 20th pixel of the high side (i.e., the y-axis of the pixel coordinate system), then the position of the target object can be (10, 20). If the center of the target area is the 100th pixel of the long side of the first video frame image and the 70th pixel of the high side, then the position of the target object can be (100, 70).

[0070] It should be noted that the terminal device can determine the central position of the target area as the position of the target object in the first video frame image, or the terminal device can also determine any point in the target area as the position of the target object in the first video frame image. The embodiments of the present disclosure do not limit this.

[0071] Next, in combination with Figure 4 , the process of the terminal device determining the position of the target object in the first video frame image will be described.

[0072] Figure 4 This is a schematic diagram of a process for determining the position of a target object provided by an embodiment of the present disclosure. Please refer to Figure 4 , including: the first video frame image. Among them, the first video frame image may include a portrait. The terminal device ( Figure 4(not shown) can process the first video frame image based on a face detection algorithm to determine the face image in the first video frame image and the target area corresponding to the face image. The terminal device can determine the center of the target area as the position of the face in the first video frame image. In this way, the terminal device can accurately determine the position of the target object in the first video frame image, thereby improving the accuracy of video editing.

[0073] Among them, the terminal device can perform smoothing processing on the position of the target object to obtain the target position in the first video frame image, which can avoid the window jitter of the generated cropped video and thus improve the display effect of the video.

[0074] Among them, the terminal device performs smoothing processing on the position of the target object to obtain the target position in the first video frame image. Specifically, it can be: store the position of the target object in the cache, and perform smoothing processing on the position of the target object in the cache and multiple positions to obtain the target position in the first video frame image.

[0075] Among them, the cache can include multiple positions of the target object in multiple video frame images. For example, the video to be edited may include multiple first video frame images. Therefore, the terminal device can determine the position of the target object in each first video frame image and store multiple positions in the cache. The terminal device can perform smoothing processing on the multiple positions to obtain the target position corresponding to each first video frame image. For example, the terminal device can store the position 1 of the target object in the first video frame A and the position 2 of the target object in the first video frame B in the cache. The terminal device can perform smoothing processing on the position 1 and the position 2 to obtain the target position a corresponding to the position 1 and the target position b corresponding to the position 2. Among them, the target position a can be the target position of the first video frame A, and the target position b can be the target position of the first video frame B. In this way, since the target position is the position after smoothing processing, the window jitter of the cropped video generated by the terminal device is small, improving the video playback effect.

[0076] It should be noted that the terminal device can perform smoothing processing on multiple positions in the cache to obtain the target position, or when the position difference between the target object in two adjacent first video frame images is greater than a preset threshold, perform smoothing processing on these two positions. The embodiments of the present disclosure do not limit this.

[0077] It should be noted that the terminal device can perform smoothing processing on the position of the target object in the cache and multiple positions based on any feasible implementation manner. The embodiments of the present disclosure do not limit this.

[0078] S203. Crop the first video frame image according to the crop size and the target position to obtain the target video frame image.

[0079] Among them, the target video frame image can be the video frame image after cropping the first video frame image, and the target video frame image can include the target object. For example, based on the target position and the cropping size, the terminal device can crop out the target video frame image including the target object from the first video frame image, and the size of the target video frame image is the same as the preset cropping size. In this way, the cropping accuracy of the target video frame image and the efficiency of video editing can be improved.

[0080] Among them, the terminal device can obtain the target video frame image based on the following feasible implementation: obtain the target position of the first video frame image in the cache, take the target position as the cropping center, and crop the first video frame image according to the cropping size to obtain the target video frame image. Since the preset cropping size matches the service requirements (such as a 16:9 video, etc.), not only can the efficiency of video cropping be improved, but also the accuracy of video cropping can be improved.

[0081] Optionally, the terminal device can determine the cropping area in the first video frame image according to the target position and the cropping size, and perform cropping processing on the cropping area to obtain the target video frame image. For example, the cropping area can include the target object, and the image in the cropping area can be the target video frame image.

[0082] Among them, the size of the cropping area is the same as the cropping size. For example, if the preset cropping size is a 16:9 size, the size of the cropping area can also be a 16:9 size; if the preset cropping size is 800*600, the size of the cropping area can also be 800*600.

[0083] Next, in combination with Figure 5A - Figure 5B , the cropping area will be described.

[0084] Figure 5A This is a schematic diagram of a cropping area provided by an embodiment of the present disclosure. Please refer to Figure 5A , including the first video frame image. Among them, the first video frame image can include a portrait. If the target object is the face in the portrait, the terminal device ( Figure 5A not shown) can determine the center of the face as the target position of the target object in the first video frame image, that is, point A is the target position of the first video frame image. If the preset cropping size is 9:16 (aspect ratio, other cropping sizes in this embodiment of the present disclosure can also be aspect ratios, and this embodiment of the present disclosure will not elaborate further), the terminal device takes point A as the center, determines a 16:9 area, and determines this area as the cropping area. Among them, the cropping area can include the face image, and the center of the cropping area can be the center of the face.

[0085] Figure 5BAnother schematic diagram of a cropping area provided by an embodiment of the present disclosure. Please refer to Figure 5B , including a first video frame image. Among them, the first video frame image may include a portrait. If the target object is the face in the portrait, the terminal device ( Figure 5B not shown) may determine the center of the face as the target position of the target object in the first video frame image, that is, point A is the target position of the first video frame image. If the preset cropping size is 600*800 (number of pixels), the terminal device may determine an 800*600 area in the first video frame image based on point A and determine this area as the cropping area. Among them, the cropping area includes a face image, and the center of the face image may be located on the central axis of the cropping area (if it is located at the center of the cropping area, sufficient pixels cannot be obtained, and the center of the face image may also be located at any position in the cropping area, and the embodiments of the present disclosure do not limit this), so that the accuracy of cropping can be improved and the video display effect can be improved.

[0086] Optionally, the electronic device may determine the upper left vertex, the upper right vertex, the lower left vertex, and the lower right vertex in the first video frame image based on the target position and the cropping size, and determine the area enclosed by the upper left vertex, the upper right vertex, the lower left vertex, and the lower right vertex as the cropping area. For example, the coordinates of the upper left vertex, the upper right vertex, the lower left vertex, and the lower right vertex may be the pixel coordinates of the four vertices of the cropping area in the first video frame image. For example, if the cropping area is a rectangle, the terminal device may determine the pixel coordinates of the upper left vertex of the cropping area in the first video frame image as the upper left vertex coordinates, the terminal device may determine the pixel coordinates of the upper right vertex of the cropping area in the first video frame image as the upper right vertex coordinates, the terminal device may determine the pixel coordinates of the lower left vertex of the cropping area in the first video frame image as the lower left vertex coordinates, and the terminal device may determine the pixel coordinates of the lower right vertex of the cropping area in the first video frame image as the lower right vertex coordinates.

[0087] Optionally, after the terminal device determines the upper left vertex, the upper right vertex, the lower left vertex, and the lower right vertex, it may determine the area enclosed by the upper left vertex, the upper right vertex, the lower left vertex, and the lower right vertex as the cropping area. For example, if the coordinates of the upper left vertex are (1, 3), the coordinates of the upper right vertex are (3, 3), the coordinates of the lower left vertex are (1, 0), and the coordinates of the lower right vertex are (3, 0), the terminal device may determine the area enclosed by (1, 3), (3, 3), (1, 0), and (3, 0) as the cropping area.

[0088] It should be noted that the terminal device can determine the cropping area in the first video frame image based on any other feasible implementation manner, and the embodiments of the present disclosure do not limit this.

[0089] Next, in combination with Figure 6 , the process of the terminal device determining the target video frame image will be described.

[0090] Figure 6 FIG. is a schematic diagram of a process for determining a target video frame image provided by an embodiment of the present disclosure. Please refer to Figure 6 , including: a first video frame image. Among them, the first video frame image may include a human figure. If the target object is the face in the human figure, the terminal device ( Figure 6 not shown) may determine the center of the face as the target position of the target object in the first video frame image, that is, point A is the target position of the first video frame image.

[0091] Please refer to Figure 6 , if the preset cropping size is 800*800 (number of pixels), the terminal device takes point A as the center and determines the coordinates B of the upper left vertex, the coordinates C of the lower left vertex, the coordinates D of the upper right vertex, and the coordinates E of the lower right vertex in the first video frame image based on the cropping size of 800*800. The terminal device may determine the area enclosed by B, C, D, and E as the cropping area and perform cropping processing on the cropping area to obtain the target video frame image. Among them, the target video frame image may include a face image, and the center of the face image may be located at the center of the target video frame image. In this way, the terminal device can automatically perform cropping processing on the first video frame image to obtain the target video frame image including the target object, thereby improving the efficiency of image cropping and the efficiency of video editing.

[0092] Next, in combination with Figure 7 , the processing process of the above plug-in will be described.

[0093] Figure 7 FIG. is a processing process of a plug-in provided by an embodiment of the present disclosure. Please refer to Figure 7 , including: a plug-in. Among them, the plug-in can receive the video frame image sent by the video editing side, and the plug-in can process the video frame image, and then generate cropping data and store the cropping data in the cache. The cropping data may be the position of the target object in the video frame image, and the terminal device can perform smoothing processing on the cropping data in the cache. When the plug-in uses the cropping data, it can perform cropping processing on the video frame image according to the preset cropping size and the cropping data to obtain the target video frame image. In this way, the terminal device can achieve non-linear intelligent cropping based on this plug-in, improve the cropping efficiency of the video frame image, and thereby improve the generation efficiency of the cropped video.

[0094] S204. Generate a cropped video based on the target video frame image.

[0095] Optionally, the terminal device may perform a synthesis process on the target video frame images corresponding to each first video frame image to obtain a cropped video. For example, after the terminal device obtains the target video frame image based on the plugin, it may send the target video frame image to the video synthesis side. After the video synthesis side obtains the target video frame images corresponding to each first video frame image (the video synthesis side may also process the target video frame images frame by frame, which is not limited in this embodiment of the present disclosure), it may generate a cropped video. The cropped video includes the target object, and the display effect of the cropped video is that the camera follows the movement of the target object.

[0096] Optionally, after the terminal device obtains the target video frame image based on the plugin, it may send the target video frame image to the on-screen module. The on-screen module may display the target video frame image on the preview page. The user can preview the cropping result of the terminal device. If an error occurs, the user can correct it in time. Moreover, the user can also adjust the cropping result, which can improve the accuracy of video cropping and enhance the user experience.

[0097] The embodiments of the present disclosure provide a video editing method. The terminal device may obtain the first video frame image in the video to be edited and the preset cropping size, and determine the position of the target object in the first video frame image, store the position of the target object in the cache, where the cache includes the positions of the target object in multiple video frame images, perform smoothing processing on the position of the target object and multiple positions in the cache to obtain the target position of the first video frame image. The terminal device may obtain the target position of the first video frame image in the cache, and use the target position as the cropping center to perform cropping processing on the first video frame image according to the cropping size to obtain the target video frame image. The terminal device may generate a cropped video based on the target video frame image. In this way, since the terminal device can perform smoothing processing on the position of the target object to obtain the target position, the problem of window jitter in the cropped video can be avoided, and the video display effect can be improved. Moreover, since the terminal device can perform non-linear cropping processing on the video frame image based on the plugin, the terminal device can automatically generate a cropped video without manual participation, reducing the complexity of video cropping and improving the efficiency of video cropping.

[0098] In Figure 2Based on the illustrated embodiments, in the above video editing method, if the first video frame image is an image obtained by interval frame extraction (the terminal device extracts the first video frame image from the video to be edited based on the interval frame extraction method), then there is no target position in the cache for the multiple video frame images that have not been frame-extracted in the video to be edited. Therefore, after the terminal device smooths the position of the target object to obtain the target position in the first video frame image, the above video editing method further includes a method for cropping the second video frame image that has not been frame-extracted. Below, in combination with Figure 8 , the method for cropping the second video frame image will be described in detail.

[0099] Figure 8 FIG. is a schematic diagram of a method for cropping a second video frame image provided by an embodiment of the present disclosure. Please refer to Figure 8 , and the method flow may include:

[0100] S801. Obtain a second video frame image between the first video frame image and the next frame-extracted video frame image.

[0101] Among them, the second video frame image is the video frame image between the first video frame image and the next frame-extracted video frame image. For example, the frame extraction method of the terminal device for the video to be edited includes frame-by-frame extraction and interval frame extraction. If the frame extraction method of the terminal device for the video to be edited is frame-by-frame extraction, then each video frame image in the video to be edited is the first video frame image. Based on the above video editing method, the terminal device can determine the target video frame image corresponding to each video frame image, and then obtain the cropped video. If the frame extraction method of the video to be edited is interval frame extraction, then the video frame image frame-extracted by the terminal device in the video to be edited is the first video frame image, and the video frame image not frame-extracted by the terminal device can be the second video frame image.

[0102] For example, the frame extraction method of the terminal device for the video to be edited is interval frame extraction. If the terminal device determines video frame image 1 as the current first video frame image and video frame image 2 as the next frame-extracted video frame image, then the terminal device can determine the video frame image between video frame image 1 and video frame image 2 as the second video frame image corresponding to video frame image 1 and video frame image 2.

[0103] It should be noted that the terminal device can also determine the second video frame image that has not been frame-extracted by the terminal device based on any other feasible implementation manner, and the embodiments of the present disclosure do not limit this.

[0104] Optionally, the terminal device may determine the frame extraction method of the video to be edited based on the optical flow information between video frame images. For example, if the optical flow difference between adjacent video frame images in the video to be edited is less than a preset optical flow value, the terminal device may determine that the frame extraction method of the video to be edited is interval frame extraction. If the optical flow difference between adjacent video frame images in the video to be edited is greater than or equal to the preset optical flow value, the terminal device may determine that the frame extraction method of the video to be edited is frame-by-frame extraction. In this way, the accuracy of frame extraction of the video to be edited can be improved.

[0105] It should be noted that the terminal device may also determine the frame extraction method of the video to be edited based on any other feasible implementation manner (for example, implemented based on a trained model), and the embodiments of the present disclosure do not limit this.

[0106] Next, Figure 9 is used to illustrate the second video frame image.

[0107] Figure 9 is a schematic diagram of a second video frame image provided by an embodiment of the present disclosure. Please refer to Figure 9 , including: the video to be edited. Among them, the video to be edited includes Image 1 (the 1st frame), Image 2 (the 2nd frame), Image 3 (the 3rd frame), Image 4 (the 4th frame), Image 5 (the 5th frame),..., Image n (the nth frame). If the frame extraction method of the video to be edited by the terminal device ( Figure 9 not shown) is interval frame extraction, the first video frame obtained by the terminal device through interval frame extraction may include Image 1, Image 5,..., Image n.

[0108] Please refer to Figure 9 , since Image 5 is the next first video frame of Image 1, the terminal device may determine that the second video frame corresponding to Image 1 and Image 5 may include Image 2, Image 3, and Image 4. In this way, the terminal device can quickly and accurately determine the second video frame, improving the accuracy and efficiency of video editing.

[0109] S802. Obtain a smooth curve corresponding to the target position of the first video frame image and the target position of the next video frame image drawn.

[0110] Optionally, after the terminal device performs smoothing processing on the target position of the first video frame image and the target position of the next video frame image drawn, a smooth curve corresponding to the smoothing processing can be obtained.

[0111] It should be noted that the terminal device may determine the smooth curve according to any feasible implementation manner, and the embodiments of the present disclosure do not limit this.

[0112] S803. Determine the target position of the second video frame image according to the smooth curve.

[0113] After the terminal device determines the smooth curve, it can determine the target position of the second video frame image according to the values of the abscissa and ordinate in the smooth curve. For example, if the smooth curve is the smooth curve corresponding to target position 1 and target position 2, the terminal device can determine the abscissa and ordinate corresponding to the points on this smooth curve as the target position of the second video frame image. For example, the coordinates of target position 1 of the first video frame image A are (1, 1), the coordinates of target position 2 of the first video frame image B are (3, 3), and there is 1 second video frame image between the first video frame image A and the first video frame image B (only for example, not limited), then the target position of this second video frame image can be the coordinates (2, 2), which can avoid the problem of window jitter in the cropped video and improve the video playback effect.

[0114] S804. Crop the second video frame image according to the target position and the cropping size of the second video frame image to obtain the target video frame image corresponding to the second video frame image.

[0115] After the terminal device determines the target position of the second video frame image, it can crop the second video frame image based on the cropping method of the first video frame image to obtain the target video frame image corresponding to the second video frame image. For example, after the terminal device determines the target position of the second video frame image, the terminal device takes the target position of the second video frame image as the center and crops the second video frame image according to the cropping size, and then the target video frame image corresponding to the second video frame image can be obtained. In this way, the terminal device can quickly crop each video frame image in the video to be edited, and thus improve the efficiency of video editing.

[0116] The embodiments of the present disclosure provide a method for cropping a second video frame image. The terminal device can obtain the second video frame image between the first video frame image and the next video frame image obtained by frame extraction, obtain the smooth curve corresponding to the target position of the first video frame image and the target position of the next video frame image obtained by frame extraction, determine the target position of the second video frame image according to the smooth curve, and crop the second video frame image according to the target position and the cropping size of the second video frame image to obtain the target video frame image corresponding to the second video frame image. In this way, the terminal device can quickly determine the target position of each second video frame image according to the smooth curve, and thus improve the efficiency of video cropping.

[0117] Based on any one of the above embodiments, below, in combination with Figure 10 ., the process of the above video editing method will be described.

[0118] Figure 10Schematic diagram of a video editing method provided by an embodiment of the present disclosure. Please refer to Figure 10 , including: Video A. Among them, Video A includes the first frame image, the second frame image, and the third frame image. Among them, each frame image includes a ball and a square. In the first frame, the ball is on the left side of the square. In the second frame, the square blocks part of the ball. In the third frame, the ball is on the right side of the square. That is, in Video A, the square remains stationary, and the ball moves from the left side of the square to the right side of the square.

[0119] Please refer to Figure 10 , the terminal device ( Figure 10 not shown) can determine the target object in each frame image. Among them, the target object in each frame image is the ball. The terminal device can determine the center of the ball as the target position of the ball in each video frame, that is, Figure 10 point A in can be the target position in each frame image. If the cropping size is 600:800 (aspect ratio), the terminal device takes point A as the center, determines a 600*800 cropping area, and performs cropping processing on the image based on the cropping area to obtain the target video frame image corresponding to each video frame.

[0120] Please refer to Figure 10 , the terminal device can synthesize 3 target video frame images into Video B. Among them, the first frame image of Video B includes the ball and part of the square on the right side of the ball. The second frame image includes the square and part of the ball (the square blocks the ball). The third frame image includes the ball and part of the square on the left side of the ball. In this way, the display effect of Video B can be the display effect of the camera following the movement of the ball, and moreover, there is no need to manually edit the video, which improves the accuracy and efficiency of video cropping.

[0121] Figure 11 Schematic diagram of the structure of a video editing device provided by an embodiment of the present disclosure. Please refer to Figure 11 The video editing device 110 includes a first acquisition module 111, a determination module 112, a cropping module 113, and a generation module 114, where:

[0122] The first acquisition module 111 is used to acquire the first video frame image in the video to be edited and the preset cropping size;

[0123] The determination module 112 is used to analyze the first video frame image according to the plug-in and determine the target position in the first video frame image;

[0124] The cropping module 113 is used to perform cropping processing on the first video frame image according to the cropping size and the target position to obtain the target video frame image;

[0125] The generating module 114 is configured to generate a cropped video based on the target video frame image.

[0126] According to one or more embodiments of the present disclosure, the determining module 112 is specifically configured to:

[0127] Determine the position of the target object in the first video frame image;

[0128] Smooth the position of the target object to obtain the target position in the first video frame image.

[0129] According to one or more embodiments of the present disclosure, the determining module 112 is specifically configured to:

[0130] Store the position of the target object in a cache, where the cache includes multiple positions of the target object in multiple video frame images;

[0131] Smooth the position of the target object and the multiple positions in the cache to obtain the target position in the first video frame image.

[0132] According to one or more embodiments of the present disclosure, the determining module 112 is specifically configured to:

[0133] Determine the target area of the target object in the first video frame image;

[0134] Determine the position of the target object in the first video frame image according to the target area.

[0135] According to one or more embodiments of the present disclosure, the determining module 112 is specifically configured to:

[0136] Determine the center position of the target area as the position of the target object.

[0137] According to one or more embodiments of the present disclosure, the determining module 112 is specifically configured to:

[0138] Obtain the target position of the first video frame image from the cache;

[0139] Use the target position as the cropping center and crop the first video frame image according to the cropping size to obtain the target video frame image.

[0140] The video editing device provided by the embodiments of the present disclosure can be used to execute the technical solutions of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0141] Figure 12 This is a schematic structural diagram of another video editing device provided by the embodiments of the present disclosure. Figure 11Based on the illustrated embodiments, please refer to Figure 12 , the video editing apparatus 110 further includes a second obtaining module 115, wherein the second obtaining module 115 is configured to:

[0142] Obtain a second video frame image between the first video frame image and the next video frame image obtained by frame extraction;

[0143] Obtain a smooth curve corresponding to the target position of the first video frame image and the target position of the next video frame image obtained by frame extraction;

[0144] Determine the target position of the second video frame image according to the smooth curve.

[0145] The video editing apparatus provided by the embodiments of the present disclosure can be used to execute the technical solutions of the above method embodiments, and its implementation principles and technical effects are similar, which will not be elaborated herein.

[0146] Figure 13 The following is a schematic structural diagram of a terminal device provided by an embodiment of the present disclosure. Please refer to Figure 13 , which shows a schematic structural diagram of a terminal device 1300 suitable for implementing the embodiments of the present disclosure. Among them, the terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (Personal Digital Assistant, abbreviated as PDA), tablet computers (Portable Android Device, abbreviated as PAD), portable multimedia players (Portable Media Player, abbreviated as PMP), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 13 The terminal device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0147] As Figure 13 shown, the terminal device 1300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1301, which can execute various appropriate actions and processes according to a program stored in a read-only memory (Read Only Memory, abbreviated as ROM) 1302 or a program loaded from a storage device 1308 into a random access memory (Random Access Memory, abbreviated as RAM) 1303. In the RAM 1303, various programs and data required for the operation of the terminal device 1300 are also stored. The processing device 1301, the ROM 1302, and the RAM 1303 are connected to each other through a bus 1304. An input / output (I / O) interface 1305 is also connected to the bus 1304.

[0148] Typically, the following devices can be connected to the I / O interface 1305: an input device 1306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1309. The communication device 1309 can allow the terminal device 1300 to communicate with other devices wirelessly or wireline to exchange data. Although Figure 13 the terminal device 1300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.

[0149] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 1309, or installed from the storage device 1308, or installed from the ROM 1302. When the computer program is executed by the processing device 1301, the above functions defined in the method of the embodiment of the present disclosure are performed.

[0150] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0151] The above computer-readable medium can be included in the above terminal device; or it can exist independently without being assembled into the terminal device.

[0152] The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the terminal device, the terminal device is caused to execute the method shown in the above embodiments.

[0153] The embodiments of the present disclosure provide a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the methods that may be involved in various aspects of the above embodiments are implemented.

[0154] The embodiments of the present disclosure provide a computer program product, including a computer program. When the computer program is executed by a processor, the methods that may be involved in various aspects of the above embodiments are implemented.

[0155] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0157] The units described in the embodiments of the present disclosure can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least two Internet protocol addresses".

[0158] The functions described above herein can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.

[0159] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0160] It should be noted that the modification of "one" and "plural" mentioned in the present disclosure is illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".

[0161] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0162] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0163] For example, in response to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be executed will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as a terminal device, application program, server, or storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message. As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the terminal device.

[0164] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0165] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) should comply with the requirements of corresponding laws, regulations and related provisions. The data may include information, parameters, messages, etc., such as shear flow indication information.

[0166] In a first aspect, an embodiment of the present disclosure provides a video editing method, and the video editing method includes:

[0167] Obtain a first video frame image in the video to be edited and a preset cropping size;

[0168] Analyze the first video frame image according to the plug-in to determine a target position in the first video frame image, and crop the first video frame image according to the cropping size and the target position to obtain a target video frame image, where the target position is determined according to the position of the target object to be cropped in the first video frame image;

[0169] Generate a cropped video according to the target video frame image.

[0170] According to one or more embodiments of the present disclosure, the analyzing the first video frame image according to the plug-in to determine the target position in the first video frame image includes:

[0171] Determine the position of the target object in the first video frame image;

[0172] Perform smoothing processing on the position of the target object to obtain the target position in the first video frame image.

[0173] According to one or more embodiments of the present disclosure, the performing smoothing processing on the position of the target object to obtain the target position in the first video frame image includes:

[0174] Store the position of the target object in a cache, where the cache includes multiple positions of the target object in multiple video frame images;

[0175] Perform smoothing processing on the position of the target object and the multiple positions in the cache to obtain the target position of the first video frame image.

[0176] According to one or more embodiments of the present disclosure, the determining the position of the target object in the first video frame image includes:

[0177] Determine a target area of the target object in the first video frame image;

[0178] Determine the position of the target object in the first video frame image according to the target area.

[0179] According to one or more embodiments of the present disclosure, determining the position of the target object in the first video frame image according to the target area includes:

[0180] Determine the central position of the target area as the position of the target object.

[0181] According to one or more embodiments of the present disclosure, cropping the first video frame image according to the cropping size and the target position to obtain a target video frame image includes:

[0182] Obtain the target position of the first video frame image in the cache;

[0183] Taking the target position as the cropping center, crop the first video frame image according to the cropping size to obtain the target video frame image.

[0184] According to one or more embodiments of the present disclosure, the first video frame image is an image obtained by interval frame extraction. After smoothing the position of the target object to obtain the target position in the first video frame image, the method further includes:

[0185] Obtain a second video frame image between the first video frame image and the next video frame image obtained by frame extraction;

[0186] Obtain a smooth curve corresponding to the target position of the first video frame image and the target position of the next video frame image obtained by frame extraction;

[0187] Determine the target position of the second video frame image according to the smooth curve.

[0188] In a second aspect, an embodiment of the present disclosure provides a video editing device, which includes a first acquisition module, a determination module, a cropping module, and a generation module, where:

[0189] The first acquisition module is configured to acquire a first video frame image in a video to be edited and a preset cropping size;

[0190] The determination module is configured to analyze the first video frame image according to a plug-in to determine the target position in the first video frame image;

[0191] The cropping module is configured to crop the first video frame image according to the cropping size and the target position to obtain a target video frame image;

[0192] The generation module is configured to generate a cropped video according to the target video frame image.

[0193] According to one or more embodiments of the present disclosure, the determination module is specifically configured to:

[0194] Determine the position of the target object in the first video frame image;

[0195] Smooth the position of the target object to obtain the target position in the first video frame image.

[0196] According to one or more embodiments of the present disclosure, the determining module is specifically configured to:

[0197] Store the position of the target object in a cache, where the cache includes multiple positions of the target object in multiple video frame images;

[0198] Smooth the position of the target object and the multiple positions in the cache to obtain the target position of the first video frame image.

[0199] According to one or more embodiments of the present disclosure, the determining module is specifically configured to:

[0200] Determine the target area of the target object in the first video frame image;

[0201] Determine the position of the target object in the first video frame image according to the target area.

[0202] According to one or more embodiments of the present disclosure, the determining module is specifically configured to:

[0203] Determine the center position of the target area as the position of the target object.

[0204] According to one or more embodiments of the present disclosure, the determining module is specifically configured to:

[0205] Obtain the target position of the first video frame image from the cache;

[0206] Use the target position as the cropping center, and crop the first video frame image according to the cropping size to obtain the target video frame image.

[0207] According to one or more embodiments of the present disclosure, the video editing device further includes a second obtaining module, where the second obtaining module is configured to:

[0208] Obtain the second video frame image between the first video frame image and the next video frame image obtained by frame extraction;

[0209] Obtain the smooth curve corresponding to the target position of the first video frame image and the target position of the next video frame image obtained by frame extraction;

[0210] Determine the target position of the second video frame image according to the smooth curve.

[0211] In a third aspect, an embodiment of the present disclosure provides a terminal device including: a processor and a memory;

[0212] The memory stores computer-executable instructions;

[0213] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video editing method as described in the first aspect above and various possible aspects related to the first aspect.

[0214] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, and when the processor executes the computer-executable instructions, the video editing method as described in the first aspect above and various possible aspects related to the first aspect is implemented.

[0215] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the present disclosure.

[0216] In addition, although the operations are depicted in a specific order, this should not be construed as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0217] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A video editing method, characterized in that, Including: Obtain the first video frame image in the video to be edited and a preset cropping size; Analyze the first video frame image according to the plug-in to determine the target position in the first video frame image, and crop the first video frame image according to the cropping size and the target position to obtain a target video frame image, where the target position is determined according to the position of the target object to be cropped in the first video frame image; Generate a cropped video according to the target video frame image.

2. The method according to claim 1, wherein The analyzing the first video frame image according to the plug-in to determine the target position in the first video frame image includes: Determine the position of the target object in the first video frame image; Smooth the position of the target object to obtain the target position in the first video frame image.

3. The method according to claim 2, wherein The smoothing the position of the target object to obtain the target position in the first video frame image includes: Store the position of the target object in a cache, where the cache includes multiple positions of the target object in multiple video frame images; Smooth the position of the target object and the multiple positions in the cache to obtain the target position of the first video frame image.

4. The method according to claim 2 or 3, characterized in that, The determining the position of the target object in the first video frame image includes: Determine the target area of the target object in the first video frame image; Determine the position of the target object in the first video frame image according to the target area.

5. The method according to claim 4, wherein The determining the position of the target object in the first video frame image according to the target area includes: Determine the center position of the target area as the position of the target object.

6. The method according to any one of claims 3-5, characterized in that, The cropping the first video frame image according to the cropping size and the target position to obtain a target video frame image includes: Obtain the target position of the first video frame image in the cache; Use the target position as the cropping center and crop the first video frame image according to the cropping size to obtain the target video frame image.

7. The method according to any one of claims 2-6, characterized in that, The first video frame image is an image obtained by interval frame extraction. After smoothing the position of the target object to obtain the target position in the first video frame image, the method further includes: Obtain a second video frame image between the first video frame image and the next video frame image obtained by frame extraction; Obtain a smooth curve corresponding to the target position of the first video frame image and the target position of the next video frame image obtained by frame extraction; Determine the target position of the second video frame image according to the smooth curve.

8. A video editing device, characterized in that, Including a first acquisition module, a determination module, a cropping module, and a generation module, where: The first acquisition module is used to obtain the first video frame image in the video to be edited and a preset cropping size; The determination module is used to analyze the first video frame image according to the plug-in to determine the target position in the first video frame image; The cropping module is used to crop the first video frame image according to the cropping size and the target position to obtain a target video frame image; The generating module is configured to generate a cropped video based on the target video frame image.

9. A terminal device, characterized in that, It includes: a processor and a memory; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the processor executes the video editing method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the video editing method according to any one of claims 1-7 is implemented.