Video processing method and apparatus, electronic device and storage medium
By generating the first lock information and the second lock information, the problem of unstable object locking in complex video scenes is solved, and the stable display of the target object parts in the video is achieved, which improves the stability and accuracy of video processing.
Patent Information
- Application Number
- PCT/CN2024/138878
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-12-12
- Publication Date
- 2025-07-03
AI Technical Summary
In complex multi-object video scenarios, it is difficult for the prior art to stably lock specific parts of the video object, resulting in problems of lens shaking and locking errors.
By generating the first lock information, the position of the target video object is characterized, the object part is detected to obtain the second lock information, and the subject lock video is generated based on the information, so that the target object part is maintained in a fixed target area during playback.
Improves the stability and accuracy of video object locking, reduces lens shaking, and improves the visual sense of the video.
Smart Images

Figure CN2024138878_03072025_PF_FP_ABST
Abstract
Description
Video processing method, device, electronic device and storage medium
[0001] This application claims priority to Chinese patent application No. 202311813751.0 filed on December 26, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The embodiments of the present disclosure relate to a video processing method, device, electronic device, and storage medium. Background Art
[0003] For the object tracking effect of the video, it is possible to dynamically lock a specified object in the video, so that in the generated subject-locked video, the object is always displayed at a specified position in the video, such as being centered.
[0004] For example, object tracking effects are usually based on image recognition technology to track objects in the video. However, in complex multi-object video scenes, there is a problem of unstable locking of small objects, which leads to lens shaking. Summary of the Invention
[0005] The present disclosure provides a video processing method, including:
[0006] Based on a selection operation on a target video frame, first locking information is generated, and the first locking information is used to represent the position of the target video object selected by the selection operation in the target video frame; based on the first locking information, object part detection is performed to obtain second locking information, and the second locking information represents the position of the target object part of the target video object in the target video frame; based on the second locking information, a subject-locked video is generated, and during the playback process of the subject-locked video, the target object part is displayed in a fixed target area.
[0007] An embodiment of the present disclosure provides a video processing device, including:
[0008] A first processing module is configured to generate first locking information based on a selection operation on a target video frame, wherein the first locking information is used to represent a position of a target video object selected by the selection operation in the target video frame;
[0009] a second processing module, configured to perform object part detection based on the first locking information to obtain second locking information, wherein the second locking information represents a position of a target object part of the target video object in the target video frame;
[0010] The generating module is configured to generate a subject-locked video based on the second locking information, wherein the subject-locked video displays the target object part in a fixed target area during playback.
[0011] An embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0012] The memory stores computer-executable instructions;
[0013] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video processing method provided by any embodiment of the present disclosure.
[0014] An embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the video processing method provided by any embodiment of the present disclosure is implemented.
[0015] An embodiment of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, the video processing method provided by any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, a brief introduction will be given below to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] FIG1 is a diagram of an application scenario of a video processing method provided by an embodiment of the present disclosure;
[0018] FIG2 is a flow chart of a video processing method according to an embodiment of the present disclosure;
[0019] FIG3 is a flow chart of a specific implementation of step S101 in the embodiment shown in FIG2 ;
[0020] FIG4 is a flowchart of a specific implementation method of step S1013 in the embodiment shown in FIG3 ;
[0021] FIG5 is a schematic diagram of a process for determining first locking information provided by this embodiment;
[0022] FIG6 is a flowchart of a specific implementation of step S102 in the embodiment shown in FIG2 ;
[0023] FIG7 is a flowchart of a specific implementation of step S103 in the embodiment shown in FIG2 ;
[0024] FIG8 is a second flow chart of a video processing method provided by an embodiment of the present disclosure;
[0025] FIG9 is a flowchart of a specific implementation of step S202 in the embodiment shown in FIG8 ;
[0026] FIG10 is a schematic diagram of a process for generating a fitting curve provided by an embodiment of the present disclosure;
[0027] FIG11 is a structural block diagram of a video processing device provided by an embodiment of the present disclosure;
[0028] FIG12 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure; and
[0029] FIG13 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0032] The following explains the application scenarios of the embodiments of the present disclosure:
[0033] FIG1 is a diagram of an application scenario of the video processing method provided by an embodiment of the present disclosure. The video processing method provided by an embodiment of the present disclosure can be applied to an application with a video editing function. More specifically, it can be applied to an application scenario in which a target object in a video is tracked in real time or non-real time to generate a subject-locked video. The execution subject of this embodiment can be a terminal device running the above-mentioned application with a video editing function, or a server that deploys the server corresponding to the above-mentioned application, or other electronic devices that perform similar functions. Referring to FIG1 , taking a terminal device as an example, taking the scenario of non-real-time video processing as an example, after running the above-mentioned application with a video editing function, the terminal device loads the original video and uses the special effects tool provided in the application (the name of the special effects tool in FIG1 is Camera Effect #1) to dynamically crop the original video, thereby generating a camera effect with character A in the video as the main subject, i.e., generating a subject-locked video. In this subject-locked video, character A is always in the middle of the camera frame, thereby achieving a tracking and focusing effect on character A, improving the video's expressiveness.
[0034] For example, the tracking and locking of target people and target objects in a video is usually achieved based on image recognition technology. For example, a person or a person's face in a video is identified, and then tracked and the image is cropped to achieve the effect of subject locking. However, in complex multi-object video scenes, such as when a video contains multiple dynamically moving video objects (such as people), the similarity between different video objects is high, which increases the difficulty of video object tracking based on image recognition. For parts of video objects (such as people's faces), the image area is further reduced, making it more difficult to distinguish the object parts of different video objects, resulting in further increased tracking difficulty, resulting in unstable object part locking, causing lens shaking and even locking errors.
[0035] The embodiments of the present disclosure provide a video processing method to solve the above problems.
[0036] Referring to FIG2 , FIG2 is a flow chart illustrating a video processing method according to an embodiment of the present disclosure. The method of this embodiment can be applied to a terminal device, such as a smartphone, tablet computer, or personal computer, and can also be applied to other electronic devices with computing capabilities. For example, the video processing method includes:
[0037] Step S101: Based on a selection operation on a target video frame, first locking information is generated, where the first locking information is used to represent a position of a target video object selected by the selection operation in the target video frame.
[0038] For example, referring to the application scenario diagram shown in FIG1 , in a video editing scenario, after loading a target video to be processed, a terminal device can edit the target video starting from the first frame or a middle frame of the target video, thereby adding a camera effect, such as that shown in FIG1 , to the target video and thereby locking a video object in the target video. In this embodiment, the case of starting processing from the first frame of the target video is used as an example. The case of starting processing from an intermediate frame of the target video is similar and will not be further described. Specifically, when starting processing from the first frame of the target video, the first frame of the target video is the target video frame, i.e., the current video frame to be processed. First, illustratively, the target video frame includes at least two video objects, such as two people. In one possible implementation, the terminal device performs image recognition on the target video frame, identifies the at least two video objects, and displays them, for example, by marking the video objects with rectangular boxes. Subsequently, the user performs a selection operation on the terminal device to select one of the identified video objects and designate one of them as the target video object, i.e., the locked subject object.
[0039] More specifically, in one possible implementation, the user's selection operation can directly select a target object part of the target video object. For example, the selection operation involves clicking on the facial area of person A in the target video frame. The terminal device then determines the corresponding target video object and target object part based on the selection operation, such as obtaining identification information corresponding to the target video object and target object part. The terminal device then performs subsequent processing steps based on this identification information to lock the video subject. In another possible implementation, the user's selection operation includes at least two sub-operations, such as a first sub-operation and a second sub-operation. The first sub-operation only selects the target video object. The terminal device then displays the trackable parts of the target video object, such as the face, torso, or hands. The terminal device then determines the tracking target, i.e., the target object part, on the target video object by responding to the second sub-operation input by the user.
[0040] Furthermore, the terminal device generates corresponding first locking information based on the target video object determined by the selection operation. The first locking information is used to represent the position of the target video object selected by the selection operation in the target video frame. In one possible implementation method, the terminal device has detected the video objects in all target video frames in advance and obtained the corresponding position of each video. After the terminal device receives the selection operation, it directly obtains the corresponding object position, i.e., the first locking information, based on the target video object indicated by the selection operation.
[0041] In another possible implementation, the terminal device determines the first locking information by using the video frame before the target video frame. As shown in FIG3 , the specific implementation of step S101 includes:
[0042] Step S1011: Based on the selection operation, determine the target video object.
[0043] Step S1012: If the target video frame is not the first video frame of the target video, then obtain the cache object information of the target video object, where the cache object information represents the position of the target video object in the preceding video frame.
[0044] Step S1013: performing a matching detection on the video object in the target video frame according to the cached object information to generate first locking information.
[0045] For example, when the target video frame is not the first video frame of the target video, the target video frame is preceded by a preceding video frame. During the frame-by-frame processing of the target video frame, based on a similar processing method, the preceding video frame generates corresponding first locking information for indicating the position of the target video object in the preceding video frame, and the first locking information corresponding to the preceding video frame is stored as the corresponding cache object information. Subsequently, when processing the target video frame, the target video object in the target video frame is first detected to obtain the corresponding first position. Then, the second position represented by the cache object information corresponding to the preceding video frame is obtained, and the second position described by the cache object information is compared with the first position obtained by detecting the target video frame. If the two match, such as completely overlapping or partially overlapping, then the first position obtained by detecting the target video frame is accurate, and the first locking information is generated based on the first position. Otherwise, if the two do not match, such as completely separated or the separation distance is greater than a threshold, then the first position obtained by detecting the target video frame is inaccurate, and the target video frame is re-detected.
[0046] Furthermore, in a possible implementation, as shown in FIG4 , the specific implementation of step S1013 includes:
[0047] Step S1013A: Obtain a first object positioning frame according to the cached object information.
[0048] Step S1013B: Detect the target video frame and obtain a second object positioning frame corresponding to each video object.
[0049] Step S1013C: generating first locking information according to the degree of overlap between the first object positioning frame and each second object positioning frame.
[0050] FIG5 is a schematic diagram of a process for determining first lock information provided by this embodiment. The above steps are described below in conjunction with FIG5 . For example, in the target video, the target video frame is frame M, and its previous frame is the preceding video frame, i.e., frame M-1. First, by detecting the cached object information corresponding to frame M-1, a first object positioning frame for locating the target video object (shown as a portrait A in the figure) is obtained. Subsequently or prior to this, each video object in frame M is detected to obtain a second object positioning frame corresponding to each video object (shown as positioning frame F_1, positioning frame F_2, and positioning frame F_3 in the figure). Subsequently, the overlap between the first object positioning frame and each second object positioning frame is compared, and the positioning frame F_2 with the highest overlap with the first object positioning frame is determined as the first lock information.
[0051] In the steps of this embodiment, by obtaining the cache object information of the previous video frame, the video object detected in the currently processed target video frame is matched and detected, so as to obtain an accurate target video object, thereby avoiding video object detection errors caused by high similarity between different video objects and improving the stability and accuracy of subject locking.
[0052] Furthermore, the steps of this embodiment also include:
[0053] Step S1014: if the target video frame is the first video frame of the target video, then obtaining an object positioning frame of the target video object in the target video frame.
[0054] Step S1015: generating first locking information according to the object positioning frame of the target video object, and storing the first locking information as cache object information of the target video frame.
[0055] For example, if the target video frame is the first video frame of the target video, that is, the target video frame has no preceding video frame, or the preceding video frame of the target video frame has no cached object information (that is, the target video frame is the first video frame to which the subject lock special effect is applied), in this case, the cached object information cannot be used to verify the detection result of the target video frame. In this case, the object positioning frame (detection result) of the target video object is directly used to generate the first lock information. Afterwards, the first lock information is stored as the cached object information of the target video frame. In this way, when executing the subsequent video frame of the target video frame, the cached object information of the target video frame can be used to verify the object recognition result, thereby improving the accuracy of object recognition.
[0056] Step S102: performing object part detection based on the first locking information to obtain second locking information, where the second locking information represents the position of the target object part of the target video object in the target video frame.
[0057] For example, after obtaining the first locking information, it is equivalent to screening multiple video objects in the target video frame to obtain a screened target video object. After that, further identification and tracking are performed within the area indicated by the first locking information, thereby realizing the tracking of the target object part of the target video object.
[0058] In a possible implementation, as shown in FIG6 , a specific implementation of step S102 includes:
[0059] Step S1021: Based on the first locking information, intercept the target video frame to obtain a partial image corresponding to the target video frame.
[0060] Step S1022: performing key point detection on the local image to obtain key point information corresponding to the target object part, where the key point information represents the coordinates of the key points constituting the target object part.
[0061] Step S1023: Generate second locking information based on the key point information.
[0062] For example, after obtaining the first lock information, the target video frame is cropped based on the location (image area) described by the first lock information to extract the partial image indicated by the first lock information, i.e., an image containing only the target video object (e.g., person A in the target video frame). This partial image is then processed to track a specific target object portion within the target video object. Because the cropped partial image does not contain other target video objects, interference from similar object portions is reduced, thereby improving the accuracy of object portion recognition.
[0063] Specifically, key point detection is performed on a local image to obtain key point information corresponding to the target object part. The key point information represents the coordinates of the key points that constitute the target object part. In one possible implementation method, the key point information includes a coordinate set of multiple key points. The key points are the points that constitute the target object part in the image or video frame, such as the contour points of the face and the joint points of the human limbs. Through the key points, the description of the object parts such as the face, torso, and limbs can be achieved. More specifically, the key points can be the skeleton points that constitute the object part, or the points that constitute the contour of the object part. The specific implementation method of the key points can be set as needed, which will not be detailed here. Accordingly, based on the specific implementation method of the key points, the corresponding key point inspection logic is set to achieve the detection of the key points, which will not be repeated here.
[0064] Afterwards, based on the detection results of the key points, the key point information is obtained. The point set described by the key point information can be directly used as the second locking information, or after further abnormal point screening of the points described by the key point information, the second locking information representing the position of the target object part of the target video object in the target video frame can be generated.
[0065] Optionally, after performing key point detection on the local image to obtain key point information corresponding to the target object part, the method further includes:
[0066] Obtain confidence corresponding to the key point information; when the confidence is less than a confidence threshold, obtain cache location information corresponding to the previous video frame; and generate second locking information based on the cache location information.
[0067] Exemplarily, after obtaining the key point information, the credibility of the key point information is calculated to obtain a corresponding confidence level representing the accuracy and correctness of the key point information. The higher the confidence level, the more accurate the key point information constituting the target object part. Conversely, the higher the confidence level, the less accurate the key point information constituting the target object part. Specifically, the method for obtaining the confidence level corresponding to the key point information includes, for example, obtaining the confidence level by comparing the key point information with preset model information, wherein the model information refers to relevant descriptive information of the preset model corresponding to the target object part, and the model information can be obtained using the identification information obtained in the previous step (see the introduction in step S101). The model information can be used to describe the two-dimensional contour and shape of the target object part, and the key point information can be verified through the model information. When the contour or shape formed by the key points corresponding to the key point information is similar to the contour or shape of the target object part described by the model information, the key point information has a higher confidence level, that is, the coordinate set represented by the key point information can accurately describe the shape of the target object part and has a higher credibility; conversely, the key point information has a lower confidence level, that is, the coordinate set represented by the key point information cannot be significantly different from the shape of the target object part and has a lower credibility.
[0068] Furthermore, when the confidence level is less than the confidence level threshold, that is, when the confidence level is low, the cache location information corresponding to the preceding video frame is obtained, and the second locking information is generated based on the cache location information, thereby achieving error correction of the key point information. For example, the cache location information (key point information) generated by the preceding video frame (for example, the previous video frame of the target video frame) is directly used as the second locking information of the target video frame, or the cache location information generated by the preceding video frame is fused with the key point information of the target video frame (for example, the coordinate points are superimposed and averaged), and then the fusion result is used as the second locking information. Among them, the specific implementation method of the preceding video frame has been introduced in detail in the embodiment steps shown in Figure 3 and will not be repeated here.
[0069] In the steps of this embodiment, confidence evaluation is performed on the key point information, thereby achieving error correction of the key point information, thereby improving the accuracy of the generated second locking information and reducing camera shake.
[0070] Step S103: Based on the second locking information, a subject-locked video is generated. During playback, the subject-locked video displays the target object part in a fixed target area.
[0071] For example, after obtaining the second locking information, it is equivalent to obtaining the position description information of the target object part. Afterwards, based on the second locking information, the target video can be cropped with the position of the target object part as the center to obtain a cropped video frame. The cropped video frame has the effect of subject locking, that is, the target object part is located at the target position in the image, such as the middle of the picture.
[0072] For example, as shown in FIG7 , the specific implementation of step S103 includes:
[0073] Step S1031: Generate an image affine transformation matrix corresponding to the cropping frame according to the second locking information.
[0074] Step S1032: Generate key frame information according to the image affine transformation matrix.
[0075] Step S1033: Generate a locked video frame corresponding to the target video frame according to the key frame information.
[0076] Step S1034: Generate a subject-locked video according to the locked video frame corresponding to the target video frame.
[0077] Exemplarily, after obtaining the second lock information, a cropping frame of a fixed size centered on the target object portion is generated based on the position of the target object portion described by the second lock information. The fixed size corresponding to the cropping frame is consistent with the video screen size of the subject lock video. Specifically, based on the second lock information and a preset fixed size (screen length, height), a corresponding image affine transformation matrix is generated. Affine transformation refers to operations such as translation, rotation, scaling, shearing, and flipping used in image processing. The specific implementation of affine transformation is usually represented by an image affine transformation matrix. In this embodiment, the position of the target object portion represented by the second lock information is combined with a preset fixed size to obtain an image affine transformation matrix for processing the target video frame. Subsequently, the target video frame is processed based on the image affine transformation matrix to obtain a processed image, i.e., a locked video frame. In the locked video frame, the target object portion is displayed at a corresponding position, such as the center of the image.
[0078] The above steps of this embodiment are then repeated for the subsequent video frames of the target video frame until the end of the target video is reached. The subsequent video frame can be the first video frame after the target video frame or the Nth video frame after the target video frame. This generates a corresponding locked video frame for each subsequent video frame. Finally, the processed locked video frames are merged to generate a subject-locked video that displays the target object's location in a fixed target area during playback.
[0079] In this embodiment, a first lock information is generated based on a selection operation on a target video frame. The first lock information is used to represent the position of the target video object selected by the selection operation in the target video frame. Based on the first lock information, object part detection is performed to obtain second lock information. The second lock information represents the position of the target object part of the target video object in the target video frame. Based on the second lock information, a subject-locked video is generated. During playback, the subject-locked video displays the target object part in a fixed target area. By processing the target video frame in the target video in two stages, the first lock information and the second lock information are obtained respectively, and the video object and the object part are tracked sequentially. This allows for accurate tracking and locking of the target object part of the target video object in complex scenes, improving the video stability and accuracy of the final generated subject-locked video, reducing camera shake, and improving the video viewing experience.
[0080] Referring to FIG8 , FIG8 is a second flow chart of the video processing method provided by the embodiment of the present disclosure. Based on the embodiment shown in FIG2 , this embodiment further refines step S102 , and the video processing method includes:
[0081] Step S201: Based on a selection operation on a target video frame, first locking information is generated, where the first locking information is used to represent a position of a target video object selected by the selection operation in the target video frame.
[0082] Step S202: Generate first revised locking information based on the first locking information.
[0083] For example, referring to the process of generating the first locking information introduced in the embodiment shown in FIG2 , in one possible implementation method, if the target video frame is the first video frame of the target video, the target video frame is directly subjected to object detection, and the first locking information is obtained in combination with the identification information indicated by the selection operation; if the target video frame is not the first video frame of the target video, the first locking information is obtained based on the cache object information of a preceding video frame of the target video frame. The specific implementation method of the above steps has been introduced in the embodiment shown in FIG2 and will not be repeated here.
[0084] In another possible implementation, in this embodiment, when the target video frame is not the first video frame of the target video, the first locking information can be obtained through cache object information of multiple preceding video frames corresponding to the multiple preceding video frames. Specifically, as shown in FIG. 9 , the specific implementation of step S202 includes:
[0085] Step S2021: obtaining a preceding object positioning frame of a target video object in at least two preceding video frames, where a preceding video frame is a video frame in the target video that is located before the target video frame.
[0086] Step S2022: performing curve fitting based on the positions of at least two previous object positioning frames to obtain a predicted object positioning frame for the target video frame.
[0087] Step S2023: Fusing the predicted object positioning frame and the first locking information to obtain first corrected locking information.
[0088] For example, in the steps of this embodiment, first, first lock information is obtained by detecting the target video frame. This first lock information can be obtained by directly detecting the target video object in the target video frame or by generating it based on cached object information of a preceding video frame adjacent to the target video frame. The specific implementation method is not further described. Then, the preceding object positioning frames of the target video object in at least two preceding video frames are obtained. For example, the object positioning frames of N ordered preceding video frames preceding the target video frame are obtained, i.e., the first lock information of the N ordered preceding video frames preceding the target video frame is obtained. The first lock information of the N ordered preceding video frames preceding the target video frame is generated sequentially during the processing of these preceding video frames. The generation method is similar to the method for generating the first lock information of the target video frame in this embodiment and is not further described. Then, curve fitting is performed on the preceding object positioning frames of the N ordered preceding video frames to obtain a change curve representing the object positioning frame of the target video object. Then, the change curve is used to correct the first lock information of the currently detected target video frame to obtain first corrected lock information.
[0089] Figure 10 is a schematic diagram of a fitting curve generation process provided by an embodiment of the present disclosure. As shown in Figure 10, illustratively, the target video frame is frame M in the target video, and preceding video frames M-1, M-2, ..., MN are also included before the target video frame. On the one hand, the terminal device detects the target video frame and obtains the object positioning frame F_0, i.e., the first locking information corresponding to the target video frame. On the other hand, the terminal device obtains the preceding object positioning frames F_1, F_2, ..., F_N corresponding to the preceding video frames M-1, M-2, ..., MN based on the cached object information corresponding to the preceding video frames M-1, M-2, ..., MN. Afterwards, the terminal device performs curve fitting on the object positioning frames F_1, F_2...F_N. Specifically, the positioning points of each object positioning frame, such as the lower right corner point and center point of each object positioning frame (the lower right corner point is taken as an example in the figure), can be obtained, and curve fitting can be performed based on the above positioning points (the horizontal and vertical coordinates) to obtain a fitting curve L. Then, the fitting curve L is used to predict the position of the target video object in the target video frame (M frame) in the target video, that is, the position of the predicted object positioning frame, which is shown as point P1 in the figure; then, the average point of point P1 and the lower right corner point P0 of the object positioning frame F_0 is calculated to obtain a fusion point P2. Based on the fusion point P2, the object positioning frame corresponding to the target video frame is restored, that is, the corrected object positioning frame F_R. The set of coordinate points describing the corrected object positioning frame F_R is the first corrected locking information.
[0090] In this embodiment, the object positioning frame corresponding to the previous video frame is used to perform curve fitting of the object positioning frame's positioning points, thereby predicting and correcting the first lock information to obtain the first corrected lock information. By exploiting the periodicity of the target video content dimension, the object positioning frame in the current target video frame is predicted, further improving the accuracy of the first corrected lock information.
[0091] Step S203: Based on the first corrected locking information, the target video frame is intercepted to obtain a partial image corresponding to the target video frame.
[0092] Step S204: performing key point detection on the local image to obtain key point information corresponding to the target object part, where the key point information represents the coordinates of the key points constituting the target object part.
[0093] Step S205: Obtain the facial offset angle of the person in the partial image.
[0094] Step S206: Obtaining the posture offset angle of the person in the partial image.
[0095] Step S207: Determine the scene type according to the key point information and the image space area of the target video frame.
[0096] Step S208: Generate second locking information based on the key point information and the reference information, wherein the reference information includes at least one of a face offset angle, a posture offset angle, and a shot type.
[0097] For example, in a scenario where the target video object in the target video frame is a person, for the target object part based on the person, information specific to the portrait can be further obtained as a reference to improve the accuracy of the second locking information. Specifically, after performing key point detection on the partial image and obtaining key point information corresponding to the target object part, the facial offset angle of the person in the partial image, the posture offset angle of the person in the partial image, and the scene type can be further obtained as reference information. The specific implementation method of step S204 includes: obtaining the coordinates of a first key point and a second key point in the partial image, wherein the first key point is an eye positioning point and the second key point is a nose tip positioning point; obtaining the facial offset angle based on a first direction vector formed by the first key point and the second key point. The specific implementation method of step S205 includes: obtaining the coordinates of a third key point and a fourth key point in the partial image, wherein the third key point is a shoulder positioning point and the fourth key point is a facial positioning point; obtaining the posture offset angle based on a second direction vector formed by the third key point and the fourth key point.
[0098] Furthermore, shot types, such as close shot, mid shot, and long shot, are determined by the ratio of the outline area enclosed by key points represented by the key point information to the image space area of the target video frame. The image space area can be the total size of the target video frame or the size of a specific reference object (e.g., a fixed-size object such as a vehicle, road, or street sign) in the target video frame.
[0099] Afterwards, based on the key point information, the reference information is further combined to realize the screening of abnormal coordinate points in the key point information, improve the accuracy of the second locking information, solve the problems of abnormal values and small jitter, further enhance the smoothness and stability of the overall locking information, and achieve the detection accuracy of complex boundaries (such as when the portrait face is rotated), further reduce the problem of picture jitter.
[0100] Step S209: generating a subject-locked video based on the second locking information. During playback, the subject-locked video displays the target object part in a fixed target area.
[0101] In this embodiment, the implementation method of step S201 and step S209 is the same as the implementation method of step S101 and step S103 in the embodiment shown in Figure 2 of this disclosure. The specific implementation method of steps S203-S204 has been introduced in the embodiment shown in Figure 2 and will not be repeated here.
[0102] Corresponding to the video processing method of the above embodiment, FIG11 is a structural block diagram of a video processing device provided by an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Referring to FIG11, the video processing device 3 includes:
[0103] A first processing module 31 is configured to generate first locking information based on a selection operation on a target video frame, where the first locking information is used to represent a position of a target video object selected by the selection operation in the target video frame;
[0104] A second processing module 32 is configured to perform object part detection based on the first locking information to obtain second locking information, where the second locking information represents a position of a target object part of the target video object in the target video frame;
[0105] The generating module 33 is configured to generate a subject-locked video based on the second locking information. During playback of the subject-locked video, the target object portion is displayed in a fixed target area.
[0106] According to one or more embodiments of the present disclosure, the first processing module 31 is specifically used to: determine the target video object based on the selection operation; if the target video frame is not the first video frame of the target video, obtain the cache object information of the target video object, and the cache object information represents the position of the target video object in the previous video frame; perform matching detection on the video object in the target video frame according to the cache object information to generate first locking information.
[0107] According to one or more embodiments of the present disclosure, when the first processing module 31 performs match detection on the video objects in the target video frame according to the cached object information and generates the first locking information, it is specifically used to: obtain the first object positioning frame according to the cached object information; detect the target video frame to obtain the second object positioning frame corresponding to each video object; and generate the first locking information according to the degree of overlap between the first object positioning frame and each second object positioning frame.
[0108] According to one or more embodiments of the present disclosure, when the first processing module 31 generates the first locking information based on the selection operation for the target video frame, it is also used to: if the target video frame is the first video frame of the target video, obtain the object positioning frame of the target video object in the target video frame; generate the first locking information according to the object positioning frame of the target video object; and store the first locking information as the cached object information of the target video frame.
[0109] According to one or more embodiments of the present disclosure, the first processing module 31 is further used to: obtain a preceding object positioning frame of a target video object in at least two preceding video frames, where the preceding video frame is a video frame in the target video that is located before the target video frame; perform curve fitting based on the positions of the at least two preceding object positioning frames to obtain a predicted object positioning frame for the target video frame; fuse the predicted object positioning frame and the first locking information to obtain first corrected locking information; the second processing module 32 is specifically used to: perform object part detection based on the first corrected locking information to obtain second locking information.
[0110] According to one or more embodiments of the present disclosure, the second processing module 32 is specifically used to: based on the first locking information, intercept the target video frame to obtain a local image corresponding to the target video frame; perform key point detection on the local image to obtain key point information corresponding to the target object part, the key point information represents the coordinates of the key points constituting the target object part; based on the key point information, generate the second locking information.
[0111] According to one or more embodiments of the present disclosure, the second processing module 32 is further configured to: obtain coordinates of a first key point and a second key point in the local image, wherein the first key point is an eye positioning point and the second key point is a nose tip positioning point; obtain a facial offset angle based on a first direction vector formed by the first key point and the second key point;
[0112] The second processing module 32 is specifically configured to generate second locking information according to the key point information and the facial offset angle.
[0113] According to one or more embodiments of the present disclosure, the second processing module 32 is further used to: obtain the coordinates of the third key point and the fourth key point in the local image, wherein the third key point is the shoulder positioning point and the fourth key point is the face positioning point; obtain the posture offset angle according to the second direction vector formed by the third key point and the fourth key point; the second processing module 32 is specifically used to: generate second locking information according to the key point information and the posture offset angle.
[0114] According to one or more embodiments of the present disclosure, the second processing module 32 is further used to determine the shot type based on the key point information and the picture space area of the target video frame; the second processing module 32 is specifically used to generate second locking information based on the key point information and the shot type.
[0115] According to one or more embodiments of the present disclosure, the second processing module 32 is further used to: obtain the confidence corresponding to the key point information; when the second processing module 32 generates the second locking information based on the key point information, it is specifically used to: when the confidence is less than the confidence threshold, obtain the cache location information corresponding to the previous video frame; and generate the second locking information based on the cache location information.
[0116] According to one or more embodiments of the present disclosure, the generation module 33 is specifically used to: generate an image affine transformation matrix corresponding to the cropping frame according to the second locking information; generate key frame information according to the image affine transformation matrix; generate a locked video frame corresponding to the target video frame according to the key frame information; and generate a subject locked video according to the locked video frame corresponding to the target video frame.
[0117] The first processing module 31, the second processing module 32 and the generating module 33 are connected in sequence. The video processing device 3 provided in this embodiment can implement the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, which will not be described in detail in this embodiment.
[0118] FIG12 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. As shown in FIG12 , the electronic device 4 includes:
[0119] A processor 41, and a memory 42 communicatively connected to the processor 41;
[0120] Memory 42 stores computer-executable instructions;
[0121] The processor 41 executes the computer-executable instructions stored in the memory 42 to implement the video processing method in the embodiments shown in FIG. 2 to FIG. 10 .
[0122] Optionally, the processor 41 and the memory 42 are connected via a bus 43 .
[0123] The relevant explanations can be understood by referring to the relevant descriptions and effects corresponding to the steps in the embodiments corresponding to Figures 2 to 10, and no further details will be given here.
[0124] An embodiment of the present disclosure provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, they are used to implement the video processing method provided in any of the embodiments corresponding to Figures 2 to 10 of the present disclosure.
[0125] An embodiment of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, the video processing method provided in any one of the embodiments corresponding to Figures 2 to 10 of the present disclosure is implemented.
[0126] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device.
[0127] Referring to FIG13 , a schematic diagram of the structure of an electronic device 900 suitable for implementing an embodiment of the present disclosure is shown. The electronic device 900 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG13 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0128] As shown in FIG13 , the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0129] Typically, the following devices can be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 can allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although FIG13 shows an electronic device 900 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0130] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0131] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0132] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0133] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0134] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0136] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0137] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0138] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0139] According to one or more embodiments of the present disclosure, a video processing method is provided, including:
[0140] Based on a selection operation on a target video frame, first locking information is generated, and the first locking information is used to represent the position of the target video object selected by the selection operation in the target video frame; based on the first locking information, object part detection is performed to obtain second locking information, and the second locking information represents the position of the target object part of the target video object in the target video frame; based on the second locking information, a subject-locked video is generated, and during the playback process of the subject-locked video, the target object part is displayed in a fixed target area.
[0141] According to one or more embodiments of the present disclosure, the first locking information is generated based on the selection operation on the target video frame, including: determining the target video object based on the selection operation; if the target video frame is not the first video frame of the target video, obtaining the cache object information of the target video object, the cache object information representing the position of the target video object in the previous video frame; performing matching detection on the video object in the target video frame according to the cache object information to generate the first locking information.
[0142] According to one or more embodiments of the present disclosure, performing match detection on the video objects in the target video frame based on the cached object information to generate first locking information includes: obtaining a first object positioning frame based on the cached object information; detecting the target video frame to obtain a second object positioning frame corresponding to each of the video objects; and generating the first locking information based on the degree of overlap between the first object positioning frame and each of the second object positioning frames.
[0143] According to one or more embodiments of the present disclosure, the generation of the first locking information based on the selection operation on the target video frame also includes: if the target video frame is the first video frame of the target video, obtaining the object positioning frame of the target video object in the target video frame; generating the first locking information based on the object positioning frame of the target video object; and storing the first locking information as the cached object information of the target video frame.
[0144] According to one or more embodiments of the present disclosure, the method further includes: obtaining a preceding object positioning frame of a target video object in at least two preceding video frames, the preceding video frame being a video frame in the target video that is located before the target video frame; performing curve fitting based on the positions of at least two preceding object positioning frames to obtain a predicted object positioning frame for the target video frame; fusing the predicted object positioning frame and the first locking information to obtain first corrected locking information; performing object part detection based on the first locking information to obtain second locking information, including: performing object part detection based on the first corrected locking information to obtain second locking information.
[0145] According to one or more embodiments of the present disclosure, the object part detection is performed based on the first locking information to obtain the second locking information, including: based on the first locking information, the target video frame is intercepted to obtain a local image corresponding to the target video frame; key point detection is performed on the local image to obtain key point information corresponding to the target object part, the key point information represents the coordinates of the key points constituting the target object part; based on the key point information, the second locking information is generated.
[0146] According to one or more embodiments of the present disclosure, the method further includes: obtaining the coordinates of a first key point and a second key point in the local image, wherein the first key point is an eye positioning point and the second key point is a nose tip positioning point; obtaining a facial offset angle based on a first direction vector formed by the first key point and the second key point; generating the second locking information based on the key point information includes: generating the second locking information based on the key point information and the facial offset angle.
[0147] According to one or more embodiments of the present disclosure, the method further includes: obtaining the coordinates of a third key point and a fourth key point in the local image, wherein the third key point is a shoulder positioning point and the fourth key point is a face positioning point; obtaining a posture offset angle based on a second direction vector formed by the third key point and the fourth key point; generating the second locking information based on the key point information includes: generating the second locking information based on the key point information and the posture offset angle.
[0148] According to one or more embodiments of the present disclosure, the method further includes: determining the shot type based on the key point information and the picture space area of the target video frame; generating the second locking information based on the key point information, including: generating the second locking information based on the key point information and the shot type.
[0149] According to one or more embodiments of the present disclosure, the method further includes: obtaining the confidence corresponding to the key point information; generating the second locking information based on the key point information, including: when the confidence is less than a confidence threshold, obtaining the cache location information corresponding to the previous video frame; generating the second locking information based on the cache location information.
[0150] According to one or more embodiments of the present disclosure, a subject-locked video is generated based on the second locking information, including: generating an image affine transformation matrix corresponding to the cropping frame according to the second locking information; generating key frame information according to the image affine transformation matrix; generating a locked video frame corresponding to the target video frame according to the key frame information; and generating the subject-locked video according to the locked video frame corresponding to the target video frame.
[0151] According to one or more embodiments of the present disclosure, there is provided a video processing apparatus, including:
[0152] A first processing module is configured to generate first locking information based on a selection operation on a target video frame, wherein the first locking information is used to represent a position of a target video object selected by the selection operation in the target video frame;
[0153] a second processing module, configured to perform object part detection based on the first locking information to obtain second locking information, wherein the second locking information represents a position of a target object part of the target video object in the target video frame;
[0154] The generating module is configured to generate a subject-locked video based on the second locking information, wherein the subject-locked video displays the target object part in a fixed target area during playback.
[0155] According to one or more embodiments of the present disclosure, the first processing module is specifically used to: determine the target video object based on the selection operation; if the target video frame is not the first video frame of the target video, obtain the cache object information of the target video object, and the cache object information represents the position of the target video object in the previous video frame; perform matching detection on the video object in the target video frame according to the cache object information to generate first locking information.
[0156] According to one or more embodiments of the present disclosure, when the first processing module performs match detection on the video objects in the target video frame according to the cached object information and generates the first locking information, it is specifically used to: obtain a first object positioning frame according to the cached object information; detect the target video frame to obtain a second object positioning frame corresponding to each of the video objects; and generate the first locking information according to the degree of overlap between the first object positioning frame and each of the second object positioning frames.
[0157] According to one or more embodiments of the present disclosure, when the first processing module generates the first locking information based on the selection operation for the target video frame, it is also used to: if the target video frame is the first video frame of the target video, obtain the object positioning frame of the target video object in the target video frame; generate the first locking information according to the object positioning frame of the target video object; and store the first locking information as the cached object information of the target video frame.
[0158] According to one or more embodiments of the present disclosure, the first processing module is further used to: obtain a preceding object positioning frame of a target video object in at least two preceding video frames, where the preceding video frame is a video frame in the target video that is located before the target video frame; perform curve fitting based on the positions of at least two preceding object positioning frames to obtain a predicted object positioning frame for the target video frame; fuse the predicted object positioning frame with the first locking information to obtain first corrected locking information; the second processing module is specifically used to: perform object part detection based on the first corrected locking information to obtain second locking information.
[0159] According to one or more embodiments of the present disclosure, the second processing module is specifically used to: based on the first locking information, intercept the target video frame to obtain a local image corresponding to the target video frame; perform key point detection on the local image to obtain key point information corresponding to the target object part, the key point information represents the coordinates of the key points constituting the target object part; based on the key point information, generate the second locking information.
[0160] According to one or more embodiments of the present disclosure, the second processing module is further configured to: obtain coordinates of a first key point and a second key point in the partial image, wherein the first key point is an eye positioning point and the second key point is a nose tip positioning point; and obtain a facial offset angle based on a first direction vector formed by the first key point and the second key point;
[0161] The second processing module is specifically configured to generate the second locking information according to the key point information and the facial offset angle.
[0162] According to one or more embodiments of the present disclosure, the second processing module is further used to: obtain the coordinates of a third key point and a fourth key point in the local image, wherein the third key point is a shoulder positioning point and the fourth key point is a facial positioning point; obtain a posture offset angle based on a second direction vector formed by the third key point and the fourth key point; the second processing module is specifically used to: generate the second locking information based on the key point information and the posture offset angle.
[0163] According to one or more embodiments of the present disclosure, the second processing module is further used to: determine the shot type based on the key point information and the picture space area of the target video frame; the second processing module is specifically used to: generate the second locking information based on the key point information and the shot type.
[0164] According to one or more embodiments of the present disclosure, the second processing module is further used to: obtain the confidence corresponding to the key point information; when the second processing module generates the second locking information based on the key point information, it is specifically used to: when the confidence is less than the confidence threshold, obtain the cache location information corresponding to the previous video frame; and generate the second locking information based on the cache location information.
[0165] According to one or more embodiments of the present disclosure, the generation module is specifically used to: generate an image affine transformation matrix corresponding to the cropping frame based on the second locking information; generate key frame information based on the image affine transformation matrix; generate a locked video frame corresponding to the target video frame based on the key frame information; and generate the subject locked video based on the locked video frame corresponding to the target video frame.
[0166] According to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one processor and a memory;
[0167] The memory stores computer-executable instructions;
[0168] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the video processing method provided by any embodiment of the present disclosure.
[0169] According to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the video processing method provided by any embodiment of the present disclosure is implemented.
[0170] According to one or more embodiments of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the video processing method provided by any embodiment of the present disclosure is implemented.
[0171] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0172] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0173] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A video processing method, comprising: Generating first locking information based on a selection operation for a target video frame, where the first locking information is used to characterize the position of a target video object selected by the selection operation in the target video frame; Performing object part detection based on the first locking information to obtain second locking information, where the second locking information characterizes the position of a target object part of the target video object in the target video frame; and Generating a main body locked video based on the second locking information, where during the playback of the main body locked video, the target object part is displayed in a fixed target area.
2. The method according to claim 1, wherein, The generating first locking information based on a selection operation for a target video frame includes: Determining the target video object based on the selection operation; In response to the target video frame not being the first video frame of the target video, obtaining cached object information of the target video object, where the cached object information characterizes the position of the target video object in a previous video frame; and Performing matching detection on the video objects in the target video frame according to the cached object information to generate the first locking information.
3. The method according to claim 2, wherein, The performing matching detection on the video objects in the target video frame according to the cached object information to generate the first locking information includes: Obtaining a first object positioning frame according to the cached object information; Detecting the target video frame to obtain second object positioning frames corresponding to the respective video objects; and Generating the first locking information according to the coincidence degree between the first object positioning frame and each of the second object positioning frames.
4. The method according to claim 2 or 3, wherein The generating first locking information based on a selection operation for a target video frame further includes: In response to the target video frame being the first video frame of the target video, obtaining an object positioning frame of the target video object in the target video frame; Generating the first locking information according to the object positioning frame of the target video object; and Storing the first locking information as the cached object information of the target video frame.
5. The method according to any one of claims 1-4, further comprising: Obtaining previous object positioning frames of the target video object in at least two previous video frames, where the previous video frames are video frames in the target video that are before the target video frame; Performing curve fitting based on the positions of at least two of the previous object positioning frames to obtain a predicted object positioning frame for the target video frame; Fusing the predicted object positioning frame and the first locking information to obtain first corrected locking information; and The performing object part detection based on the first locking information to obtain second locking information includes: Performing the object part detection based on the first corrected locking information to obtain the second locking information.
6. The method according to any one of claims 1-4, wherein, The performing object part detection based on the first locking information to obtain second locking information includes: Based on the first locking information, intercepting the target video frame to obtain a partial image corresponding to the target video frame; Perform key point detection on the local image to obtain key point information corresponding to the target object part, where the key point information represents the coordinates of the key points constituting the target object part; and Generate the second locking information based on the key point information.
7. The method according to claim 6, further comprising: Obtain the coordinates of a first key point and a second key point in the local image, where the first key point is an eye positioning point and the second key point is a nose tip positioning point; Obtain a face offset angle according to a first direction vector formed by the first key point and the second key point; and The generating the second locking information based on the key point information includes: Generate the second locking information according to the key point information and the face offset angle.
8. The method according to claim 6, further comprising: Obtain the coordinates of a third key point and a fourth key point in the local image, where the third key point is a shoulder positioning point and the fourth key point is a face positioning point; Obtain a posture offset angle according to a second direction vector formed by the third key point and the fourth key point; and The generating the second locking information based on the key point information includes: Generate the second locking information according to the key point information and the posture offset angle.
9. The method according to claim 6, further comprising: Determine a scene type according to the key point information and the screen space area of the target video frame; and The generating the second locking information based on the key point information includes: Generate the second locking information according to the key point information and the scene type.
10. The method according to claim 6, further comprising: Obtain the confidence corresponding to the key point information; and The generating the second locking information based on the key point information includes: When the confidence is less than a confidence threshold, obtain cached part information corresponding to a previous video frame; and Generate the second locking information based on the cached part information.
11. According to the method according to any one of claims 1-10, wherein The generating the main body locked video based on the second locking information includes: Generate an image affine transformation matrix corresponding to a cropping frame according to the second locking information; Generate key frame information according to the image affine transformation matrix; Generate a locked video frame corresponding to the target video frame according to the key frame information; and Generate the main body locked video according to the locked video frame corresponding to the target video frame.
12. A video processing device, comprising: A first processing module configured to generate first locking information based on a selection operation for a target video frame, where the first locking information is used to represent the position of a target video object selected by the selection operation in the target video frame; A second processing module configured to perform object part detection based on the first locking information to obtain second locking information, where the second locking information represents the position of a target object part of the target video object in the target video frame; and A generation module, configured to generate a subject-locked video based on the second locking information, wherein, during the playback of the subject-locked video, the target object part is displayed in a fixed target area.
13. An electronic device, comprising: A processor and a memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the video processing method according to any one of claims 1 to 11.
14. A computer-readable storage medium, wherein, Computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the video processing method according to any one of claims 1 to 11 is implemented.
15. A computer program product, comprising a computer program, wherein, When the computer program is executed by the processor, the video processing method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Target tracking and displaying method and device in panoramic video
CN105843541A
Tracking object of interest in an omnidirectional video
CN108292364A
Tracking method and device, equipment and storage medium
CN114463373A
Live video character tracking method and device, equipment and storage medium
CN114466218A
Video processing method and device, electronic equipment and storage medium
CN117956235A