Human body key point-based video frame alignment method and device, equipment and medium

By acquiring user videos and special effects resource packages, extracting key points of the human body, calculating the angles and pose distances of human joints, and constructing path curves, the problems of large computational load and low accuracy in existing technologies are solved, achieving high-precision video frame alignment.

CN116016810BActive Publication Date: 2025-12-19HANGZHOU YUNXIANG TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211599101.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-12
Publication Date
2025-12-19
Estimated Expiration
2042-12-12

AI Technical Summary

Technical Problem

Existing methods for aligning video frames between special effects videos and user videos involve large computational loads and the recognition process can only resolve the trigger logic, resulting in poor accuracy in aligning continuous actions.

Method used

By acquiring user videos and special effects resource packages, key points of the human body are extracted, the angles and pose distances of human joints are calculated, path curves are constructed, and mapping processing is performed to achieve video frame alignment.

Benefits of technology

It improves the accuracy of video frame alignment between special effects videos and user videos, reduces computational load, and improves the accuracy of continuous alignment actions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116016810B_ABST
    Figure CN116016810B_ABST
Patent Text Reader

Abstract

The application relates to a human body key point-based video frame alignment method and device, equipment and medium, wherein the method comprises the following steps: acquiring a user video and a special effect resource package; extracting human body key points in the user video, and calculating human body joint angles based on the human body key points; calculating human body posture distances of video frames in the user video and a model video based on the human body joint angles, model joint angles, model features and the human body key points; constructing path curves corresponding to the user video and the model video based on the human body posture distances; and performing mapping processing on the user video and a special effect video according to the path curves, so that the frames of the user video and the special effect video are aligned. The application avoids direct frame alignment of the user video and the special effect video, and is beneficial to improving the video frame alignment accuracy of the special effect video and the user video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video image processing, and in particular to a video frame alignment method and device based on human body key points, equipment and medium. BACKGROUND

[0002] With the development of mobile Internet and artificial intelligence, short videos have been widely used, and creative play based on video content is more and more complex. One of the play is that professional designers design special effect videos and scenes, and users cooperate with the special effect video for shooting according to the instructions, and finally output a short video with a movie special effect. In these scenes, some special effect contents to be added and human body actions have a corresponding relationship in time. For example, similar to the special effect of launching a shock wave in an animation, the special effect and the action are continuous in time, and the special effect content is also continuously changed with the continuous change of the action, and the special effect content includes not only the light wave on the hand, but also the scene special effect. In this case, one way is to cooperate with the special effect video by the character, and the trouble of this method is that the time point is not easy to control, and the effect of the shooting is distorted, and ordinary users are difficult to do well; another way is to use video editing software, and adjust the special effect video by using the curve speed to superimpose it into the target period of the shooting video. The trouble of this method is that the operation is time-consuming, and the operation on the mobile phone is more troublesome, which is not consistent with the use scene of the mobile phone. Therefore, when making a special effect video, the video frame alignment of the special effect video and the user video is a problem to be solved urgently.

[0003] The existing video frame alignment method of the special effect video and the user video is based on the coordinates of the human body key points, uses deep learning to extract features, and then performs time sequence action recognition. However, the calculation amount of the deep learning model itself is relatively large, and the frame number of the video is also relatively large, which leads to large calculation pressure, and the recognition process can only solve the triggering logic, which leads to poor alignment accuracy of continuous actions. Therefore, there is an urgent need for a method capable of improving the video frame alignment accuracy of the special effect video and the user video. SUMMARY

[0004] The purpose of the embodiments of the present application is to provide a video frame alignment method, device and equipment based on human body key points, and medium, to improve the video frame alignment accuracy of the special effect video and the user video.

[0005] In order to solve the above technical problems, the embodiments of the present application provide a video frame alignment method based on human body key points, comprising:

[0006] Obtaining a user video and a special effect resource package, wherein the special effect resource package includes a model video, a special effect video, a model feature and a model joint angle;

[0007] extract a human key point in the user video, and calculate a human joint angle based on the human key point;

[0008] calculate a human posture distance of video frames in the user video and the model video based on the human joint angle, the model joint angle, the model feature, and the human key point;

[0009] construct a path curve corresponding to the user video and the model video based on the human posture distance;

[0010] map the user video and the special effect video according to the path curve, so that frames of the user video and the special effect video are aligned.

[0011] To solve the above technical problems, an embodiment of the present application provides a video frame alignment device based on a human key point, comprising:

[0012] a user video acquisition module configured to acquire a user video and a special effect resource package, wherein the special effect resource package comprises a model video, a special effect video, a model feature, and a model joint angle;

[0013] a human joint angle calculation module configured to extract a human key point in the user video, and calculate a human joint angle based on the human key point;

[0014] a human posture distance calculation module configured to calculate a human posture distance of video frames in the user video and the model video based on the human joint angle, the model joint angle, the model feature, and the human key point;

[0015] a path curve construction module configured to construct a path curve corresponding to the user video and the model video based on the human posture distance;

[0016] a video frame alignment module configured to map the user video and the special effect video according to the path curve, so that frames of the user video and the special effect video are aligned.

[0017] To solve the above technical problems, one technical solution adopted by the present application is to provide a computer device, comprising one or more processors, and a memory configured to store one or more programs, so that the one or more processors implement the video frame alignment method based on a human key point described in any of the above.

[0018] To solve the above technical problems, one technical solution adopted by the present application is a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the video frame alignment method based on a human key point described in any of the above.

[0019] The embodiment of the present application provides a video frame alignment method based on human body key points, a device, equipment and a medium. The method comprises the following steps: obtaining a user video and a special effect resource package, wherein the special effect resource package comprises a model video, a special effect video, model features and model joint angles; extracting human body key points in the user video, and calculating human body joint angles based on the human body key points; calculating human body posture distances of video frames in the user video and the model video based on the human body joint angles, the model joint angles, the model features and the human body key points; constructing path curves corresponding to the user video and the model video based on the human body posture distances; and performing mapping processing on the user video and the special effect video according to the path curves, so that the frames of the user video and the special effect video are aligned. The embodiment of the present application calculates the human body posture distances of the user video and the model video, constructs path curves based on the human body posture distances, and finally aligns the video frames of the user video and the special effect video according to the path curves, thereby avoiding direct frame alignment of the user video and the special effect video, and facilitating improvement of the video frame alignment accuracy of the special effect video and the user video. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the scheme in the present application, the drawings required in the embodiment description of the present application will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0021] Figure 1 is an implementation flowchart of the video frame alignment method based on human body key points provided by the embodiment of the present application;

[0022] Figure 2 is an implementation flowchart of a sub-process of the video frame alignment method based on human body key points provided by the embodiment of the present application;

[0023] Figure 3 is another implementation flowchart of a sub-process of the video frame alignment method based on human body key points provided by the embodiment of the present application;

[0024] Figure 4 is a schematic diagram of target joint points of a human body in a user video provided by the embodiment of the present application;

[0025] Figure 5 is another implementation flowchart of a sub-process of the video frame alignment method based on human body key points provided by the embodiment of the present application;

[0026] Figure 6 is another implementation flowchart of a sub-process of the video frame alignment method based on human body key points provided by the embodiment of the present application;

[0027] Figure 7 is another implementation flowchart of the sub-process in the video frame alignment method based on human key points provided by the embodiments of the present application;

[0028] Figure 8 is another implementation flowchart of the sub-process in the video frame alignment method based on human key points provided by the embodiments of the present application;

[0029] Figure 9 is a schematic diagram of the video frame alignment device provided by the embodiments of the present application;

[0030] Figure 10 is a schematic diagram of the computer device provided by the embodiments of the present application. DETAILED DESCRIPTION

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application; the description and the drawings are to be regarded as illustrative in nature and are not intended to limit the application; the terminology used in the description and the claims of the present application and the above description of the drawings includes the terms specifically mentioned above as well as any equivalents thereof.

[0032] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. It is expressly understood that the embodiments described herein are merely examples from a much larger number of embodiments that can be claimed.

[0033] In order to enable persons skilled in the art to better understand the schemes of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings.

[0034] The present application will be described in detail below in conjunction with the drawings and embodiments.

[0035] It should be noted that the video frame alignment method based on human key points provided by the embodiments of the present application is generally executed by a server, and accordingly, the video frame alignment device based on human key points is generally configured in the server.

[0036] Please refer to Figure 1 , Figure 1 shows a specific embodiment of the video frame alignment method based on human key points.

[0037] It should be noted that the method of the present application is not limited to the order of the flow shown, and includes the following steps: Figure 1

[0038] S1: Obtain a user video and a special effect resource package, wherein the special effect resource package includes a model video, a special effect video, model features, and model joint angles.

[0039] Specifically, when creating a special effect video, the creator will first shoot a video based on a model (referred to as a model video), and then create a special effect based on this video (referred to as a special effect video). The two videos are superimposed to obtain the target video. The user shoots based on the special effect video (referred to as the user video), and the final effect video is obtained after superimposing the user video and the special effect video. Because the model video and the special effect video are matched, the embodiment of the present application first aligns the user video and the model video, and then aligns the special effect video and the user video based on the curve obtained by alignment, thereby obtaining the final effect video.

[0040] Further, the template video can have only one version, or multiple versions can be shot, and the multiple versions of the model video will have more robust effects in subsequent alignment. The multiple versions refer to the same model or different models shooting the same action sequence. By extracting the features of all feature videos, a model video is selected as a template video for making a special effect, and other model videos are aligned with the template video. The features of the aligned template video and other model videos corresponding to the other model videos are used as model features. The model joint angles refer to the joint angles of the aligned template video and other model videos corresponding to the other model videos.

[0041] S2: Extract human key points in the user video, and calculate human joint angles based on the human key points.

[0042] Specifically, the human key points refer to the joint key points in the human posture. The human joint angles refer to the angles corresponding to the joint points in the human posture.

[0043] Please refer to Figure 2 , Figure 2 A specific implementation of step S2 is shown as follows:

[0044] S21: Extract human key points in the user video based on a human skeleton point detection method.

[0045] Specifically, different human skeleton point detection methods are implemented by different detectors, and the detectors include detectors for detecting two-dimensional points or three-dimensional points. The embodiment of the present application can implement video frame alignment based on two-dimensional points or three-dimensional points. ​

[0046] S22: Select a preset joint of the human body as a target joint, and generate a joint vector corresponding to the target joint.

[0047] Please refer to Figure 3 and Figure 4 , Figure 3 shows a specific embodiment of step S22, Figure 4 is a schematic diagram of a target joint of a human body in a user video provided by the embodiment of the application, which is described in detail as follows:

[0048] S221: Select a preset joint of the human body as a target joint, wherein the target joint includes the right elbow, the left elbow, the left shoulder, the right shoulder, the left hip, the right hip, the left knee, and the right knee.

[0049] S222: For any target joint, obtain a human body key point corresponding to the target joint as an origin key point.

[0050] S223: Obtain two human body key points closest to the origin key point as basic key points.

[0051] S224: Subtract the coordinates of the two basic key points from the coordinates of the origin key point to obtain a joint vector.

[0052] Specifically, the preset joint includes the right elbow, the left elbow, the left shoulder, the right shoulder, the left hip, the right hip, the left knee, and the right knee. Since each target joint detects a human body key point, the human body key point can be represented by a coordinate point. For any target joint, obtain a human body key point corresponding to the target joint as an origin key point, obtain two human body key points closest to the origin key point as basic key points; then subtract the coordinates of the two basic key points from the coordinates of the origin key point to obtain a joint vector, until all target joints are calculated, thereby obtaining all joint vectors. As shown in Figure 4 the left shoulder as a target joint, the joints closest to the left shoulder are the left elbow and the left shoulder, so the left shoulder, the left elbow, and the key points corresponding to the left shoulder are obtained, the left elbow coordinate point is subtracted from the left shoulder coordinate point to obtain a vector v, and the right shoulder coordinate point is subtracted from the left shoulder coordinate point to obtain a vector w, so the vector v and the vector w are the joint vectors of the left shoulder. In the same way, each target joint is calculated to obtain the joint vectors corresponding to all target joints. It should be noted that the joint vector here is a column vector.

[0053] S23: Based on the joint vector, calculate the human body joint angle corresponding to the target joint by a first preset formula.

[0054] Specifically, since the joint vectors corresponding to all target joints have been obtained through the above steps, the human joint angle corresponding to the target joint is calculated through a first preset formula.

[0055] The first preset formula is as follows: Wherein, a is the human joint angle, v and w are the joint vectors.

[0056] S3: Based on the human joint angle, the model joint angle, the model feature, and the human key point, the human posture distance of the video frames in the user video and the model video is calculated.

[0057] Please refer to Figure 5 , Figure 5 An embodiment of step S3 is shown as follows:

[0058] S31: The angle difference between the human joint angle and the model joint angle is calculated through a second preset formula to obtain a target angle difference.

[0059] S32: Based on the target angle difference, the model feature, and the human key point, the Chebyshev distance of the video frames in the user video and the model video is calculated to obtain the human posture distance.

[0060] The second preset formula is as follows: Wherein, a is the human joint angle, b is the model joint angle, d a (a, b) is the target angle difference.

[0061] Specifically, the angle difference between the human joint angle and the model joint angle is calculated through the second preset formula to obtain the target angle difference; then the Chebyshev distance is used as the distance measurement of two human postures, so as to calculate the Chebyshev distance of the video frames in the user video and the model video, and obtain the human posture distance. The specific calculation formula is as follows: F (x) , F (y) represents the extracted human key point and the target angle difference, F (x) [i] represents one of the human key points and the target angle difference.

[0062] Please refer to Figure 6 , Figure 6 An embodiment before step S3 is shown as follows:

[0063] S3A: The human key points corresponding to the same frames in the user video and the multiple model videos are obtained to obtain the user same frame key points and the multiple model same frame key points.

[0064] S3B: Calculate the human posture distance between each of the plurality of model same-frame key points and the user same-frame key points to obtain a posture distance set.

[0065] S3C: Take the posture distance with the shortest distance in the posture distance set as a target distance, and take the posture distances in the posture distance set other than the target distance as to-be-deleted distances.

[0066] S3D: Delete the model video corresponding to the to-be-deleted distance.

[0067] Specifically, since there can be multiple model videos, if the user video is compared with each of the model videos, the calculation amount is large, so the embodiment of the present application filters out the model videos similar to the user video, and excludes the redundant model videos, thereby reducing the subsequent calculation amount. First, the user same-frame key points and the plurality of model same-frame key points are obtained by acquiring the human key points corresponding to the same frames in the user video and the plurality of model videos, then the human posture distance between each of the plurality of model same-frame key points and the user same-frame key points is calculated to obtain a posture distance set, and finally the posture distance with the shortest distance in the posture distance set is taken as a target distance, and the posture distances in the posture distance set other than the target distance are taken as to-be-deleted distances. The human posture distance calculation formula is: wherein F (x) and F are the user same-frame key points and the model same-frame key points, respectively.

[0068] S4: Based on the human posture distance, a path curve corresponding to the user video and the model video is constructed.

[0069] Please refer to Figure 7 , Figure 7 for a specific implementation of step S4, which is described in detail as follows:

[0070] S41: A two-dimensional grid is constructed according to the video frame numbers of the user video and the model video.

[0071] S42: Based on the human posture distance, a distance array of the two-dimensional grid is constructed.

[0072] S43: Based on the distance data, elements with minimum values are sequentially obtained from the two-dimensional grid, and based on the elements with minimum values, a path curve is constructed.

[0073] Specifically, assuming that the frame number of the input video is N u , the frame number of the model feature is N v , a two-dimensional grid N u × N v is constructed, the left upper corner coordinate in the two-dimensional grid is (1, 1), and the right lower corner coordinate is (N u , Nv ). The goal of the embodiments of the present application is to find a monotone increasing path from the top-left corner to the bottom-right corner, such that the cost of this path is the minimum among all possible paths. With this path, it is easy to get the function that aligns the frames of the user video to the frames of the model video. Then define the distance array Element D[x, y] defines the distance between the x-th frame of the user video and the y-th frame of the model feature. All arrays are initialized to -1. When reading D[x, y], if its value is -1, then calculate the distance between the x-th frame of the user video and the y-th frame of the model feature, otherwise take the value directly. In a specific embodiment of constructing the path curve, 1: initialize (N u , N v ), p u ← N u , p v ← N v (P is initialized to an empty queue); 2. If p u = 1 and p u = 1, go to step 11, otherwise next step; 3. Initialize an empty set S; 4. If p u > 1, add the pair (D[p u - 1, p v ], (-1, 0)) to S; 5. If p v > 1, add the pair (D[p u , p v - 1], (0, -1)) to S; 6. If p u > 1 and p v > 1, add the pair (D[p u - 1, p v - 1], (-1, -1)) to S; 7. Take the minimum value of the first member of each element in set S as the value of the element, take out the element e, and perform the following steps: 8. Add e[2], the second member of e, as the first element of P; 9. e[2] itself is also a pair, i.e. e[2], update p u ← p u + δ u , p v ← p v + δ v ; 10. Go to step 2; 11. P is the optimal path (path curve).

[0074] S5: According to the path curve, map the user video and the special effect video to make the frames of the user video and the special effect video aligned.

[0075] Specifically, a mapping relationship between the user video and the special effect video is constructed, such as constructing a mapping relationship of 1:N according to the path curve P. u → 1:N v Finally, the frame alignment of the user video and the special effect video is performed based on the mapping relationship.

[0076] Please refer to Figure 8 , Figure 8 An embodiment before step S1 is shown as follows:

[0077] S1A: Obtain a plurality of model videos, and perform feature extraction on the plurality of model videos to obtain initial features.

[0078] S1B: Calculate model joint node angles based on the initial features.

[0079] S1C: Determine a template video from the plurality of model videos, and perform frame alignment processing on the template video and the plurality of model videos, and obtain features in the frame-aligned template video and the plurality of model videos as model features.

[0080] S1D: Package the model features, the model joint node angles, and the plurality of model videos to obtain a special effect resource package.

[0081] Specifically, the embodiment of the present application first constructs a special effect resource package, and when a user needs to perform special effect processing later, the special effect resource package can be directly downloaded and used. In this embodiment, a plurality of model videos are obtained, and features are extracted from the plurality of model videos to obtain initial features. Based on the initial features, model joint node angles are calculated. A template video is determined from the plurality of model videos, and frame alignment processing is performed on the template video and the plurality of model videos, and features in the frame-aligned template video and the plurality of model videos are obtained as model features. It should be noted that the feature extraction method of this embodiment adopts the same method as steps S2, the calculation method of the model joint node angles adopts the same method as steps S22-S23, and the frame alignment processing method adopts the same method as steps S3-S5. To avoid repetition, this will not be repeated here.

[0082] In the embodiment, a user video and an effect resource package are acquired, wherein the effect resource package includes a model video, an effect video, model features, and model joint angle; human body key points in the user video are extracted, and human body joint angles are calculated based on the human body key points; human body posture distances of video frames in the user video and the model video are calculated based on the human body joint angles, the model joint angles, the model features, and the human body key points; path curves corresponding to the user video and the model video are constructed based on the human body posture distances; and the user video and the effect video are mapped according to the path curves, so that the frames of the user video and the effect video are aligned. The human body posture distances of the user video and the model video are calculated, and the path curves are constructed based on the human body posture distances, and finally the video frames of the user video and the effect video are aligned according to the path curves, so that direct frame alignment of the user video and the effect video is avoided, and the video frame alignment accuracy of the effect video and the user video is improved.

[0083] Please refer to Figure 9 , as an implementation of the method shown in Figure 1 , the application provides an embodiment of a video frame alignment device based on human body key points, which corresponds to the method embodiment shown in Figure 1 , and the device can be applied to various electronic devices.

[0084] As shown in Figure 9 , the video frame alignment device based on human body key points in the embodiment includes a user video acquisition module 61, a human body joint angle calculation module 62, a human body posture distance calculation module 63, a path curve construction module 64, and a video frame alignment module 65, wherein:

[0085] The user video acquisition module 61 is configured to acquire a user video and an effect resource package, wherein the effect resource package includes a model video, an effect video, model features, and model joint angles.

[0086] The human body joint angle calculation module 62 is configured to extract human body key points in the user video, and calculate human body joint angles based on the human body key points.

[0087] The human body posture distance calculation module 63 is configured to calculate human body posture distances of video frames in the user video and the model video based on the human body joint angles, the model joint angles, the model features, and the human body key points.

[0088] The path curve construction module 64 is configured to construct path curves corresponding to the user video and the model video based on the human body posture distances.

[0089] The video frame alignment module 65 is configured to map the user video and the effect video according to the path curves, so that the frames of the user video and the effect video are aligned.

[0090] Further, the user video acquisition module 61 further comprises:

[0091] An initial feature extraction module, configured to acquire a plurality of model videos, and perform feature extraction on the plurality of model videos to obtain initial features;

[0092] A model joint angle calculation module, configured to calculate model joint angles based on the initial features;

[0093] A model feature generation module, configured to determine a template video from the plurality of model videos, and perform frame alignment processing on the template video and the plurality of model videos, and acquire features of the frame-aligned template video and the plurality of model videos as model features;

[0094] An effect resource package generation module, configured to package the model features, the model joint angles, and the plurality of model videos to obtain an effect resource package.

[0095] Further, the human body joint angle calculation module 62 comprises:

[0096] A human body key point extraction unit, configured to extract human body key points in the user video based on a human body skeleton point detection manner;

[0097] A target joint selection unit, configured to select a preset joint of the human body as a target joint, and generate a joint vector corresponding to the target joint;

[0098] A joint angle calculation unit, configured to calculate a human body joint angle corresponding to the target joint by a first preset formula based on the joint vector.

[0099] Further, the target joint selection unit comprises:

[0100] A target joint determination subunit, configured to select a preset joint of the human body as a target joint, wherein the target joint comprises a right elbow, a left elbow, a left shoulder, a right shoulder, a left hip, a right hip, a left knee, and a right knee;

[0101] An origin key point acquisition subunit, configured to acquire a human body key point corresponding to any target joint as an origin key point;

[0102] A basic key point acquisition subunit, configured to acquire two human body key points closest to the origin key point as basic key points;

[0103] A joint vector generation subunit, configured to subtract coordinates of the two basic key points from coordinates of the origin key point to obtain a joint vector.

[0104] Further, the human body posture distance calculation module 63 comprises:

[0105] a target angle difference calculation unit configured to calculate an angle difference between the human joint angle and the model joint angle by a second preset formula to obtain a target angle difference;

[0106] a human posture distance generation unit configured to calculate a Chebyshev distance of video frames in the user video and the model video based on the target angle difference, the model feature and the human key point to obtain a human posture distance;

[0107] The second preset formula is: Wherein, a is the human joint angle, b is the model joint angle, d a (a, b) is the target angle difference.

[0108] Further, the human posture distance calculation module 63 comprises:

[0109] a same frame key point acquisition unit configured to acquire human key points corresponding to the same frame in the user video and the multiple model videos to obtain user same frame key points and multiple model same frame key points;

[0110] a posture distance set generation unit configured to calculate human posture distances between the multiple model same frame key points and the user same frame key points respectively to obtain a posture distance set;

[0111] a distance to be deleted identification unit configured to take the posture distance with the shortest distance in the posture distance set as a target distance, and take posture distances in the posture distance set except the target distance as distances to be deleted;

[0112] a model video deletion unit configured to delete the model video corresponding to the distance to be deleted.

[0113] Further, the path curve construction module 64 comprises:

[0114] a two-dimensional network framework unit configured to construct a two-dimensional grid according to the number of video frames of the user video and the model video;

[0115] a distance data construction unit configured to construct a distance array of the two-dimensional grid based on the human posture distance;

[0116] a path curve generation unit configured to sequentially acquire elements with minimum values from the two-dimensional grid based on the distance data, and construct a path curve based on the elements with minimum values.

[0117] To solve the above technical problems, the embodiment of the present application further provides a computer device. For details, please refer to Figure 10 , Figure 10 The basic structure block diagram of the computer device of the present embodiment is shown in the figure.

[0118] The computer device 7 includes a memory 71, a processor 72, and a network interface 73 which are communicatively connected by a system bus. It should be noted that the computer device 7 is only shown with three components, the memory 71, the processor 72, and the network interface 73, but it should be understood that not all of the shown components are required to be implemented, and more or less components can be alternatively implemented. Among them, those skilled in the art can understand that the computer device herein is a device capable of automatically performing numerical calculation and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0119] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The computer device can interact with the user through a keyboard, a mouse, a remote controller, a touchpad, a voice control device, and the like.

[0120] The memory 71 includes at least one type of readable storage medium, which includes a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, and the like. In some embodiments, the memory 71 can be an internal storage unit of the computer device 7, such as a hard disk or a memory of the computer device 7. In other embodiments, the memory 71 can also be an external storage device of the computer device 7, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like. Of course, the memory 71 can also include both the internal storage unit and the external storage device of the computer device 7. In the present embodiment, the memory 71 is generally used to store an operating system and various application software installed in the computer device 7, such as program codes of the method for aligning video frames based on human key points, and the like. In addition, the memory 71 can also be used to temporarily store various data that have been output or will be output.

[0121] The processor 72 may, in some embodiments, be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 72 is generally used to control the overall operation of the computer device 7. In the present embodiment, the processor 72 is configured to run program codes or process data stored in the memory 71, such as program codes of the above-mentioned human key point based video frame alignment method, to implement various embodiments of the human key point based video frame alignment method.

[0122] The network interface 73 may include a wireless network interface or a wired network interface, and is generally used to establish a communication connection between the computer device 7 and other electronic devices.

[0123] The present application also provides another embodiment, i.e., to provide a computer readable storage medium, which stores a computer program, and the computer program can be executed by at least one processor to make the at least one processor execute the steps of a human key point based video frame alignment method as described above.

[0124] From the above description of the embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be realized by means of software and necessary general hardware platforms, of course, they can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the methods of various embodiments of the present application.

[0125] Obviously, the above-described embodiments are only some of the embodiments of the present application, not all the embodiments, and the preferred embodiments of the present application are given in the drawings, but do not limit the patent scope of the present application. The present application can be realized in many different forms, and conversely, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing specific embodiments, or make equivalent replacements to some technical features. Any equivalent structure made by using the contents of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the scope of the patent protection of the present application.

Claims

1. A human key point based video frame alignment method, characterized in that, The method comprises: obtaining a user video and a special effect resource package, wherein the special effect resource package comprises a model video, a special effect video, model features and model joint angle points; extracting human body key points in the user video and calculating human body joint angle points based on the human body key points; calculating human body posture distances of video frames in the user video and the model video based on the human body joint angle points, the model joint angle points, the model features and the human body key points; constructing path curves corresponding to the user video and the model video based on the human body posture distances; mapping the user video and the special effect video according to the path curves to align frames of the user video and the special effect video; the constructing path curves corresponding to the user video and the model video based on the human body posture distances comprises: constructing a two-dimensional grid according to the number of video frames of the user video and the model video; constructing a distance array of the two-dimensional grid based on the human body posture distances; obtaining elements with minimum values from the two-dimensional grid in sequence based on the distance array, and constructing the path curves based on the elements with minimum values; before the calculating human body posture distances of video frames in the user video and the model video based on the human body joint angle points, the model joint angle points, the model features and the human body key points, the method further comprises: obtaining human body key points of the same frames in the user video and a plurality of the model videos to obtain user same frame key points and a plurality of model same frame key points; performing human body posture distance calculation on a plurality of the model same frame key points and the user same frame key points respectively to obtain a posture distance set; taking a posture distance with the shortest distance in the posture distance set as a target distance, and obtaining posture distances in the posture distance set except the target distance as to-be-deleted distances; deleting model videos corresponding to the to-be-deleted distances.

2. The human keypoint-based video frame alignment method of claim 1, wherein, the extracting human body key points in the user video and calculating human body joint angle points based on the human body key points comprises: extracting the human body key points in the user video based on a human body skeleton point detection method; selecting a preset joint of a human body as a target joint and generating a joint vector corresponding to the target joint; calculating the human body joint angle corresponding to the target joint through a first preset formula based on the joint vector.

3. The human keypoint-based video frame alignment method of claim 2, wherein, the selecting a preset joint of a human body as a target joint and generating a joint vector corresponding to the target joint comprises: selecting a preset joint of a human body as the target joint, wherein the target joint comprises a right elbow, a left elbow, a left shoulder, a right shoulder, a left hip, a right hip, a left knee and a right knee; for any target joint, obtaining a human body key point corresponding to the target joint as an origin key point; obtaining two human body key points closest to the origin key point as basic key points; subtracting coordinates of the two basic key points from coordinates of the origin key point to obtain the joint vector.

4. The human keypoint-based video frame alignment method of claim 1, wherein, The human joint angle, the model joint angle, the model feature, and the human key point are used to calculate a human posture distance of video frames in the user video and the model video. A second preset formula is used to calculate an angle difference between the human joint angle and the model joint angle, to obtain a target angle difference. The target angle difference, the model feature, and the human key point are used to calculate a Chebyshev distance of video frames in the user video and the model video, to obtain the human posture distance. The second preset formula is: Wherein, a is the human joint angle, b is the model joint angle, The target angle difference is the target angle difference.

5. A human keypoint based video frame alignment apparatus, characterized by, The method comprises the following steps: A user video acquisition module is configured to acquire a user video and a special effect resource package, wherein the special effect resource package comprises a model video, a special effect video, a model feature, and a model joint angle. A human joint angle calculation module is configured to extract human key points in the user video and calculate a human joint angle based on the human key points. A human posture distance calculation module is configured to calculate a human posture distance of video frames in the user video and the model video based on the human joint angle, the model joint angle, the model feature, and the human key point. A path curve construction module is configured to construct a path curve corresponding to the user video and the model video based on the human posture distance. A video frame alignment module is configured to map the user video and the special effect video according to the path curve, so that the frames of the user video and the special effect video are aligned. The path curve construction module comprises: A two-dimensional network architecture unit is configured to construct a two-dimensional grid according to the number of video frames of the user video and the model video. A distance array construction unit is configured to construct a distance array of the two-dimensional grid based on the human posture distance. A path curve generation unit is configured to sequentially obtain elements with minimum values from the two-dimensional grid based on the distance array, and construct the path curve based on the elements with minimum values. Before the human posture distance calculation module, the video frame alignment device based on human key points further comprises: A same frame key point acquisition unit is configured to acquire human key points corresponding to the same frames in the user video and a plurality of model videos, to obtain user same frame key points and a plurality of model same frame key points. A posture distance set generation unit is configured to calculate human posture distances of a plurality of model same frame key points and the user same frame key points respectively, to obtain a posture distance set. A distance to be deleted identification unit is configured to take a posture distance with the shortest distance in the posture distance set as a target distance, and take posture distances in the posture distance set other than the target distance as distances to be deleted. A model video deletion unit is configured to delete a model video corresponding to the distance to be deleted.

6. A computer device, comprising: The device comprises a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the video frame alignment method based on human key points.

7. A computer readable storage medium characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the human body key point based video frame alignment method in any one of claims 1 to 4.

Citation Information

Patent Citations

  • GIF graph generation method and device, electronic equipment and storage medium

    CN110910478A