A key point recognition method and device, electronic equipment and storage medium
By predicting multiple locations and candidate regions of key points in video frames using a pre-set model and combining the final locations of adjacent frames, the problem of facial key point jitter and inaccurate prediction in videos is solved, thus improving stability and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2022-03-17
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies for identifying facial key points in videos suffer from several drawbacks. First, the weighted smoothing coefficient requires manual adjustment, which is inflexible. Second, the jitter problem persists, and third, when the facial position changes significantly between adjacent frames, the key point prediction is inaccurate.
The system predicts at least three predicted positions for each key point in the target video frame using a pre-set model, determines candidate regions, and selects the final position from the candidate regions based on the final positions of adjacent frames, thus avoiding manual parameter adjustment.
It significantly reduces jitter, improves the accuracy and flexibility of keypoint prediction, and ensures the stability of keypoints in video frames.
Smart Images

Figure CN116797957B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image recognition technology, and in particular to a method, apparatus, electronic device, and storage medium for identifying key points. Background Technology
[0002] Facial landmarks can represent specific locations on a face, such as pupil points, nose points, and contour points. When recognizing facial landmarks in videos, there are cases where the positional change of the face is very small between two adjacent frames, but the predicted landmarks differ; this can manifest as landmark jitter. Landmark jitter affects subsequent face-related tasks and therefore urgently needs to be addressed.
[0003] In existing technologies, when identifying facial landmarks in a video, after predicting the facial landmarks of the current frame, the facial landmarks are usually weighted and smoothed with the facial landmarks of the adjacent frames that are earlier in the time sequence to obtain the final facial landmarks of the current frame in order to reduce jitter.
[0004] The shortcomings of existing technologies include at least the following: the weighted smoothing coefficients need to be manually adjusted, which lacks flexibility; smoothing can only slightly alleviate jitter, and the jitter problem still exists; when the position of the face changes significantly between adjacent frames, there will be a problem that the predicted key points will deviate greatly from the actual specific position, resulting in inaccurate key point prediction. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, and storage medium for key point identification. It eliminates the need for manual parameter adjustment, significantly reduces jitter, and ensures the accuracy of key point prediction.
[0006] In a first aspect, embodiments of this disclosure provide a method for identifying key points, including:
[0007] Input the target video frame into a preset model, and output at least three predicted positions of each key point in the target video frame based on the preset model;
[0008] Based on at least three predicted positions of each key point, a candidate region for each key point in the target video frame is determined;
[0009] Based on the final position of each key point in the adjacent frame that is earlier in the time sequence than the target video frame, the final position of the corresponding key point in the target video frame is determined from each candidate region.
[0010] Secondly, embodiments of this disclosure also provide a key point identification device, including:
[0011] The prediction location determination module is used to input the target video frame into a preset model and output at least three predicted locations of each key point in the target video frame based on the preset model.
[0012] The candidate region determination module is used to determine the candidate region of each key point in the target video frame based on at least three predicted positions of each key point;
[0013] The final position determination module is used to determine the final position of the corresponding key point in the target video frame from each of the candidate regions based on the final position of each key point in an adjacent frame that is earlier in the time sequence than the target video frame.
[0014] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0015] One or more processors;
[0016] Storage device for storing one or more programs.
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the key point identification method as described in any embodiment of this disclosure.
[0018] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform a key point identification method as described in any of the embodiments of this disclosure.
[0019] The technical solution of this disclosure involves inputting a target video frame into a preset model, outputting at least three predicted positions of each key point in the target video frame based on the preset model, determining candidate regions of each key point in the target video frame based on the at least three predicted positions of each key point, and determining the final position of the corresponding key point in the target video frame from each candidate region based on the final position of each key point in an adjacent frame that is earlier in the time sequence than the target video frame.
[0020] By using a pre-defined model, three possible locations for each keypoint in the target video frame can be predicted. From these three pre-defined locations, candidate regions where each keypoint is likely to appear in the target video frame can be obtained. Since the positions of keypoints are correlated between adjacent video frames, the position with the highest stability for each keypoint can be selected from the candidate regions based on its final position in the previous adjacent frame. This method not only ensures the accuracy of keypoint prediction and significantly reduces jitter, but also avoids manual parameter adjustments, improving the method's flexibility. Attached Figure Description
[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0022] Figure 1 This is a flowchart illustrating a key point identification method provided in Embodiment 1 of this disclosure;
[0023] Figure 2 This is a schematic diagram showing the positions of each key point in a key point identification method provided in Embodiment 1 of this disclosure.
[0024] Figure 3 This is a schematic diagram of the structure of the preset model in the key point identification method provided in Embodiment 2 of this disclosure;
[0025] Figure 4 This is a schematic diagram of the structure of the preset model in the key point identification method provided in Embodiment 2 of this disclosure;
[0026] Figure 5 This is a flowchart illustrating the process of determining the parameters of the fully connected layer of a preset model in a key point identification method provided in Embodiment 2 of this disclosure.
[0027] Figure 6 This is a schematic diagram of the structure of a key point identification device provided in Embodiment 3 of this disclosure;
[0028] Figure 7 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of this disclosure. Detailed Implementation
[0029] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0030] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0031] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0032] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0033] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0034] Example 1
[0035] Figure 1 This is a flowchart illustrating a key point recognition method provided in Embodiment 1 of this disclosure. This embodiment is applicable to key point recognition in videos, such as facial key point recognition in videos. The method can be executed by a key point recognition device, which can be implemented in software and / or hardware and can be configured in electronic devices, such as mobile phones, computers, etc.
[0036] like Figure 1 As shown, the key point identification method provided in this embodiment may include:
[0037] S110. Input the target video frame into the preset model, and output at least three predicted positions of each key point in the target video frame based on the preset model.
[0038] In the process of developing the key point recognition method provided in this disclosure, two types of experiments were conducted: 1. Inputting nearly static adjacent frames into the same model to output the positions of key points; 2. Inputting the same video frame into different models to output the positions of key points. The result of Experiment 1 was that the positions of key points in nearly static adjacent frames were different. The result of Experiment 2 was that the positions of key points output by different models for the same video frame were also different.
[0039] The conclusions that can be drawn from the above experiments include at least the following: for the same video frame, there can be multiple reasonable prediction positions for the same key point; the jitter phenomenon that occurs when using the same model to predict the position of key points in adjacent frames that are basically the same tends to be consistent with the jitter phenomenon that occurs when using different models to predict the position of key points in the same video frame.
[0040] Based on this, in this embodiment of the disclosure, different models can be used to predict at least three possible locations for each keypoint in the target video frame, thus obtaining at least three predicted locations. These predicted locations can be considered as reasonable locations where each keypoint may appear, and can be considered equivalent to using the same model to predict the possible jitter locations of each keypoint when predicting the target video frame after predicting the previous video frame.
[0041] The preset model in this embodiment may include: multiple different models, and / or the same model containing multiple output branches. The preset model may be obtained in advance through supervised, semi-supervised, or unsupervised training, and the sample videos used during training are of the same video type as the target video frame. The keypoint labels during training are the same in number and category as the keypoints in the predicted target video frame.
[0042] For different types of videos, corresponding preset models can be used to predict different numbers and categories of keypoints. For example, for facial videos, a corresponding preset model can predict 68 facial keypoints, and the categories of keypoints can include, for example, pupil points, nose points, corner points of the mouth points, and facial contour points. As another example, for body videos, a corresponding preset model can predict 17 body keypoints, and the categories of keypoints can include, for example, shoulder points, elbow points, wrist points, hip points, and knee points.
[0043] In this embodiment, at least three predicted positions for each keypoint in the target video frame can be predicted based on a preset model. For example, when there are 10 keypoints in the target video frame, if three predicted positions can be predicted for each keypoint, then a total of 30 predicted positions can be obtained for 10 keypoints.
[0044] S120. Based on at least three predicted positions of each key point, determine the candidate region of each key point in the target video frame.
[0045] For each keypoint, candidate regions in the target video with a high probability of appearing can be determined based on at least three predicted locations. The methods for determining candidate regions may include, but are not limited to: connecting the outermost of the at least three predicted locations to form a line and using the enclosed area as a candidate region; or fitting a preset shape to the at least three predicted locations and using the area within the preset shape as a candidate region, where the candidate shape may be, for example, a triangle, a circle, etc. Other methods for determining candidate regions based on predicted locations can also be applied here, and will not be exhaustive.
[0046] S130. Based on the final position of each key point in the adjacent frame that is earlier than the target video frame in the time sequence, determine the final position of the corresponding key point in the target video frame from each candidate region.
[0047] In this embodiment, the adjacent frame that is earlier than the target video frame in the time sequence can be considered as the previous video frame of the target video frame. After obtaining the final position of each key point in the previous video frame of the target video frame, for each key point, a position with less jitter compared to the final position of the key point in the previous video frame can be determined in the corresponding candidate region, and this position can be used as the final position of the key point in the target video frame, thereby significantly reducing jitter.
[0048] Furthermore, by selecting the final location from the candidate regions of each key point, the problem of large deviations between the predicted key points and the actual specific locations can be avoided, thus improving prediction accuracy. In addition, compared with the traditional method of identifying key points using weighted smoothing, this embodiment does not require manual adjustment of the weighting parameters, improving the flexibility of the method.
[0049] For example, Figure 2 This is a schematic diagram showing the location of each key point in a key point identification method provided in Embodiment 1 of this disclosure. Figure 2 In this paper, taking the key points of the eyebrow region in a facial video as an example, the paper illustrates the multiple predicted positions of each key point in the target video frame, as well as the final positions of each key point in adjacent frames that are earlier than the target video in time sequence.
[0050] See Figure 2 The key points in the eyebrow region can be categorized into the inner brow, the arch, and the tail. The predicted position of the inner brow in the target video frame can include three positions: A1, A2, and A3. The triangle formed by these three positions can be considered as the candidate region for the inner brow in the target video frame. Similarly, the predicted position of the arch can include three positions: B1, B2, and B3. The triangle formed by these three positions can be considered as the candidate region for the arch. The predicted position of the tail can include three positions: C1, C2, and C3. Furthermore, the final position of the inner brow in an adjacent frame that is earlier in the target video sequence can be represented by A', the final position of the arch in that adjacent frame can be represented by B', and the final position of the tail in that adjacent frame can be represented by C'.
[0051] In some optional implementations, the final position of the corresponding key point in the target video frame is determined from each candidate region based on the final position of each key point in the adjacent frame that is earlier in the time sequence than the target video frame. This includes: determining whether the final position of each key point in the adjacent frame is within the corresponding candidate region; and determining the final position of the corresponding key point in the target video frame from each candidate region based on the determination result.
[0052] Specifically, for each keypoint, it can be determined whether its final position in adjacent frames falls within the corresponding candidate region in the target video frame. The boundary of the candidate region may or may not be included within the candidate region; this is not limited here. The determination result can include whether the keypoint's final position in adjacent frames falls within the corresponding candidate region, or whether the keypoint's final position in adjacent frames does not fall within the corresponding candidate region.
[0053] Based on the judgment results, the final position of the corresponding keypoint in the target video frame is determined from each candidate region. This can include: if the final position of each keypoint in an adjacent frame is within the corresponding candidate region, then the final position of each keypoint in the adjacent frame can be used as the final position in the target video frame. This can minimize jitter and ensure the stability of the keypoints.
[0054] If there are keypoints in adjacent frames whose final positions are not within the corresponding candidate regions, it can be assumed that these points have undergone significant displacement between the adjacent frames and the target video frame. In this case, the change in the keypoint position can be considered not as jitter, but as keypoint movement. In this situation, determining the final position of the corresponding keypoint in the target video frame from each candidate region based on the judgment result can include: determining the target point from each keypoint based on the judgment result; the final position of this target point in the adjacent frames is not within the corresponding candidate region. Any position within the candidate region corresponding to the target point is taken as the final position of the target point in the target video frame. Furthermore, for all keypoints other than the target point, their final positions in adjacent frames can also be taken as their final positions in the target video frame.
[0055] In some implementations, after determining the target point based on the judgment result, the target position with the smallest distance from the final position of the target point in the adjacent frame can be determined from the candidate region corresponding to the target point, and this target position can be used as the final position of the target point in the target video frame. This ensures that the movement distance of each target point is minimized, improving the uniformity and stability of each key point in the target video frame.
[0056] For example, see again Figure 2If A' and C' are within their respective candidate regions, then A' and C' can be taken as the final positions of the brow head and brow tail in the target video frame, respectively. If B' is not within its corresponding candidate region, it can be assumed that the brow peak has undergone a significant displacement. In this case, the target position b with the smallest distance from B' can be determined from the candidate region of the brow peak, and b can be taken as the final position of the brow peak in the target video frame.
[0057] In some optional implementations, this method is applied to facial processing applications, with the key point being facial key points. These facial processing applications may include, but are not limited to, facial recognition applications, expression recognition applications, and facial effect addition applications.
[0058] In these optional implementations, facial processing applications can capture facial videos in real time or read pre-stored facial videos from a preset storage space. The steps for facial processing applications to identify facial key points may include: inputting video frames of the facial video into a preset model in a temporal sequence, with the currently input video frame serving as the target video frame; outputting at least three predicted positions for each facial key point in the target video frame based on the preset model; determining candidate regions for each facial key point in the target video frame based on the at least three predicted positions; and determining the final position of the corresponding facial key point in the target video frame from each candidate region based on the final position of each facial key point in the previous input frame. The final position of each facial key point in the first video frame can be any position within the candidate region, such as the center position or any predicted position, and is not limited here.
[0059] Furthermore, keypoint recognition methods can be applied not only to facial processing applications but also to other types of videos. The process of recognizing keypoints in other types of videos can be referenced from the process of recognizing facial keypoints in facial videos, and will not be exhaustively detailed here.
[0060] The technical solution of this disclosure involves inputting a target video frame into a preset model, outputting at least three predicted positions of each key point in the target video frame based on the preset model, determining candidate regions of each key point in the target video frame based on the at least three predicted positions of each key point, and determining the final position of the corresponding key point in the target video frame from each candidate region based on the final position of each key point in an adjacent frame that is earlier in the time sequence than the target video frame.
[0061] By using a pre-defined model, three possible locations for each keypoint in the target video frame can be predicted. From these three pre-defined locations, candidate regions where each keypoint is likely to appear in the target video frame can be obtained. Since the positions of keypoints are correlated between adjacent video frames, the position with the highest stability for each keypoint can be selected from the candidate regions based on its final position in the previous adjacent frame. This method not only ensures the accuracy of keypoint prediction and significantly reduces jitter, but also avoids manual parameter adjustments, improving the method's flexibility.
[0062] Example 2
[0063] This embodiment can be combined with various optional schemes in the key point recognition method provided in the above embodiments. The key point recognition method provided in this embodiment details the steps for generating multiple predicted locations for each key point. By connecting different fully connected layers after the same feature extraction layer of the preset model, multiple predicted locations for each key point can be output based on each fully connected layer. Furthermore, processing the feature image through deep convolutional layers between the input fully connected layers can increase the diversity of the prediction results, ensuring that the candidate region is within a suitable range.
[0064] Furthermore, this embodiment describes in detail the steps for determining the coefficients of each fully connected layer. By changing the quality of the sample video frames, multiple adjustment frames can be obtained. Since the positions of the same key points in each adjustment frame are theoretically consistent, by using multiple fully connected layers to determine multiple predicted positions for each adjustment frame, and determining the jitter vector for each fully connected layer based on the predicted positions of the same key points in each adjustment frame, the coefficients of each fully connected layer can be trained based on the angle between the jitter vectors. This allows for the output of robust candidate regions under video frames of different quality based on the coefficients of each fully connected layer.
[0065] For example, Figure 3 This is a schematic diagram of the structure of a preset model in a key point identification method provided in Embodiment 2 of this disclosure. See also... Figure 3 In some optional implementations, outputting at least three predicted positions for each keypoint in the target video frame based on a preset model may include:
[0066] Based on the feature extraction layer of the preset model (represented by FE in the figure), the feature image of the target video frame (represented by f in the figure) is determined; the feature image f is input into at least three fully connected layers of the preset model (represented by FC in the figure), and at least three predicted positions of each key point are output based on the at least three fully connected layers FC (represented by landmark-1, landmark-2, ..., landmark-N in the figure).
[0067] Among these alternative implementations, by sharing the feature extraction layer of a pre-defined model and then connecting different fully connected layers, it is possible not only to output multiple predicted locations for each keypoint based on each fully connected layer, but also to simplify model deployment.
[0068] For example, Figure 4 This is a schematic diagram of the structure of a preset model in a key point identification method provided in Embodiment 2 of this disclosure. Figure 4 for Figure 3 Some improvements based on the original design; for details not described, please refer to [link to relevant documentation]. Figure 3 See also Figure 4 In some further implementations, before inputting the feature image into at least three fully connected layers of the predefined model, the following may also be included:
[0069] The feature image f is input into the depthwise convolutional layer corresponding to each fully connected layer (represented by depth-wise in the figure); correspondingly, the feature image is input into at least three fully connected layers of the preset model, including: inputting the convolution result of each depthwise convolutional layer into the corresponding fully connected layer FC.
[0070] Among these alternative implementations, processing the feature image through deep convolutional layers between fully connected input layers can increase the diversity of prediction results and ensure that the candidate region is within a suitable range.
[0071] For example, Figure 5 This is a flowchart illustrating the process of determining the parameters of the fully connected layer of a preset model in a key point identification method provided in Embodiment 2 of this disclosure. See also... Figure 5 In some alternative implementations, the coefficients of at least three fully connected layers can be determined through the following steps:
[0072] First, the image quality of the sample video frames (represented by S0 in the figure) is adjusted to obtain each adjusted frame (represented by S1 and S2 in the figure). Figure 5 In the diagram, processing steps related to S1 are represented by solid lines, and processing steps related to S2 are represented by dashed lines. Adjusting the image quality of the sample video frames may include, but is not limited to: adding noise to the sample video frames, smoothing and denoising the sample video frames, increasing the resolution of the sample video frames, and changing the brightness of the sample video frames. Each adjusted frame can be a video frame after the sample video has undergone image quality adjustment according to the same adjustment type but different adjustment degrees; or it can be a video frame after the sample video has undergone image quality adjustment according to different adjustment types.
[0073] Secondly, the feature images of each adjustment frame (represented by f1 and f2 in the figure) are input into at least three fully connected layers of the preset model (including FC1, FC2, and FC3 in the figure). The adjustment frames S1 and S2 can be feature extracted based on the same feature extraction layer FE of the preset model, yielding corresponding feature images f1 and f2 respectively. Furthermore, after determining the feature images, before inputting them into the fully connected layers, they can be processed by a deep convolutional layer to increase the diversity of the prediction results.
[0074] Next, based on at least three fully connected layers, the predicted positions of each keypoint in each adjustment frame are output. For example, for keypoint 1 in S0, based on the fully connected layers FC1, FC2, and FC3 of the same preset model, the three predicted positions of the keypoint in adjustment frame S1 (landmark-1, landmark-2, and landmark-3) are output according to feature image f1, and the three predicted positions of the keypoint in adjustment frame S2 (landmark-4, landmark-5, and landmark-6) are output according to feature image f2. Furthermore, the predicted positions of other keypoints besides keypoint 1 in each adjustment frame can be determined in the same way.
[0075] Then, for each fully connected layer, the jitter vector is determined based on the predicted position of the same keypoints in each adjustment frame. For example, for... Figure 5 In the fully connected layer FC1, landmark-1 and landmark-4 are the predicted positions of keypoint 1 in adjustment frames S1 and S2. The jitter vector V1 of the fully connected layer FC1 can be obtained by subtracting the position coordinates of landmark-1 and landmark-4. The direction of the jitter vector can be, for example, from landmark-1 to landmark-4, or from landmark-4 to landmark-1. Similarly, the jitter vectors V2 and V3 for the fully connected layers FC2 and FC3 can be determined, and the directions of V2 and V3 are determined in the same way as V1.
[0076] Finally, the coefficients of each fully connected layer are adjusted to achieve a preset angle between the dither vectors of each fully connected layer. The preset angle can be set to different values depending on the number of fully connected layers. For example, when the number of fully connected layers is... Figure 5 To ensure that the predicted positions based on the outputs of the three fully connected layers have the same probability of occurrence in the video frame, the preset angle can be set to 120 degrees. At this point, the parameters of the three fully connected layers FC1, FC2, and FC3 can be adjusted with the goal of ensuring that the included angles between V1, V2, and V3 all approach 120 degrees.
[0077] Furthermore, in some implementations, the number of adjustment frames can be three or more. In this case, each fully connected layer can output three or more preset positions for the same keypoint. Then, after outputting the predicted positions of each keypoint in each adjustment frame based on at least three fully connected layers, adjusting the parameters of each fully connected layer can include:
[0078] First, for each fully connected layer, initial jitter vectors can be determined between each pair of preset positions. Then, the jitter vector of the corresponding fully connected layer can be determined based on each initial jitter vector. For example, the average, maximum, or minimum vectors of each initial jitter vector can be used as the jitter vector of the corresponding fully connected layer. Finally, the coefficients of each fully connected layer can be adjusted with the goal of satisfying a preset angle between the jitter vectors of each fully connected layer.
[0079] Alternatively, after determining the initial jitter vectors for each fully connected layer, the following approach can be adopted: randomly selecting the initial jitter vectors for each fully connected layer, and adjusting the coefficients of each fully connected layer with the goal of ensuring that the angle between the selected initial jitter vectors satisfies a preset angle. Furthermore, other methods for adjusting the coefficients of fully connected layers can also be applied here, but will not be exhaustive.
[0080] The technical solution of this disclosure provides a detailed description of the steps for generating multiple predicted locations for each keypoint. By connecting different fully connected layers after the same feature extraction layer of a preset model, multiple predicted locations for each keypoint can be output based on each fully connected layer. Furthermore, processing the feature image through deep convolutional layers between the input fully connected layers can increase the diversity of the prediction results and ensure that the candidate region is within a suitable range.
[0081] Furthermore, this embodiment describes in detail the steps for determining the coefficients of each fully connected layer. By changing the quality of the sample video frames, multiple adjustment frames can be obtained. Since the positions of the same key points in each adjustment frame are theoretically consistent, by using multiple fully connected layers to determine multiple predicted positions for each adjustment frame, and determining the jitter vector for each fully connected layer based on the predicted positions of the same key points in each adjustment frame, the coefficients of each fully connected layer can be trained based on the angle between the jitter vectors. This allows for the output of robust candidate regions under video frames of different quality based on the coefficients of each fully connected layer.
[0082] The key point identification method provided in this embodiment belongs to the same concept as the key point identification method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and the same technical features have the same beneficial effects in this embodiment and the above embodiments.
[0083] Example 3
[0084] Figure 6 This is a schematic diagram of a key point recognition device provided in Embodiment 3 of this disclosure. This embodiment is applicable to key point recognition in videos, such as facial key point recognition in videos.
[0085] like Figure 6 As shown, the key point recognition device provided in this embodiment may include:
[0086] The prediction location determination module 610 is used to input the target video frame into a preset model and output at least three predicted locations of each key point in the target video frame based on the preset model.
[0087] The candidate region determination module 620 is used to determine the candidate region of each key point in the target video frame based on at least three predicted positions of each key point.
[0088] The final position determination module 630 is used to determine the final position of the corresponding key point in the target video frame from each candidate region based on the final position of each key point in the adjacent frame that is earlier than the target video frame in the time sequence.
[0089] In some alternative implementations, the final location determination module can be used for:
[0090] Determine whether the final position of each key point in the adjacent frames is within the corresponding candidate region;
[0091] Based on the judgment results, the final position of the corresponding key point in the target video frame is determined from each candidate region.
[0092] In some alternative implementations, the final location determination module can be used for:
[0093] If the final position of each keypoint in the adjacent frame is within the corresponding candidate region, then the final position of each keypoint in the adjacent frame is taken as the final position in the target video frame.
[0094] In some alternative implementations, the final location determination module can be used for:
[0095] Based on the judgment results, the target point is determined from each key point; the final position of the target point in the adjacent frame is not within the corresponding candidate area;
[0096] From the candidate region corresponding to the target point, determine the target position that is the smallest distance from the final position of the target point in the adjacent frame, and take the target position as the final position of the target point in the target video frame.
[0097] In some alternative implementations, the predicted location determination module can be used for:
[0098] Based on the feature extraction layer of the preset model, the feature image of the target video frame is determined;
[0099] The feature image is input into at least three fully connected layers of a preset model, and at least three predicted positions of each key point are output based on the at least three fully connected layers.
[0100] In some alternative implementations, the predicted location determination module can also be used for:
[0101] Before inputting the feature image into at least three fully connected layers of the preset model, the feature image is input into the deep convolutional layer corresponding to each fully connected layer.
[0102] Accordingly, the feature image is input into at least three fully connected layers of the preset model, including: inputting the convolution result of each depth convolutional layer into the corresponding fully connected layer.
[0103] In some alternative implementations, the key point recognition device may further include:
[0104] The coefficient determination module can determine the coefficients of at least three fully connected layers through the following steps:
[0105] The image quality of the sample video frames is adjusted to obtain each adjusted frame;
[0106] Input the feature images of each adjustment frame into at least three fully connected layers of the preset model;
[0107] Based on at least three fully connected layers, the predicted positions of each keypoint in each adjustment frame are output;
[0108] For each fully connected layer, the jitter vector is determined based on the predicted position of the same key point in each adjustment frame.
[0109] The coefficients of each fully connected layer are adjusted with the goal of satisfying a preset angle between the jitter vectors of each fully connected layer.
[0110] In some alternative implementations, the key point recognition device can be applied to facial processing applications, where the key points are facial key points.
[0111] The key point identification device provided in this disclosure can execute the key point identification method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of the method execution.
[0112] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0113] Example 4
[0114] The following is for reference. Figure 7 It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 7 The diagram below shows the structure of the terminal device or server 700. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0115] like Figure 7 As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 706 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0116] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0117] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 709, or installed from storage device 706, or installed from ROM 702. When the computer program is executed by processing device 701, it performs the functions defined in the key point identification method of embodiments of this disclosure.
[0118] The electronic device provided in this embodiment and the key point identification method provided in the above embodiments belong to the same disclosed concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0119] Example 5
[0120] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the key point identification method provided in the above embodiments.
[0121] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory (FLASH), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0122] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0123] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0124] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:
[0125] Input the target video frame into the preset model, and output at least three predicted positions of each key point in the target video frame based on the preset model; determine the candidate region of each key point in the target video frame based on the at least three predicted positions of each key point; determine the final position of the corresponding key point in the target video frame from each candidate region based on the final position of each key point in the adjacent frame that is earlier in the time sequence than the target video frame.
[0126] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0127] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0128] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units and modules do not, in certain circumstances, constitute a limitation on the unit or module itself.
[0129] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0130] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0131] According to one or more embodiments of this disclosure, [Example 1] provides a method for identifying key points, the method comprising:
[0132] Input the target video frame into a preset model, and output at least three predicted positions of each key point in the target video frame based on the preset model;
[0133] Based on at least three predicted positions of each key point, a candidate region for each key point in the target video frame is determined;
[0134] Based on the final position of each key point in the adjacent frame that is earlier in the time sequence than the target video frame, the final position of the corresponding key point in the target video frame is determined from each candidate region.
[0135] According to one or more embodiments of this disclosure, [Example 2] provides a method for identifying key points, further comprising:
[0136] In some optional implementations, determining the final position of the corresponding key point in the target video frame from each candidate region based on the final position of each key point in a neighboring frame that is earlier in the time sequence than the target video frame includes:
[0137] Determine whether the final position of each key point in the adjacent frame is within the corresponding candidate region;
[0138] Based on the judgment results, the final position of the corresponding key point in the target video frame is determined from each of the candidate regions.
[0139] According to one or more embodiments of this disclosure, [Example 3] provides a method for identifying key points, further comprising:
[0140] In some optional implementations, determining the final position of the corresponding key point in the target video frame from each of the candidate regions based on the judgment result includes:
[0141] If the final position of each key point in the adjacent frame is within the corresponding candidate region, then the final position of each key point in the adjacent frame is taken as the final position in the target video frame.
[0142] According to one or more embodiments of this disclosure, [Example 4] provides a method for identifying key points, further comprising:
[0143] In some optional implementations, determining the final position of the corresponding key point in the target video frame from each of the candidate regions based on the judgment result includes:
[0144] Based on the judgment results, a target point is determined from each of the key points; the final position of the target point in the adjacent frame is not within the corresponding candidate region;
[0145] From the candidate region corresponding to the target point, determine the target position that is the smallest distance from the final position of the target point in the adjacent frame, and take the target position as the final position of the target point in the target video frame.
[0146] According to one or more embodiments of this disclosure, [Example 5] provides a method for identifying key points, further comprising:
[0147] In some optional implementations, the step of outputting at least three predicted positions of each key point in the target video frame based on the preset model includes:
[0148] Based on the feature extraction layer of the preset model, the feature image of the target video frame is determined;
[0149] The feature image is input into at least three fully connected layers of the preset model, and at least three predicted positions of each key point are output based on the at least three fully connected layers.
[0150] According to one or more embodiments of this disclosure, Example Six provides a method for identifying key points, further comprising:
[0151] In some alternative implementations, before inputting the feature image into at least three fully connected layers of the preset model, the method further includes:
[0152] The feature images are respectively input into the depth convolutional layers corresponding to each of the fully connected layers;
[0153] Accordingly, inputting the feature image into at least three fully connected layers of the preset model includes: inputting the convolution result of each depth convolutional layer into the corresponding fully connected layer.
[0154] According to one or more embodiments of this disclosure, [Example Seven] provides a method for identifying key points, further comprising:
[0155] In some alternative implementations, the coefficients of the at least three fully connected layers are determined by the following steps:
[0156] The image quality of the sample video frames is adjusted to obtain each adjusted frame;
[0157] The feature images of each adjustment frame are input into at least three fully connected layers of the preset model;
[0158] Based on the at least three fully connected layers, the predicted positions of each keypoint in each adjustment frame are output;
[0159] For each fully connected layer, the jitter vector is determined based on the predicted position of the same key point in each adjustment frame.
[0160] The coefficients of each fully connected layer are adjusted with the goal of satisfying a preset angle between the dither vectors of each fully connected layer.
[0161] According to one or more embodiments of this disclosure, [Example Eight] provides a method for identifying key points, further comprising:
[0162] In some alternative implementations, applied to facial processing applications, the key points are facial key points.
[0163] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0164] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0165] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A method for identifying key points, characterized in that, include: Input the target video frame into a preset model, and output at least three predicted positions of each key point in the target video frame based on the preset model; Based on at least three predicted positions of each key point, a candidate region for each key point in the target video frame is determined; Based on the final position of each key point in a neighboring frame that is earlier in the time sequence than the target video frame, the final position of the corresponding key point in the target video frame is determined from each candidate region, including: Determine whether the final position of each key point in the adjacent frame is within the corresponding candidate region; If the final position of each key point in the adjacent frame is within the corresponding candidate region, then the final position of each key point in the adjacent frame is taken as the final position in the target video frame. If the final position of each key point in the adjacent frame is not within the corresponding candidate area, then the target point is determined from each key point according to the judgment result. From the candidate region corresponding to the target point, any position within the candidate region corresponding to the target point is taken as the final position of the target point in the target video frame.
2. The method according to claim 1, characterized in that, The step of selecting any position within the candidate region corresponding to the target point as the final position of the target point in the target video frame includes: From the candidate region corresponding to the target point, determine the target position that is the smallest distance from the final position of the target point in the adjacent frame, and take the target position as the final position of the target point in the target video frame.
3. The method according to claim 1, characterized in that, The step of outputting at least three predicted positions for each key point in the target video frame based on the preset model includes: Based on the feature extraction layer of the preset model, the feature image of the target video frame is determined; The feature image is input into at least three fully connected layers of the preset model, and at least three predicted positions of each key point are output based on the at least three fully connected layers.
4. The method according to claim 3, characterized in that, Before inputting the feature image into at least three fully connected layers of the preset model, the method further includes: The feature images are respectively input into the depth convolutional layers corresponding to each of the fully connected layers; Accordingly, inputting the feature image into at least three fully connected layers of the preset model includes: inputting the convolution result of each depth convolutional layer into the corresponding fully connected layer.
5. The method according to claim 3, characterized in that, The coefficients of the at least three fully connected layers are determined through the following steps: The image quality of the sample video frames is adjusted to obtain each adjusted frame; The feature images of each adjustment frame are input into at least three fully connected layers of the preset model; Based on the at least three fully connected layers, the predicted positions of each keypoint in each adjustment frame are output; For each fully connected layer, the jitter vector is determined based on the predicted position of the same key point in each adjustment frame. The coefficients of each fully connected layer are adjusted with the goal of satisfying a preset angle between the dither vectors of each fully connected layer.
6. The method according to any one of claims 1-5, characterized in that, This is applied to facial processing applications, where the key points are facial key points.
7. A key point identification device, characterized in that, include: The prediction location determination module is used to input the target video frame into a preset model and output at least three predicted locations of each key point in the target video frame based on the preset model. The candidate region determination module is used to determine the candidate region of each key point in the target video frame based on at least three predicted positions of each key point; The final position determination module is used to determine the final position of the corresponding key point in the target video frame from each of the candidate regions based on the final position of each key point in an adjacent frame that is earlier in the time sequence than the target video frame, including: Determine whether the final position of each key point in the adjacent frame is within the corresponding candidate region; If the final position of each key point in the adjacent frame is within the corresponding candidate region, then the final position of each key point in the adjacent frame is taken as the final position in the target video frame. If the final position of each key point in the adjacent frame is not within the corresponding candidate area, then the target point is determined from each key point according to the judgment result. From the candidate region corresponding to the target point, any position within the candidate region corresponding to the target point is taken as the final position of the target point in the target video frame.
8. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the key point identification method as described in any one of claims 1-6.
9. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the key point identification method as described in any one of claims 1-6.