Image key point smoothing methods, apparatus, readable media, and electronic devices

By grouping and standardizing the key points in the video frames, the problems of inaccurate smoothing of facial key points and high computational cost are solved, enabling effective application on platforms with limited computing power.

CN115660983BActive Publication Date: 2026-05-26GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
Filing Date
2022-10-25
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing technologies, the smoothing effect of facial key points is inaccurate or computationally intensive, making it difficult to use on computing platforms with limited computing power.

Method used

By grouping the key points in the video frames, the key point groups corresponding to each video frame are obtained. Then, the standardization process is performed to eliminate the influence of rotation and translation factors, and the distance data of the standard key point groups between different video frames is determined and smoothed.

Benefits of technology

It improves the accuracy of image keypoint smoothing results and reduces the computational load, making it suitable for computing platforms with lower computational performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115660983B_ABST
    Figure CN115660983B_ABST
Patent Text Reader

Abstract

This disclosure provides an image keypoint smoothing method and apparatus, a computer-readable medium, and an electronic device, relating to the field of image processing technology. The method includes: acquiring image keypoints corresponding to video frames; grouping the image keypoints to obtain image keypoint groups corresponding to each video frame; standardizing the image keypoint groups to obtain standard image keypoint groups; determining distance data between the standard image keypoint groups of different video frames; and smoothing the image keypoints based on the distance data to obtain smoothed image keypoints. This disclosure considers the local information of image keypoints for smoothing and eliminates lag caused by translation and rotation through standardization, thereby improving the accuracy of the smoothing results and reducing computational load.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image processing technology, and specifically to an image key point smoothing method, an image key point smoothing device, a computer-readable medium, and an electronic device. Background Technology

[0002] With the continuous development of science and technology, the advantages of computer vision technology in daily life are becoming increasingly prominent. Computers can extract relevant features, such as image key points, from related videos or image sequences. Facial key points are a set of coordinate points obtained by analyzing and calculating facial images using a parametric model. However, a series of facial key points acquired in streaming video frames are easily affected by the parametric model, facial movement, and the environment in which the face is located, resulting in a jittery visual effect on the facial key points.

[0003] Currently, among the relevant facial landmark stabilization solutions, either the smoothing effect of facial landmarks is inaccurate, or the computational load is too large, making it unusable on computing platforms with limited computing power. Summary of the Invention

[0004] The purpose of this disclosure is to provide an image keypoint smoothing method, an image keypoint smoothing device, a computer-readable medium, and an electronic device, thereby improving the accuracy of image keypoint smoothing results and reducing computational load to at least a certain extent.

[0005] According to a first aspect of this disclosure, an image keypoint smoothing method is provided, comprising:

[0006] Obtain the key points of the image corresponding to the video frame;

[0007] The key points of the image are grouped to obtain the key point groups corresponding to each video frame;

[0008] The image key point group is standardized to obtain a standard image key point group;

[0009] The distance data of the standard image key point group between different video frames is determined, and the image key points are smoothed according to the distance data to obtain smoothed image key points.

[0010] According to a second aspect of this disclosure, an image keypoint smoothing apparatus is provided, comprising:

[0011] The key point acquisition module is used to acquire the key points of the image corresponding to the video frame;

[0012] The key point grouping module is used to group the key points of the image to obtain the key point groups corresponding to each video frame.

[0013] The key point standardization module is used to standardize the image key point group to obtain a standard image key point group.

[0014] The key point smoothing module is used to determine the distance data of the standard image key point group between different video frames, and to smooth the image key points according to the distance data to obtain smoothed image key points.

[0015] According to a third aspect of this disclosure, a computer-readable medium is provided that stores a computer program thereon, which, when executed by a processor, implements the method described above.

[0016] According to a fourth aspect of this disclosure, an electronic device is provided, characterized in that it comprises:

[0017] Processor; and

[0018] Memory is used to store one or more programs, which, when executed by one or more processors, cause the one or more processors to perform the methods described above.

[0019] One embodiment of this disclosure provides an image keypoint smoothing method that groups image keypoints in video frames to obtain image keypoint groups for each video frame. These image keypoint groups are then standardized to obtain standard image keypoint groups. Distance data between standard image keypoint groups in different video frames can then be determined, and the image keypoints are smoothed based on this distance data to obtain smoothed image keypoints. On one hand, grouping image keypoints and smoothing them on a group-by-group basis considers the local information of keypoints in different regions of the video frame, effectively improving the accuracy of the smoothed image keypoints and resulting in better visual effects. On the other hand, standardizing image keypoint groups eliminates the computational difficulty caused by translation and rotation, effectively reducing the computational load. This allows for use on computing platforms with lower computational performance, broadening its applicability.

[0020] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0021] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0022] Figure 1 A schematic diagram of an exemplary system architecture to which embodiments of the present disclosure may be applied is shown;

[0023] Figure 2 The schematic diagram illustrates a flowchart of an image keypoint smoothing method according to an exemplary embodiment of the present disclosure;

[0024] Figure 3 This schematic diagram illustrates a method for grouping key points in an image according to an exemplary embodiment of the present disclosure.

[0025] Figure 4 This schematically illustrates a process diagram for translating and rotating a group of key points in an image according to an exemplary embodiment of the present disclosure;

[0026] Figure 5 This schematic diagram illustrates the principle of standardizing image key point groups according to an exemplary embodiment of the present disclosure.

[0027] Figure 6 This schematically illustrates a flowchart of a process for determining distance data of a standard image keypoint group between different video frames in an exemplary embodiment of the present disclosure.

[0028] Figure 7 This schematic diagram illustrates the principle of determining distance data of a standard image keypoint group between different video frames in an exemplary embodiment of the present disclosure.

[0029] Figure 8 This schematic diagram illustrates the composition of an image key point smoothing apparatus in an exemplary embodiment of the present disclosure;

[0030] Figure 9 A schematic diagram of an electronic device to which embodiments of the present disclosure may be applied is shown. Detailed Implementation

[0031] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0032] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0033] Figure 1 A schematic diagram of a system architecture for an exemplary application environment in which an image keypoint smoothing method and apparatus according to embodiments of the present disclosure can be applied is shown.

[0034] like Figure 1 As shown, system architecture 100 may include one or more of terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables. Terminal devices 101, 102, and 103 may be various electronic devices with image processing capabilities, including but not limited to desktop computers, portable computers, smartphones, and tablets. It should be understood that... Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, there can be any number of terminal devices, networks, and servers. For example, server 105 could be a server cluster composed of multiple servers.

[0035] The image keypoint smoothing method provided in this embodiment is generally executed by terminal devices 101, 102, and 103, and correspondingly, the image keypoint smoothing device is generally disposed in terminal devices 101, 102, and 103. However, those skilled in the art will readily understand that the image keypoint smoothing method provided in this embodiment can also be executed by server 105, and correspondingly, the image keypoint smoothing device can also be disposed in server 105. This exemplary embodiment does not impose any special limitations on this. For example, in one exemplary embodiment, the image keypoints corresponding to video frames may be uploaded to server 105 by terminal devices 101, 102, and 103. After the server generates smoothed image keypoints using the image keypoint smoothing method provided in this embodiment, it transmits the smoothed image keypoints to terminal devices 101, 102, and 103 for display.

[0036] One related technique proposes using the coordinates of preset facial key points on the t-th frame image in the frame image coordinate system and the face image marked by a face bounding box on the t-th frame image to determine the position of the face image in the (t+1)-th frame. Based on this, a dense optical flow image is calculated, and then the position of the preset facial key points on the (t+1)-th frame image is determined by combining the dense optical flow image. However, this scheme requires the calculation of dense optical flow, which involves a large amount of computation, making it difficult to apply to computing platforms with lower computational requirements, such as mobile terminals, thus limiting its applicability.

[0037] Another related technique proposes determining whether a face in the current video frame has moved relative to the face in the previous video frame. If the face in the current video frame has moved relative to the face in the previous video frame, the facial key points located in the current video frame are determined as valid facial key points in the current video frame. If the face in the current video frame has not moved relative to the face in the previous video frame, the result of a weighted sum of the facial key points located in the current video frame and the corresponding valid facial key points in the previous video frame is determined as the valid facial key points in the current video frame. However, the techniques used in this approach to determine whether there is movement and the weighting strategy are relatively simple, only considering the global changes in the key point positions and lacking the utilization of local information. The accuracy of the key point positions obtained after jitter removal is poor, and it is difficult to handle situations where the face as a whole has been translated or rotated in real-world scenarios.

[0038] Based on one or more problems in related technologies, this disclosure first provides an image key point smoothing method. The following describes the image key point smoothing method of the exemplary embodiment of this disclosure in detail, taking the execution of the method by a terminal device as an example.

[0039] Figure 2 This illustration shows a flowchart of an image keypoint smoothing method according to an exemplary embodiment, which may include the following steps S210 to S240:

[0040] In step S210, the key points of the image corresponding to the video frame are obtained.

[0041] In an exemplary embodiment, a video frame refers to a continuous video stream acquired in any manner. For example, a video frame may be a continuous video stream acquired in real time by the image acquisition module of a terminal device, or a continuous video stream transmitted by other terminal devices through wired or wireless communication. This example embodiment does not impose any special limitation on the source of the video frame.

[0042] Image key points refer to feature key points identified in a video frame in any way. For example, the video frame can be input into a pre-trained artificial intelligence model to obtain the image key points corresponding to the video frame. Alternatively, the image key points corresponding to the video frame can be obtained through image feature descriptors such as scale-invariant feature transform (SIFT) and histogram of oriented gradients (HOG). This example embodiment does not impose any special limitations on this.

[0043] The video frame can be input into the local image feature recognition module to obtain the image key points corresponding to the video frame, or the video frame can be sent to the corresponding server and the image key points corresponding to the video frame can be obtained from the server. There are no special restrictions on the method of obtaining the image key points corresponding to the video frame here.

[0044] In step S220, the image key points are grouped to obtain image key point groups corresponding to each video frame.

[0045] In an exemplary embodiment, grouping refers to the process of dividing multiple image keypoints into different sets according to their correlation in a video frame. For example, a video frame may contain a face image, and the image keypoints may be the facial keypoints corresponding to the face image. Among the facial keypoints, the motion changes of facial keypoints located in different facial features are different. For example, even in the eye area, the upper and lower eyelids have relative motion, and the multiple keypoints corresponding to the upper eyelid have a certain "rigidity" in motion. Therefore, the image keypoints corresponding to the upper eyelid can be considered to be correlated in the video frame, and the image keypoints corresponding to the upper eyelid can be divided into the same set of keypoints. The video frame may also contain other image content with local correlation. For example, the video frame may contain an animal image. The image keypoints corresponding to any limb in the animal image have a certain "rigidity" in motion. Therefore, the image keypoints corresponding to a single limb can be divided into the same set of keypoints. Of course, this is only an illustrative example, and grouping can also be image keypoint groups determined by other division methods. This example embodiment does not specifically limit this.

[0046] Image key points can be grouped using pre-defined grouping data or by using a pre-trained artificial intelligence model. This example embodiment does not impose any special limitations on the grouping method for image key points.

[0047] The same set of image key points in different video frames can be encoded. For example, taking facial key points in a video frame as an example, the same encoding can be set for the upper eyelid key point group in the current video frame and the upper eyelid key point group in the corresponding historical video frame. In this way, in subsequent processing, the same set of upper eyelid key points in the previous and next video frames can be quickly located through encoding, thereby improving the smoothing efficiency of image key points.

[0048] In step S230, the image key point group is standardized to obtain a standard image key point group.

[0049] In an exemplary embodiment, normalization processing refers to the process of eliminating motion interference factors of an image keypoint group in different video frames. For example, normalization processing may involve rotating the entire image keypoint group to eliminate rotation factors in different video frames, or translating the entire image keypoint group to eliminate translation factors in different video frames. Of course, normalization processing may also involve translation, rotation, and scale normalization of the entire image keypoint group, etc., and this example embodiment does not impose any special limitations on this.

[0050] By standardizing the image keypoint group, the image keypoints in the same standard image keypoint group do not need to consider interference factors such as rotation and translation. In the process of smoothing the image keypoints, only the jitter of the image keypoints needs to be considered, which effectively reduces the amount of computation.

[0051] In step S240, distance data of the standard image key point group between different video frames is determined, and the image key points are smoothed according to the distance data to obtain smoothed image key points.

[0052] In an exemplary embodiment, distance data refers to the jitter displacement between different video frames belonging to the same set of standard image keypoints. For example, the distance between corresponding image keypoints in the same set of standard image keypoints can be calculated, and the distance between each image keypoint can be used as the distance data of the standard image keypoints between different video frames. Alternatively, the average distance between each image keypoint can be used as the distance data of the standard image keypoints between different video frames. This example embodiment does not impose any special limitations on this.

[0053] After obtaining the distance data of standard image keypoint groups between different video frames, the degree of deformation of the standard image keypoint groups in the preceding and following video frames can be determined based on this distance data. For example, a pre-set distance threshold can be obtained. If the distance data is determined to be less than the distance threshold, it can be considered that the degree of deformation of the standard image keypoint groups in the preceding and following video frames is small. In this case, the average value of the position coordinates of two image keypoints at the same point in the standard image keypoint group can be used as the position coordinates of the smoothed image keypoint in the current video frame. Then, the reverse operation of the standardization process can be performed on the image keypoint to obtain the final smoothed image keypoint corresponding to the current video frame. If the distance data is determined to be greater than or equal to the distance threshold, it can be considered that the degree of deformation of the standard image keypoint groups in the preceding and following video frames is large. In this case, the standard image keypoint group of the current video frame can be directly used as the smoothed result. Then, the reverse operation of the standardization process can be performed on the standard image keypoint group of the current video frame to obtain the final smoothed image keypoint corresponding to the current video frame.

[0054] By grouping keypoints in video frames, keypoint groups are obtained for each video frame. These keypoint groups are then standardized to obtain standard keypoint groups. Distance data between these standard keypoint groups across different video frames can then be determined. Based on this distance data, smoothed keypoints are obtained. On one hand, grouping keypoints and smoothing them on a group basis considers the local information of keypoints in different regions of the video frame, effectively improving the accuracy of the smoothed keypoints and resulting in better visual effects. On the other hand, standardizing keypoint groups eliminates computational difficulties caused by translation and rotation, effectively reducing computational load and allowing for use on computing platforms with lower performance, thus broadening its applicability.

[0055] Steps S210 to S240 will be described in detail below.

[0056] In one exemplary embodiment, a video frame may contain a face image, and the image key points can be the facial key points corresponding to the face image. Of course, a video frame may also contain other locally related image content, such as animal images or vehicle images. In this case, the image key points can be the animal image key points corresponding to the animal image or the vehicle image key points corresponding to the vehicle image. This example embodiment does not impose any special limitations on this. For ease of understanding, the following description uses facial key points as an example. It is readily understood by those skilled in the art that this example uses facial key points and does not imply that this example embodiment can only be used for facial key points, nor should it impose any special limitations on this example embodiment.

[0057] Optionally, the key points of the image can be grouped to obtain the key point group corresponding to each video frame. Specifically, facial feature data can be obtained, and the key points of the face can be grouped according to the facial feature data to obtain the key point group corresponding to each video frame.

[0058] Among them, facial feature data refers to pre-set data used to divide facial key points. For example, facial feature data can be grouped based on eyes, eyebrows, nose, mouth, and ears; facial feature data can also be grouped based on eyebrows, upper eyelids, lower eyelids, upper lip, lower lip, and nose; of course, facial feature data can also be grouped based on other division methods. The specific settings can be customized according to the actual situation. This example embodiment does not impose any special limitations on this.

[0059] Taking facial feature data grouped based on eyebrows, upper eyelids, lower eyelids, upper lip, lower lip, and nose as an example, the resulting image keypoint groups can include eyebrow keypoint groups, upper eyelid keypoint groups, lower eyelid keypoint groups, upper lip keypoint groups, lower lip keypoint groups, and nose keypoint groups. Grouping facial keypoints using pre-defined facial feature data further reduces computational load compared to other grouping methods.

[0060] It is understandable that this example only uses facial key points as an example. If the video frame contains other image content, when grouping the image key points, the corresponding grouping criteria can be set based on the image features in the image content. For example, if the image content is an animal image, the grouping criteria can be set based on the animal features in the animal image. This will not be elaborated on here.

[0061] It is understood that in this embodiment, a neural network model for grouping facial key points can be pre-trained by using eyebrows, upper eyelids, lower eyelids, upper lip, lower lip, and nose as grouping criteria, and the key points of the image can be grouped using this neural network model. This example embodiment does not impose any special limitations on this.

[0062] Figure 3 This illustration schematically shows a diagram of grouping key points in an image according to an exemplary embodiment of the present disclosure.

[0063] refer to Figure 3As shown, assuming the current video frame can contain a face image, the corresponding image keypoints of the current video frame can be face keypoints 310. Specifically, face keypoints 310 can be grouped according to preset facial feature data to obtain eyebrow keypoint group 320, upper eyelid keypoint group (left and right) 330, lower eyelid keypoint group (left and right) 340, nose keypoint group 350, upper lip keypoint group 360, and lower lip keypoint group 370. As those skilled in the art can easily understand, due to motion characteristics, the image keypoints in eyebrow keypoint group 320, upper eyelid keypoint group (left and right) 330, lower eyelid keypoint group (left and right) 340, nose keypoint group 350, upper lip keypoint group 360, and lower lip keypoint group 370 are basically relatively rigid and can be considered as rigid line segments. Therefore, multiple image keypoints in each divided image keypoint group can be processed uniformly as a whole, effectively reducing the computational load.

[0064] In an exemplary embodiment, the standardization of image keypoint groups can be achieved through the following steps to obtain a standard image keypoint group:

[0065] The keypoint group of an image can be translated and rotated to obtain the translated and rotated keypoint group, and then the translated and rotated keypoint group can be normalized to obtain the standard keypoint group.

[0066] Translation and rotation processing refers to the process of translating and rotating image key point groups. For example, the image key point group can be rotated first, and then the rotated image key point group can be translated to eliminate the influence of the rotation angle and translation amount of the image key point group between different video frames on the smoothing result. Of course, the image key point group can be translated first, and then the translated image key point group can be rotated. This example embodiment is not limited to this.

[0067] Normalization refers to the process of unifying the scale of each key point in an image key point group to a certain range in order to eliminate the influence of the scale of the image key point group between different video frames on the smoothing result. For example, the image key point group between different video frames can be normalized to a scale of (0, 1), or the image key point group between different video frames can be normalized to a scale of (-1, 1). This example embodiment does not make any special limitation on this.

[0068] Optional, can be done through Figure 4 The steps described in the document implement translation and rotation processing of the image keypoint group to obtain the translated and rotated image keypoint group. (Refer to...) Figure 4 As shown, it can specifically include:

[0069] Step S410: Determine the rotation angle and center offset corresponding to the image key point group;

[0070] Step S420: Rotate the image key point group according to the rotation angle;

[0071] Step S430: The rotated image key point group is translated according to the center offset to obtain the translated and rotated image key point group.

[0072] The rotation angle refers to the minimum angle between the relatively rigid line segment formed by the keypoint group and the horizontal direction. It can be calculated by taking the line connecting the first and last keypoints in the keypoint group as the main direction of the entire keypoint group, and then calculating the minimum angle between this main direction and the horizontal direction as the corresponding rotation angle. Since the keypoints in a keypoint group are generally not straight lines, approximating the overall direction of the keypoint group as the line connecting the first and last keypoints further reduces the computational load and improves the efficiency of smoothing results.

[0073] After calculating the rotation angle, the image keypoint group can be rotated as a whole to obtain the rotated image keypoint group. In this way, the main direction of the image keypoint group in different video frames is horizontal. When measuring keypoint jitter or calculating smoothing results, the rotation of keypoints does not need to be considered, effectively reducing the amount of calculation.

[0074] Center offset refers to the distance from the center point of a relatively rigid line segment formed by a group of image key points to the origin of the standardized coordinate system. The mean coordinate of the position coordinates of all image key points in the image key point group can be used as the center point position coordinate of the image key point group. Then, the distance between the center point position coordinate and the origin coordinate can be used as the center offset of the image key point group.

[0075] After calculating the center offset, all image keypoints in the rotated image keypoint group can be moved using this center offset so that the position coordinates of all image keypoints in the image keypoint group fall within the specified range. In this way, when measuring keypoint jitter or calculating smoothing results, the translation problem of keypoints does not need to be considered, effectively reducing the amount of calculation.

[0076] By performing translation and rotation processing on key point groups in images, potential translation and rotation issues may exist for key point groups belonging to the same group in different video frames. This eliminates the need to consider translation and rotation issues when measuring key point jitter or calculating smoothing results, effectively reducing computational complexity and workload, eliminating the computational performance limitations of application platforms, and expanding the applicability range.

[0077] Optionally, the normalization of the key point group of the translated and rotated image can be achieved through the following steps:

[0078] The scale data corresponding to the key point group of the image can be determined, and the coordinates of the key points of the translated and rotated image can be normalized according to the scale data to obtain the standard key point group of the image.

[0079] Here, scale data refers to data that measures the length of the relatively rigid line segment corresponding to the key point group of an image. For example, scale data can be the distance between key point groups of an image, or the total number of pixels occupied by key point groups of an image. Of course, it can also be other types of data that can measure the length of the relatively rigid line segment corresponding to key point groups of an image. This example embodiment does not make any special limitations on this.

[0080] The coordinates of key points in the translated and rotated image can be normalized based on the scale data, unifying the coordinates of key points in each key point group to a certain range, thereby reducing the amount of computation and improving computational efficiency.

[0081] Figure 5 This illustration schematically demonstrates a principle diagram for standardizing image key point groups in an exemplary embodiment of the present disclosure.

[0082] refer to Figure 5 As shown, taking the upper eyelid keypoint group 501 as an example, the process of standardizing the image keypoint group to obtain a standard image keypoint group is explained:

[0083] Step S510: The line connecting the first and last image key points in the upper eyelid key point group 501 can be taken as the main direction of the entire image key point group. Then, the minimum angle between the main direction and the horizontal direction is calculated as the rotation angle 502 corresponding to the upper eyelid key point group 501.

[0084] Step S520: Based on the determined rotation angle 502, rotate the upper eyelid key point group 501 as a whole in a clockwise direction to obtain the rotated image key point group 503.

[0085] In step S530, the mean coordinates of the position coordinates of all image key points in the upper eyelid key point group 501 can be used as the center point position coordinates of the upper eyelid key point group 501, and then the distance between the center point position coordinates and the origin coordinates can be used as the center offset 504 corresponding to the image key point group.

[0086] Step S540: Move all image key points in the rotated image key point group 503 by center offset 504 so that the position coordinates of all image key points in the rotated image key point group 503 fall within the specified range, and obtain the rotated and translated image key point group 505.

[0087] Step S550: The scale data corresponding to the image key point group 505 after rotation and translation can be determined. For example, the scale data corresponding to the image key point group can be represented by the Euclidean distance between the last point and the first point. The coordinates of the image key point group 505 after rotation and translation are normalized according to the scale data to obtain the standard image key point group 506.

[0088] It should be noted that, Figure 5 The standardized processing procedures described are merely illustrative examples and should not impose any special limitations on this example embodiment.

[0089] In one exemplary embodiment, it can be achieved through Figure 6 The steps in the document determine the distance data of standard image keypoint groups between different video frames, referencing... Figure 6 As shown, it can specifically include:

[0090] Step S610: Determine the first standard image key point group in the current video frame, and determine the second standard image key point group corresponding to the first standard image key point group in the historical video frames.

[0091] Step S620: Remap the key points in the first standard image key point group to the second standard image key point group, and calculate the first distance data from the first standard image key point group to the second standard image key point group;

[0092] Step S630: Remap the key points in the second standard image key point group to the first standard image key point group, and calculate the second distance data from the second standard image key point group to the first standard image key point group;

[0093] Step S640: Determine the distance data of the standard image keypoint group between different video frames based on the first distance data and the second distance data.

[0094] The first standard image key point group refers to a group of standard image key points in the current video frame. For example, the first standard image key point group may be the standard image key point group corresponding to the upper eyelid in the current video frame. The second standard image key point group refers to the standard image key point group corresponding to the first standard image key point group in the historical video frame corresponding to the current video frame. For example, the second standard image key point group may be the standard image key point group corresponding to the upper eyelid in the historical video frame.

[0095] It should be noted that in this embodiment, the "first standard image key point group" and the "second standard image key point group" are only used to distinguish the standard image key point groups that are in the same group in the current video frame and the historical video frames. They have no special meaning and should not impose any special limitations on this example embodiment.

[0096] The first distance data refers to the sum of the distances between key points in the first standard image key point group and key points in the second standard image key point group after remapping key points in the first standard image key point group to the second standard image key point group (of course, it can also be the maximum distance, average distance, etc. between key points in the first standard image key point group and key points in the second standard image key point group, without special limitation here); the second distance data refers to the sum of the distances between key points in the second standard image key point group and key points in the first standard image key point group after remapping key points in the second standard image key point group to the first standard image key point group (of course, it can also be the maximum distance, average distance, etc. between key points in the second standard image key point group and key points in the first standard image key point group, without special limitation here).

[0097] It will be readily understood by those skilled in the art that the first distance data here is not equal to the second distance data. If the distance between the first standard image key point group and the second standard image key point group is directly calculated, such calculation may introduce errors due to the different horizontal coordinates. By remapping the standard image key point groups in different video frames and determining the distance data between the standard image key point groups in different video frames based on the first distance data and the second distance data, the accuracy of the distance data can be effectively improved, thereby improving the accuracy of the smoothing result and improving the visual appearance of the smoothed image key points in the video frame.

[0098] Figure 7 The illustration shows a schematic diagram of the principle of determining distance data of a standard image keypoint group between different video frames in an exemplary embodiment of the present disclosure.

[0099] refer to Figure 7As shown, for the current video frame (such as the t-th frame in a continuous video stream), the distance data between the image key point group in the current video frame and the image key point group belonging to the same group in the historical video frames is calculated. The specific calculation method is as follows: Assume that the black dots represent the image key point positions in the historical video frames (such as the t-1-th frame in a continuous video stream), and the white dots represent the image key point positions in the current video frame. The distance between these two sets of image keypoints can be defined as max{d(P(t), P(t-1)), d(P(t-1), P(t))}, where d(P(t), P(t-1)) represents the distance from the first standard image keypoint set P(t) of the current video frame to the second standard image keypoint set P(t-1) of the historical video frame, and d(P(t-1), P(t)) represents the distance from the second standard image keypoint set P(t-1) to the first standard image keypoint set P(t). It is important to note that the distance data here does not satisfy the commutative law; that is, d(P(t), P(t-1)) is not actually equal to d(P(t-1), P(t)). This is because when calculating the distance between the two sets of standard image keypoints, it is necessary to remap the position coordinates of the image keypoints in one set to the other set.

[0100] Optionally, the maximum value of the two distance data is used as the distance data between the first standard image key point group and the second standard image key point group. Of course, the average value of the two distance data can also be used as the distance data between the first standard image key point group and the second standard image key point group, which is also within the protection scope of this embodiment.

[0101] Taking the calculation of d(P(t-1), P(t)) as an example, we don't directly calculate the Euclidean distance between the corresponding image keypoints and then sum them, because this would introduce errors due to different x-coordinates. Here, we use the calculation... Figure 7Taking the distance between keypoints in image 1 as an example, since the x-coordinates of point 1 in the two sets of standard image keypoints are different, point 1 is first aligned. In d(P(t-1), P(t)), the x-coordinate of the keypoint in P(t) is used as the reference to find the corresponding x-coordinate value of that point in the P(t-1) sequence. A simple way is to use points 0 and 2 in P(t-1) as an approximation, that is, to use linear interpolation to obtain the coordinates P(t-1, y1) in P(t-1) whose x-coordinate is the x-coordinate value of point 1 in P(t). Then, the absolute value of the distance between P(t-1, y1) and P(t, y2) is calculated, which is the distance between these two points. After calculating the distance between image keypoints, the distance between the two sets of standard image keypoints is obtained by summing all the distances. Similarly, calculating d(P(t), P(t-1)) follows a similar approach, the only difference being that it is mapped to the x-coordinate of P(t-1), which will not be elaborated here.

[0102] Optionally, linear interpolation is used when remapping points based on the x-coordinate. Of course, non-linear interpolation can also be used for point remapping. This example embodiment does not impose any special limitations on this.

[0103] Optionally, when calculating the distance data between the first set of key points in the first standard image and the second set of key points in the second standard image, the overall approach can be to first use polynomial fitting and then compare the differences in various parameters, or after polynomial fitting, uniformly sample within the x range and then calculate the differences between the obtained y values. In this case, the distance between the key points in the first set of key points in the first standard image and the key points in the second set of key points in the second standard image can satisfy the commutative law.

[0104] Optionally, if the distance data of the standard image keypoint group between different video frames is determined to be less than a preset distance threshold, it can be considered that the deformation degree of the standard image keypoint group in the preceding and following video frames is small. At this time, the keypoint positions in the first standard image keypoint group and the keypoint positions in the second standard image keypoint group can be averaged to obtain a smoothed standard image keypoint group. Then, the smoothed standard image keypoint group can be denormalized to obtain the smoothed image keypoints in the current video frame.

[0105] In this context, denormalization refers to the restoration process after normalization. For example, during normalization, image keypoint groups can be rotated in the positive direction based on rotation angle, translated in the positive direction based on center offset, and normalized based on scale data. Denormalization can be rotated in the opposite direction based on rotation angle, translated in the opposite direction based on center offset, and scaled back based on scale data, ultimately obtaining the image keypoints in the video frame coordinate system.

[0106] Optionally, if the distance data of the standard image key point group between different video frames is determined to be greater than or equal to the preset distance threshold, it can be considered that the deformation of the standard image key point group in the preceding and following video frames is large. In this case, it is impossible to smooth the image key points in the current video frame, and the original image key points in the current video frame can be directly used as the output image key points.

[0107] In summary, this exemplary embodiment allows for the grouping of image keypoints in video frames to obtain image keypoint groups corresponding to each video frame. These image keypoint groups are then standardized to obtain standard image keypoint groups. This enables the determination of distance data between standard image keypoint groups across different video frames, and the image keypoints are then smoothed based on this distance data to obtain smoothed image keypoints. On one hand, grouping image keypoints and smoothing them on a group basis considers the local information of keypoints in different regions of the video frame, effectively improving the accuracy of the smoothed image keypoints and resulting in better visual effects. On the other hand, standardizing image keypoint groups eliminates the computational difficulty caused by translation and rotation, effectively reducing the computational load. This allows for use on computing platforms with lower performance, broadening its applicability.

[0108] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0109] Further reference Figure 8 As shown, this example embodiment also provides an image keypoint smoothing device 800, including a keypoint acquisition module 810, a keypoint grouping module 820, a keypoint normalization module 830, and a keypoint smoothing module 840. Wherein:

[0110] The key point acquisition module 810 is used to acquire the key points of the image corresponding to the video frame;

[0111] The key point grouping module 820 is used to group the key points of the image to obtain the key point group corresponding to each video frame;

[0112] The key point standardization module 830 is used to standardize the image key point group to obtain a standard image key point group;

[0113] The key point smoothing module 840 is used to determine the distance data of the standard image key point group between different video frames, and to smooth the image key points according to the distance data to obtain smoothed image key points.

[0114] In one exemplary embodiment, the key point standardization module 830 may include:

[0115] The translation and rotation unit can be used to perform translation and rotation processing on the image key point group to obtain the translated and rotated image key point group;

[0116] The scale normalization unit can be used to normalize the translated and rotated image key point group to obtain a standard image key point group.

[0117] In one exemplary embodiment, the translation and rotation unit can be used for:

[0118] Determine the rotation angle and center offset corresponding to the key point group of the image;

[0119] The image key point group is rotated according to the rotation angle;

[0120] The rotated image key point group is translated according to the center offset to obtain the translated and rotated image key point group.

[0121] In one exemplary embodiment, the scale normalization unit can be used to:

[0122] Determine the scale data corresponding to the image key point group;

[0123] The coordinates of the translated and rotated image keypoint group are normalized based on the scale data to obtain a standard image keypoint group.

[0124] In one exemplary embodiment, image key points may include facial key points, and the key point grouping module 820 may be used for:

[0125] Obtain facial feature data;

[0126] The facial key points are grouped according to the facial feature data to obtain the image key point group corresponding to each video frame.

[0127] The image key point group includes at least one of the following: eyebrow key point group, upper eyelid key point group, lower eyelid key point group, upper lip key point group, lower lip key point group, and nose key point group.

[0128] In one exemplary embodiment, the key point smoothing module 840 can be used to:

[0129] Determine the first standard image key point group in the current video frame, and determine the second standard image key point group corresponding to the first standard image key point group in the historical video frames;

[0130] Remap the key points in the first standard image key point group to the second standard image key point group, and calculate the first distance data from the first standard image key point group to the second standard image key point group;

[0131] Remap the key points in the second standard image key point group to the first standard image key point group, and calculate the second distance data from the second standard image key point group to the first standard image key point group;

[0132] Distance data for the standard image keypoint group between different video frames is determined based on the first distance data and the second distance data.

[0133] In one exemplary embodiment, the key point smoothing module 840 can be used to:

[0134] If it is determined that the distance data is less than a preset distance threshold, the key point positions in the first standard image key point group and the key point positions in the second standard image key point group are averaged to obtain a smoothed standard image key point group.

[0135] The smoothed standard image key point group is denormalized to obtain the smoothed image key points in the current video frame.

[0136] The specific details of each module in the above-mentioned device have been described in detail in the method section of the implementation. For any undisclosed details, please refer to the implementation content of the method section, and therefore will not be repeated here.

[0137] Those skilled in the art will understand that various aspects of this disclosure can be implemented as a system, method, or program product. Therefore, various aspects of this disclosure can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software aspects, collectively referred to herein as a "circuit," "module," or "system."

[0138] Exemplary embodiments of this disclosure provide an electronic device for implementing an image keypoint smoothing method, which may be... Figure 1 The terminal devices 101, 102, 103, or server 105 are included. The electronic device includes at least a processor and a memory, the memory being used to store executable instructions of the processor, the processor being configured to perform an image keypoint smoothing method by executing the executable instructions.

[0139] The following is based on Figure 9 Taking the electronic device 900 as an example, the construction of the electronic device in this disclosure will be described by way of example. Figure 9 The electronic device 900 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0140] like Figure 9 As shown, the electronic device 900 is presented in the form of a general-purpose computing device. The components of the electronic device 900 may include, but are not limited to: at least one processing unit 910, at least one storage unit 920, a bus 930 connecting different system components (including storage unit 920 and processing unit 910), and a display unit 940.

[0141] The storage unit 920 stores program code, which can be executed by the processing unit 910, causing the processing unit 910 to perform the image key point smoothing method described in this specification.

[0142] Storage unit 920 may include readable media in the form of volatile storage units, such as random access memory (RAM) 921 and / or cache memory 922, and may further include read-only memory (ROM) 923.

[0143] Storage unit 920 may also include a program / utility 924 having a set (at least one) program module 925, such program module 925 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0144] Bus 930 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0145] Electronic device 900 can also communicate with one or more external devices 970 (e.g., sensor devices, Bluetooth devices, etc.), and with one or more devices that enable users to interact with electronic device 900, and / or with any device that enables electronic device 900 to communicate with one or more other computing devices (e.g., routers, modems, etc.). This communication can be performed via input / output (I / O) interface 950. Furthermore, electronic device 900 can also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via network adapter 960. As shown, network adapter 960 communicates with other modules of electronic device 900 via bus 930. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, data backup storage systems, and sensor modules (e.g., gyroscope sensors, magnetometers, accelerometers, distance sensors, proximity sensors, etc.).

[0146] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0147] Exemplary embodiments of this disclosure also provide a computer-readable storage medium having a program product stored thereon capable of implementing the methods described above in this specification. In some possible embodiments, various aspects of this disclosure may also be implemented as a program product including program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure.

[0148] It should be noted that the computer-readable medium disclosed herein may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0149] In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can transmit, propagate, or transfer a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.

[0150] Furthermore, program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0151] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0152] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A method for smoothing key points in an image, characterized in that, include: Obtain the key points of the image corresponding to the video frame; The key points of the image are grouped to obtain the key point groups corresponding to each video frame; The image key point group is standardized to obtain a standard image key point group; Determine the distance data between the standard image key point groups between different video frames, and smooth the image key points according to the distance data to obtain smoothed image key points; The step of determining the distance data of the standard image keypoint group between different video frames includes: Determine the first standard image key point group in the current video frame, and determine the second standard image key point group corresponding to the first standard image key point group in the historical video frames; Remap the key points in the first standard image key point group to the second standard image key point group, and calculate the first distance data from the first standard image key point group to the second standard image key point group; The key points in the second standard image key point group are remapped to the first standard image key point group, and the second distance data from the second standard image key point group to the first standard image key point group is calculated; the distance data between the standard image key point groups of different video frames is determined based on the first distance data and the second distance data. The step of smoothing the image key points based on the distance data to obtain smoothed image key points includes: If the distance data is determined to be less than a preset distance threshold, the key point positions in the first standard image key point group and the key point positions in the second standard image key point group are averaged to obtain a smoothed standard image key point group; the smoothed standard image key point group is then denormalized to obtain the smoothed image key points in the current video frame. If the distance data is determined to be greater than or equal to a preset distance threshold, then the standard image key point group of the current video frame is used as the smoothed result, and the standard image key point group of the current video frame is denormalized to obtain the smoothed image key points corresponding to the current video frame.

2. The method according to claim 1, characterized in that, The step of standardizing the image keypoint group to obtain a standard image keypoint group includes: The image key point group is translated and rotated to obtain the translated and rotated image key point group; The translated and rotated image key point group is normalized to obtain a standard image key point group.

3. The method according to claim 2, characterized in that, The step of performing translation and rotation processing on the image keypoint group to obtain the translated and rotated image keypoint group includes: Determine the rotation angle and center offset corresponding to the key point group of the image; The image key point group is rotated according to the rotation angle; The rotated image key point group is translated according to the center offset to obtain the translated and rotated image key point group.

4. The method according to claim 2, characterized in that, The normalization process for the translated and rotated image keypoint group to obtain a standard image keypoint group includes: Determine the scale data corresponding to the image key point group; The coordinates of the translated and rotated image keypoint group are normalized based on the scale data to obtain a standard image keypoint group.

5. The method according to claim 1, characterized in that, The image key points include facial key points. The process of grouping the image key points to obtain image key point groups corresponding to each video frame includes: Obtain facial feature data; The facial key points are grouped according to the facial feature data to obtain the image key point group corresponding to each video frame. The image key point group includes at least one of the following: eyebrow key point group, upper eyelid key point group, lower eyelid key point group, upper lip key point group, lower lip key point group, and nose key point group.

6. An image key point smoothing device, characterized in that, include: The key point acquisition module is used to acquire the key points of the image corresponding to the video frame; The key point grouping module is used to group the key points of the image to obtain the key point groups corresponding to each video frame. The key point standardization module is used to standardize the image key point group to obtain a standard image key point group. A key point smoothing module is used to determine the distance data of the standard image key point group between different video frames, and to smooth the image key points according to the distance data to obtain smoothed image key points; The step of determining the distance data of the standard image keypoint group between different video frames includes: A first standard image keypoint group is determined in the current video frame, and a second standard image keypoint group corresponding to the first standard image keypoint group in historical video frames is determined; keypoints in the first standard image keypoint group are remapped to the second standard image keypoint group, and a first distance data from the first standard image keypoint group to the second standard image keypoint group is calculated; keypoints in the second standard image keypoint group are remapped to the first standard image keypoint group, and a second distance data from the second standard image keypoint group to the first standard image keypoint group is calculated; distance data between the standard image keypoint groups in different video frames is determined based on the first distance data and the second distance data. The step of smoothing the image key points based on the distance data to obtain smoothed image key points includes: If the distance data is determined to be less than a preset distance threshold, the key point positions in the first standard image key point group and the key point positions in the second standard image key point group are averaged to obtain a smoothed standard image key point group; the smoothed standard image key point group is then denormalized to obtain the smoothed image key points in the current video frame. If the distance data is determined to be greater than or equal to a preset distance threshold, then the standard image key point group of the current video frame is used as the smoothed result, and the standard image key point group of the current video frame is denormalized to obtain the smoothed image key points corresponding to the current video frame.

7. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 5.

8. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the method of any one of claims 1 to 5 by executing the executable instructions.