Method and system for predicting aging effect BAED on face dynamics
Patent Information
- Application Number
- PCT/CN2025/081505
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2026-09-17
Smart Images

Figure CN2025081505_17092026_PF_FP_ABST
Abstract
Description
METHOD AND SYSTEM FOR PREDICTING AGING EFFECT BAED ON FACE DYNAMICSFIELD OF THE INVENTION
[0001] The present disclosure relates to the field of image processing. More specifically, the disclosure relates to a method and system for predicting an aging effect based on face dynamics.BACKGROUND OF THE INVENTION
[0002] Aging is a natural process that affects individuals at varying degrees and rates. People are increasingly aware of the importance of taking measures in advance to combat aging. In order to effectively slow down the aging process, it is desirable to be able to predict the facial zones that are likely to show aging signs / effects before they appear or become obvious, and to use corresponding anti-aging product and / or treatment accordingly. Current methods mainly predict an aging effect of a user by analyzing static images of the user’s face. These approaches focus on the static features of the face, such as skin texture and pigmentation. However, they ignore dynamic features of the face, which are also very important for predicting the aging effect.
[0003] People make a variety of different expressions in their daily lives (e.g., smile, surprise, sadness, etc. ) . Changes in facial expressions result in a certain amount of squeezing or stretching of facial zones. Studies have shown that if a facial zone is squeezed or stretched more frequently or to a greater extent during different facial expressions, that facial zone is more likely to have an aging effect (e.g., wrinkle growing or sagging) . Unfortunately, these dynamic changes in the face have not been adequately considered in current methods for aging effect prediction.
[0004] Therefore, there is a need to predict an aging effect based on face dynamics to further improve the accuracy of prediction.SUMMARY OF THE INVENTION
[0005] The summary is provided to introduce a selection of concepts in a simplified form that are further described below in detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.
[0006] Various aspects and features of the disclosure are described in further detail below.
[0007] According to a first aspect of the disclosure, there is provided a method for predicting an aging effect based on face dynamics, comprising: acquiring a set of face images of a user, wherein the set of face images comprises a neutral image showing no expression and one or more expressive images, each expressive image showing a specific expression; generating a face mesh based on a detected face of the user in the set of face images, wherein the face mesh comprises a set of landmark points indicating one or more key face features of the user; selecting one or more sub-regions from the face mesh; determining a face dynamic parameter of each sub-region, wherein the face dynamic parameter of each sub-region indicates dynamic change of the sub-region; and predicting an aging effect of the user based on the face dynamic parameters of the one or more sub-regions.
[0008] In an embodiment, the aging effect comprises at least one of wrinkle growing, sagging, losing radiance, losing smoothness, or losing plainness.
[0009] In an embodiment, acquiring the set of face images further comprises: capturing multiple images or one or more videos of the user that show a variety of expressions as well as no expression; and extracting the set of face images based on the captured images or videos.
[0010] In an embodiment, the key face features comprise eyes, eyebrows, mouth, and nose of the user.
[0011] In an embodiment, selecting one or more sub-regions from the face mesh further comprises: selecting a subset of landmark points from the set of landmark points; and obtaining the one or more sub-regions based on the subset of landmark points.
[0012] In an embodiment, determining a face dynamic parameter of each sub-region further comprises: determining a geometric parameter of the sub-region, wherein the geometric parameter comprises at least one of an area, a perimeter, or a centroid of the sub-region; and determining the face dynamic parameter based on a variation of the geometric parameter of the sub-region.
[0013] In an embodiment, the variation of the geometric parameter of the sub-region is determined by comparing values of the geometric parameter of the sub-region from the one or more expressive images with a value of the geometric parameter of the sub-region from the neutral image.
[0014] In an embodiment, a result of the prediction indicates a possibility of having an aging effect in each sub-region, wherein a higher value of the face dynamic parameter of a sub-region indicates a higher possibility of having an aging effect in the sub-region, and a lower value of the face dynamic parameter of a sub-region indicates a lower possibility of having an aging effect in the sub-region.
[0015] In an embodiment, the method further comprises displaying one image of the set of face images and a result of the prediction associated with the one or more sub-regions on a display device.
[0016] In an embodiment, the method further comprises: recommending an anti-aging product and / or treatment based on a result of the prediction.
[0017] In an embodiment, the method further comprises: inputting the set of face images to an aging scoring model to obtain a predicted score, wherein the predicted score describes one or more aging signs on the user’s face, and wherein the aging effect of the user is further predicted based on the predicted score.
[0018] According to a second aspect of the disclosure, there is provided a system for predicting an aging effect based on face dynamics, comprising: an image acquisition unit configured to acquire a set of face images of a user, wherein the set of face images comprises a neutral image showing no expression and one or more expressive images, each expressive image showing a specific expression; a face mesh generation unit configured to generate a face mesh based on a detected face of the user in the set of face images, wherein the face mesh comprises a set of landmark points indicating one or more key face features of the user; a parameter determination unit configured to: select one or more sub-regions from the face mesh; determine a face dynamic parameter of each sub-region, wherein the face dynamic parameter of each sub-region indicates dynamic change of the sub-region; and an aging effect prediction unit configured to predict an aging effect of the user based on the face dynamic parameters of the one or more sub-regions.
[0019] In an embodiment, the parameter determination unit is further configured to determine a face dynamic parameter of each sub-region by: determining a geometric parameter of the sub-region, wherein the geometric parameter comprises at least one of an area, a perimeter, or a centroid of the sub-region; and determining the face dynamic parameter based on a variation of the geometric parameter of the sub-region.
[0020] In an embodiment, the system further comprising a post-processing unit configured to: display one image of the set of face images and a result of the prediction associated with the one or more sub-regions on a display device; and / or recommend an anti-aging product and / or treatment based on a result of the prediction.
[0021] According to a third aspect of the disclosure, there is provided a non-transitory computer-readable medium having computer-executable instructions stored thereon that, in response to execution by one or more processors of a computer system, cause the computer system to perform operations comprising: acquiring a set of face images of a user, wherein the set of face images comprises a neutral image showing no expression and one or more expressive images, each expressive image showing a specific expression; generating a face mesh based on a detected face of the user in the set of face images, wherein the face mesh comprises a set of landmark points indicating one or more key face features of the user; selecting one or more sub-regions from the face mesh; determining a face dynamic parameter of each sub-region, wherein the face dynamic parameter of each sub-region indicates dynamic change of the sub-region; and predicting an aging effect of the user based on the face dynamic parameters of the one or more sub-regions.BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The above and other aspects, features, and benefits of various embodiments of the disclosure will become more fully apparent, by way of example, from the following detailed description with reference to the accompanying drawings, in which like reference numerals or letters are used to designate like or equivalent elements. The drawings are illustrated for facilitating better understanding of the embodiments of the disclosure and not necessarily drawn to scale, in which:
[0023] FIG. 1 illustrates a system architecture diagram for predicting an aging effect based on face dynamics according to various aspects of the present disclosure;
[0024] FIG. 2 illustrates a flowchart of a method for predicting an aging effect based on face dynamics according to various aspects of the present disclosure;
[0025] FIG. 3 illustrates a simplified flowchart of a method for predicting an aging effect based on face dynamics in combination with an aging scoring model according to various aspects of the present disclosure;
[0026] FIG. 4 illustrates a schematic diagram showing examples of face landmarks according to various aspects of the present disclosure;
[0027] FIG. 5 illustrates a schematic diagram showing a set of face images of a user as well as dynamic change of the sub-regions according to various aspects of the present disclosure;
[0028] FIG. 6 illustrates a schematic diagram of an exemplary process for predicting an aging effect based on face dynamics according to various aspects of the present disclosure; and
[0029] FIG. 7 illustrates a block diagram of a system for predicting an aging effect based on face dynamics according to various aspects of the present disclosure.DETAILED DESCRIPTION
[0030] Embodiments herein will be described in detail hereinafter with reference to the accompanying drawings, in which embodiments are shown. These embodiments herein may, however, be embodied in many different forms and should not be construed as being limited to the embodiments set forth herein. The elements of the drawings are not necessarily to scale relative to each other. Like numbers refer to like elements throughout.
[0031] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms "a" , "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" , "comprising" , "includes" and / or "including" when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0032] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meanings as commonly understood. It will be further understood that a term used herein should be interpreted as having a meaning consistent with its meaning in the context of this specification and the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0033] The present technology is described below with reference to block diagrams and / or flowchart illustrations of methods, apparatus (systems) and / or computer program products according to the present embodiments. It is understood that blocks of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, may be implemented by computer program instructions. These computer program instructions may be provided to a processor, controller or controlling unit of a general-purpose computer, special purpose computer, and / or other programmable data processing apparatus, such that the instructions, which execute via the processor of the computer and / or other programmable data processing apparatus, create means for implementing the functions / acts specified in the block diagrams and / or flowchart block or blocks.
[0034] Accordingly, the present technology may be embodied in hardware and / or in software (including firmware, resident software, micro-code, etc. ) . Furthermore, the present technology may take the form of a computer program product on a computer-usable or computer-readable storage medium having computer-usable or computer-readable program code embodied in the medium for use by or in connection with an instruction execution system. In the context of this document, a computer-usable or computer-readable medium may be any medium that may contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.
[0035] Embodiments herein will be described below with reference to the drawings.
[0036] Current methods for predicting an aging effect have not adequately considered dynamic changes in the face. The present disclosure introduces a novel approach that predicts the aging effect by numerically quantifies these dynamic changes. The technical solution of the present disclosure selects one or more sub-regions based on a user′s face mesh, quantifies the face dynamics when the user′sexpression changes, and predicts an aging effect of the user based on the face dynamics. By fully considering the dynamic changes in the face between specific expressions and no expression, this approach improves the accuracy of aging effect prediction.
[0037] FIG. 1 illustrates a diagram of a system architecture 100 for predicting an aging effect based on face dynamics according to various aspects of the present disclosure.
[0038] The system architecture 100 first acquires multiple images or one or more videos of a user (102) . The images or videos of the user show a variety of expressions of the user as well as no expression.
[0039] For the acquired images, 2D face images may be readily obtained. In the case of videos, key frames may be extracted from the videos to form a series of 2D face images (104) .
[0040] After obtaining the 2D face images, face object detection and face landmark detection may be performed (106) . For instance, this process may be performed by a face detection tool, such as MediaPipe or Dlib. With these tools, not only is the user’s face detected, but also a plurality of face landmarks may be identified. Face landmarks are key points of the user’s face that identify specific areas of the face and are used to locate and identify facial structures. Each face landmark may be represented by a corresponding landmark point that describes the coordinates of the face landmark on the face.
[0041] Upon the detection of the face landmarks, a face mesh may be generated (108) . The face mesh may be generated by connecting adjacent landmark points, which serve as the vertices of multiple polygons. These polygons, when combined, form the face mesh that reflects the geometry of the face.
[0042] It should be understood that the “user” in 102 may refer to a single one user or multiple users. In a single-user scenario (e.g., images or videos of a single one user are captured in 102) , only one face of the user may be detected, and consequently, only one face mesh may be generated in 108.
[0043] On the other hand, in a multiple-user scenario (e.g., images or videos of multiple users are captured in 102) , there could be more than one user or more than one face present in one or more of the images or videos, and thus multiple faces (from multiple users) may be detected. In this case, for any detected face of a user, a corresponding face mesh may be generated based on the detected face of that user.
[0044] For the sake of clarity and simplicity, the majority of the present disclosure is described under a single-user scenario (i.e., assume that there is a single user, generate one face mesh for that user and predict an aging effect for that user) . However, it is important to emphasize that these descriptions are not limited to the single-user context. They are equally applicable to a multi-user scenario. In such a scenario, a corresponding face mesh may be generated for each user based on the detected face of that user, and an aging effect may be predicted for each user accordingly.
[0045] Next, key facial sub-regions may be selected from the face mesh (110) . Each sub-region corresponds to a facial area / zone of the user. In some examples, the sub-regions may be selected based on clinical data indicating which parts of a face are prone to aging. In other examples, the sub-regions may be selected based on the user’s interest. For instance, if the user is interested in the areas around the corners of the eyes and mouth, these areas may be used as the basis for selection of the sub-regions.
[0046] The face dynamic parameter of each sub-region may then be calculated (112) . In some embodiments, the face dynamic parameter of a sub-region may indicate variation of a geometric parameter (e.g., an area, a perimeter, or a centroid) of the sub-region. That is, the face dynamic parameter reflects the dynamic changes in the user’s face between specific expressions and no expression of the user.
[0047] By calculating the facial dynamic parameters of the sub-regions, facial zones where squeezing and / or stretching occurs between specific expressions and no expression may be identified (114) . For example, if the area of a same sub-region increases, the sub-region may be considered as undergoing stretching; conversely, if the area of a same sub-region decreases, the sub-region may be considered as undergoing squeezing.
[0048] In some embodiments, if the change (increase or decrease) of the area of a sub-region is minimal, the sub-region may be considered as not undergoing stretching or squeezing. For instance, a specific threshold may be configured. If the change of the area of a sub-region exceeds the threshold, the sub-region may be considered as undergoing stretching or squeezing, otherwise it may be considered as not undergoing stretching or squeezing. In some other examples, two separate thresholds may be configured. If the increase of the area of a sub-region exceeds a first threshold, the sub-region may be considered as undergoing stretching, and if the magnitude of the decrease of the area of a sub-region exceeds a second threshold, the sub-region may be considered as undergoing squeezing.
[0049] The aging effect of the user may then be predicted based on the face dynamic parameter (or based on the identified squeezing and stretching zones) (116) . Specifically, a higher value of the face dynamic parameter of a sub-region indicates that the sub-region has undergone a greater degree of squeezing or stretching between specific expressions and no expression, thus it may be predicted that there is a higher possibility of having an aging effect (such as wrinkle growing or sagging) in the sub-region. Likewise, a lower value of the face dynamic parameter of a sub-region indicates that the sub-region has undergone a smaller degree of squeezing or stretching between specific expressions and no expression, thus it may be predicted that there is a lower possibility of having an aging effect in the sub-region.
[0050] After the prediction is completed, the prediction result may be displayed (118) . In addition, an anti-aging product and / or treatment may be recommended to the user based on the prediction result (120) .
[0051] While FIG. 1 shows a specific architecture 100 for predicting an aging effect based on face dynamics, it should be understood that the present disclosure is not limited thereto. Those of ordinary skill in the art will recognize that architectures different than the architecture 100 could also be implemented.
[0052] FIG. 2 illustrates a flowchart of a method 200 for predicting an aging effect based on face dynamics according to various aspects of the present disclosure.
[0053] As illustrated in FIG. 2, the method 200 begins at block 202. At block 202, a set of face images of a user may be acquired, wherein the set of face images comprises a neutral image showing no expression and one or more expressive images, each expressive image showing a specific expression.
[0054] In the context of the present disclosure, a neutral image refers to an image of the user with no facial expression, and an expressive image refers to an image of the user with a facial expression such as smile, surprise, sadness, etc. Specifically, when the user′sfacial muscles are in a relaxed state (without obvious action such as eyebrow-raising, lifting or drooping of the corners of the mouth) and there are no obvious emotional features on the face (e.g., joy, anger, sadness, happiness, etc. ) , the resulting image of the user′sface can be regarded as a neutral image with no expression.
[0055] In some embodiments, acquiring the set of face images further comprises: capturing multiple images or one or more videos of the user that show a variety of expressions as well as no expression; and extracting the set of face images based on the captured images or videos.
[0056] In some examples, if the user has previously captured images or videos showing a variety of expressions as well as no expression, then the set of face images may be acquired based on the previously captured images or videos. In some other examples, if the user does not have such images or videos, the user may be instructed to capture new images or videos showing a variety of expressions as well as no expression.
[0057] The set of face images may be obtained based on the captured images or videos. For the captured video, key frames may be extracted from the video to generate 2D images. Among the captured images and those extracted from video, there may be images that show repetitive expressions and / or do not contain the user’s face. In such cases, a preprocessing may be performed to delete these images. The remaining images may constitute the set of face images.
[0058] In some embodiments, for the captured images or videos, it is necessary to verify whether these images or videos meet a certain requirement (e.g., resolution requirement, expression quantity requirement, etc. ) . If the captured images or videos do not meet the requirement, the user may be prompted to recapture images or videos until the recaptured images or videos meet the requirement.
[0059] At block 204, a face mesh may be generated based on a detected face of the user in the set of face images, wherein the face mesh comprises a set of landmark points indicating one or more key face features of the user. In some embodiments, the user’s face in the set of face images may be detected using a face detection tool, such as Dlib and MediaPipe.
[0060] In some embodiments, the key face features may comprise, but not limited to, eyes, eyebrows, mouth, and nose of the user.
[0061] In some embodiments, the face detection tool such as Dlib and MediaPipe may also be used to generate the face mesh.
[0062] Dlib and MediaPipe are both open source tools that can be used for face detection as well as face landmark detection. Dlib is a C++ open source package containing machine learning algorithms for a wide range of applications in the fields of machine learning and computer vision. MediaPipe is an integrated tool library of machine learning algorithms containing a variety of models for face detection, face landmark detection, gesture recognition. Specifically, MediaPipe may perform face (landmark) detection by applying a machine learning algorithm consisting of two real-time deep learning neural networks to firstly compute face locations and secondly predict the approximate face geometry via regression. Both Dlib and MediaPipe are well-known in the field of machine learning and will not be further elaborated here.
[0063] As an example, FIG. 4 illustrates a schematic diagram 400 showing examples of face landmarks detected utilizing Dlib and MeidaPipe. Specifically, the left side of FIG. 4 shows an example of face landmarks detected utilizing Dlib, and the right side of FIG. 4 shows an example of face landmarks detected utilizing MediaPipe. As can be seen from FIG. 4, the face landmarks detected utilizing Dlib includes 68 landmark points, whereas the face landmarks detected utilizing MeidaPipe includes 468 landmark points. Despite the difference in the quantity of landmark points, both sets of landmark points cover the key features of the face including eyes, eyebrows, nose and mouth.
[0064] While FIG. 4 shows face landmarks detected utilizing Dlib and MediaPipe, it should be understood that Dlib and MediaPipe are only exemplary tools. In practice, other tools (e.g., other open source tools or non-open source tools) may also be used by those of ordinary skill in the art to detect user’s face and face landmarks.
[0065] After detecting the face landmarks, the face mesh may be generated by connecting the adjacent landmark points to form multiple polygons. As depicted on the right side of FIG. 4, an example of such a face mesh is generated by connecting the adjacent landmark points of the 468 landmark points.
[0066] Returning to FIG. 2, after generating the face mesh, the method 200 proceeds to block 206. At block 206, one or more sub-regions may be selected from the face mesh.
[0067] In some embodiments, selecting one or more sub-regions from the face mesh may comprise: selecting a subset of landmark points from the set of landmark points; and obtaining the one or more sub-regions based on the subset of landmark points.
[0068] In some embodiments, the one or more sub-regions may be selected based on a predetermined rule. For example, facial zones that are prone to aging may be identified based on clinical data, and sub-regions may be selected based on those facial zones. In another example, sub-regions may be selected based on a user′sinterest. For instance, the user may be interested in the areas around the corners of the eyes and mouth, and sub-regions may be selected based on the facial areas of interest to the user.
[0069] In one embodiment, the landmark points that are closest to the boundary of a facial zone prone to aging or of interest to the user may be selected. The selected landmark points may then be connected, thereby obtaining a sub-region corresponding to the facial zone. For instance, each image in FIG. 4 illustrates two sub-regions of a user’s face mesh. The left image in FIG. 4 shows one sub-region at the outer corner of the left eye and another sub-region to the left of the corner of the mouth. Similarly, the right image in FIG. 4 shows one sub-region at the outer corner of the left eye and another sub-region on the left cheek.
[0070] Returning to FIG. 2, at block 208, a face dynamic parameter of each sub-region may be determined, wherein the face dynamic parameter of each sub-region indicates dynamic change of the sub-region.
[0071] In some embodiments, a face dynamic parameter of each sub-region may be determined by a variation of a geometric parameter of the sub-region. In some examples, the geometric parameter of a sub-region may comprise at least one of an area, a perimeter, or a centroid of the sub-region. To determine the variation of the geometric parameter of a sub-region, values of the geometric parameter of the sub-region from the one or more expressive images may be compared with a value of the geometric parameter of the sub-region from the neutral image.
[0072] In some embodiments, the face dynamic parameter of each sub-region may be defined by the magnitude of change of the sub-region between expressive images and neutral image. In the case where the geometric parameter is the area of a sub-region, for a given sub-region, the difference between the area of the sub-region from each expressive image and the area of the sub-region from the neutral image may be calculated. For instance, the calculation may be performed via the Dlib or MediaPipe. All the calculated differences may then be processed (e.g., averaged, weighted and averaged, etc. ) to derive a single value that represents the overall magnitude of area change, and this value may be regarded as the face dynamic parameter of the sub-region. For the case where the geometric parameter is the perimeter of a sub-region as well as the case where the geometric parameter is the centroid of a sub-region, the face dynamic parameter of the sub-region may be determined in a similar way and indicates the overall magnitude of the perimeter change of the sub-region as well as the overall magnitude of the centroid change (i.e., displacement of the centroid) of the sub-region, respectively.
[0073] In some other embodiments, the face dynamic parameter of each sub-region may be defined by the frequency of change of the sub-region. In the case where the geometric parameter is the area of a sub-region, for a given sub-region, the difference between the area of the sub-region from each expressive image and the area of the sub-region from the neutral image may be calculated. Subsequently, the proportion of the differences above a predetermined threshold to all the differences may be determined. This proportion may be regarded as the frequency of change of the area of the sub-region, i.e., the face dynamic parameter of the sub-region. The higher the proportion calculated for a sub-region is, the more frequent the sub-region undergoes relatively large changes. For the case where the geometric parameter is the perimeter of a sub-region as well as the case where the geometric parameter is the centroid of a sub-region, the face dynamic parameter of the sub-region may be determined in a similar way.
[0074] If the geometric parameter includes a single parameter (one of an area, a perimeter, or a centroid of a sub-region) , the face dynamic parameter of a sub-region may be a scalar, and a value of the scalar indicates the variation of this single parameter of the sub-region.
[0075] If the geometric parameter includes multiple parameters (two or more of an area, a perimeter, or a centroid of a sub-region) , the face dynamic parameter may be a vector containing multiple elements, where each element corresponds to one of the multiple parameters and the value of each element indicates the variation of the corresponding parameter of the sub-region.
[0076] It should be noted that the above geometric parameters and face dynamic parameters of the sub-region are only exemplary and not limiting. In specific embodiments, those of ordinary skill in the art may also employ other geometric parameters and / or facial dynamic parameters.
[0077] To better understand how a sub-region dynamically changes between specific expressions and no expression, FIG. 5 illustrates a schematic diagram 500 showing a set of face images of a user as well as dynamic changes of the sub-regions.
[0078] The upper part of FIG. 5 shows a set of face images of a user. The leftmost image in the upper part is a neutral image showing no expression of the user, and the remaining images are expressive images each showing a specific expression of the user.
[0079] The lower part of FIG. 5 shows dynamic changes of two exemplary sub-regions between expressive images and neutral image of the set of face images. One sub-region is situated at the outer corner of the left eye, and the other sub-region is located on the left cheek. From FIG. 5 it can be seen that the geometric properties (such as area, a perimeter, or a centroid) of the same sub-region may vary between expressive images and neutral image. It should be noted that, for purpose of illustration rather than limitation, the images in FIG. 4 and FIG. 5 are from JAFFE database (Michael J. Lyons, Shigeru Akemastu, Miyuki Kamachi, Jiro Gyoba. Coding Facial Expressions with Gabor Wavelets, 3rd IEEE International Conference on Automatic Face and Gesture Recognition, pp. 200-205 (1998) ) . More information about the JAFFE database is available at: http: / / www. kasrl. org / jaffe. html.
[0080] Returning to FIG. 2, at block 210, an aging effect of the user may be predicted based on the face dynamic parameters of the one or more sub-regions.
[0081] In some embodiments, the aging effect comprises, but is not limited to, at least one of wrinkle growing, sagging, losing radiance, losing smoothness, or losing plainness.
[0082] Studies have shown that, the greater the magnitude or the higher the frequency of changes in the same facial area / zone between specific expressions and no expression, the more likely the facial area / zone is to have an aging effect.
[0083] Thus, a result of the prediction may indicate a possibility of having an aging effect in each sub-region, wherein a higher value of the face dynamic parameter of a sub-region may indicate a higher possibility of having an aging effect in the sub-region, and a lower value of the face dynamic parameter of a sub-region may indicate a lower possibility of having an aging effect in the sub-region.
[0084] In practice, the value of a face dynamic parameter may be compared with a preset threshold. If the value of the parameter of a sub-region is less than the threshold, it indicates that the dynamic change of the sub-region is quite small when the facial expressions changes, and the sub-region is less likely to experience an aging effect. Whereas if the value of the parameter of a sub-region is greater than or equal to the threshold, it indicates that the dynamic change of the sub-region is quite large when the facial expressions changes, and the sub-region is more likely to experience an aging effect. In some embodiments, the threshold may be set using a variety of approaches, including but not limited to empirical knowledge, theoretical computations, or clinical data. For instance, an average value of the face dynamic parameters for individuals within the same age group as the user may be estimated based on clinical data and serve as the threshold. Furthermore, in some embodiments, as new clinical data becomes available over time, the threshold may be dynamically adjusted.
[0085] As mentioned above, the face dynamic parameter may be a scalar or a vector including multiple elements. When the face dynamic parameter of a sub-region is a scalar, a higher value of the scalar may indicate a higher possibility of having an aging effect in the sub-region, and a lower value of the scalar may indicate a lower possibility of having an aging effect in the sub-region. In the case where the face dynamic parameter of a sub-region is a vector including multiple elements, a higher value of any of the elements may indicate a higher possibility of having an aging effect in the sub-region, and lower values of all the elements may indicate a lower possibility of having an aging effect in the sub-region.
[0086] After predicting the aging effect, a result of the prediction may also be displayed (not shown in FIG. 2) .
[0087] In some embodiments, displaying a result of the prediction may comprise displaying one image of the set of face images and a result of the prediction associated with the one or more sub-regions on a display device.
[0088] For example, a neutral image (or an expressive image) of the user may be displayed on the display device (e.g., computer display, cell phone, tablet, etc. ) , and the result of the prediction associated with the one or more sub-regions may be displayed in / near the one or more sub-regions. Additionally, in some scenarios, a simulated depiction of the predicted aging effect in the sub-regions may also be provided.
[0089] In addition to displaying the result of the prediction, an anti-aging product and / or treatment may also be recommended based on the result of the prediction (not shown in FIG. 2) . For example, if the result of the prediction indicates that sub-regions around the corners of the user′seyes have a relatively high possibility of having an aging effect, an anti-aging product and / or treatment targeting areas around the corners of the eyes may be recommended to the user.
[0090] With method 200, the dynamic features of a user’s face are fully considered when predicting an aging effect of the user, thus improving the accuracy of the prediction.
[0091] In addition to predicting the aging effect based on face dynamics (e.g., the face dynamic parameters of the sub-regions) , in some embodiments, the aging effect may be predicted based on face dynamics in combination with an aging scoring model.
[0092] Accordingly, FIG. 3 illustrates a simplified flowchart of a method 300 for predicting an aging effect based on face dynamics in combination with an aging scoring model according to various aspects of the present disclosure.
[0093] As illustrated in FIG. 3, the method 300 begins at block 302. At block 302, a set of face images of a user may be acquired, wherein the set of face images comprises a neutral image showing no expression and one or more expressive images, each expressive image showing a specific expression. The operation of block 302 corresponds to the operation of block 202.
[0094] For the set of face images acquired through block 302, a face dynamic parameter of each sub-region may be determined through blocks 304-308. Blocks 304-308 of the method 300 correspond to blocks 204-208 of the method 200. Thus, blocks 304-308 will not be described in detail here.
[0095] For the set of face images, in addition to determining the face dynamic parameter of each sub-region through blocks 304-308, these images may also be input into an aging scoring model, as illustrated by block 310.
[0096] The aging scoring model may recognize / identify aging signs in the input face images and assess the face images based on the aging signs. In the present disclosure, aging signs refer to specific skin conditions associated with aging observed on face images, such as localized wrinkles, sagging of facial areas, enlarged skin pores, skin pigmentation, and the like. In the context of the present disclosure, the term “aging sign” is similar but slightly different from the term “aging effect” . Specifically, an aging sign is a very clear and visible feature on people’s face that indicates the impact of the aging process. In contrast, aging effects encompass not only the clear aging signs, but also some invisible impacts on people′sface due to aging. Therefore, the term “aging effect” has a broader meaning than “aging sign” .
[0097] The aging scoring model may be built based on expert grading. In some embodiments, the source of expert grading may be Skin Aging Atlas (Skin Aging Atlas Vol. 1 -Caucasian Type (ISBN-13: 978-2354030018) and Skin Aging Atlas Vol. 2 -Asian Type (ISBN-13: 978-2354030339) ) . How to build an aging scoring model based on expert grading is well-known in the art, and therefore, will not be described any further.
[0098] With the aging scoring model, a predicted score for each sub-region may be obtained, wherein each predicted score describes one or more aging signs on the corresponding sub-region of the user’s face (block 312) . In some embodiments, the predicted score may be a scalar value that indicates a severity of the aging signs on the corresponding sub-region of the user’s face. Specifically, a higher value of the score may indicate less severe aging signs, while a lower value of the score may indicate more severe aging signs.
[0099] After obtaining the face dynamic parameter and the predicted score, an aging effect of the user may be predicted based on the face dynamic parameter and the score (block 314) .
[0100] Specifically, for a particular sub-region, if a value of the face dynamic parameter is high (which indicates a high possibility of having an aging effect in the sub-region) and a value of the score predicted by the aging scoring model is low (which indicates severe aging signs in the sub-region) , this means that the prediction result based on face dynamics is consistent with the result of the aging scoring model. In this case, the prediction result based on face dynamics is more reliable. For example, for a particular sub-region, the higher the value of the face dynamic parameter and the lower the value of the score, the higher the possibility of having an aging effect in the sub-region may be predicted. Conversely, the lower the value of the face dynamic parameter and the higher the value of the score, the lower the possibility of having an aging effect in the sub-region may be predicted.
[0101] By combining the face dynamics and the aging scoring model to predict the aging effect, the accuracy and stability of the prediction may be improved.
[0102] Methods 200 and 300 may be utilized in a variety of scenarios. For example, an APP may be installed on a user device (e.g., cell phone) that instructs the user to take a selfie to obtain images or videos of the user′sface, which may be used to predict the user′saging effect utilizing methods 200 or 300.
[0103] For a clearer understanding, FIG. 6 illustrates a schematic diagram of an exemplary process 600 for predicting an aging effect based on face dynamics according to various aspects of the present disclosure.
[0104] For ease of explanation, it is assumed that a mobile APP for predicting an aging effect of a user based on face dynamics is installed on the user’s cell phone.
[0105] As shown in FIG. 6, the process 600 begins with the user launching the APP (602) . Then the APP may ask the user whether the user already has images or videos that show a variety of expressions as well as no expression (604) .
[0106] If the user has such images or videos (output of 604 is “Y” ) , then the user may load these images or videos to the APP (606) . If the images or videos are stored on the user’s cell phone, the user may directly load these images or videos to the APP. Alternatively, if the images or videos are stored externally to the user’s cell phone (e.g., stored on a server) , the user may download these images or videos from the server to the APP.
[0107] For the loaded images or videos, it is necessary to verify whether these images or videos meet a certain requirement (608) . For example, the loaded images or videos may be required to meet a resolution requirement. In another example, the loaded images or videos may be required to contain a sufficient number of expressions.
[0108] If the loaded images or videos do not meet the requirement (output of 608 is “N” ) , or if the user does not have images or videos that show a variety of expressions as well as no expression (output of 604 is “N” ) , then the user may be instructed to capture new images or videos (610) . In an example, the user may be instructed to take a selfie photo without any expression to obtain a neutral image, and then take selfie photos with different expressions (e.g., smile, surprise, frown, etc. ) to obtain expressive images. In another example, the user may be asked to take a selfie video and make expressions as the APP instructed to obtain a video that meets the requirement.
[0109] After capturing the new images or videos, they may be verified to see whether these new images or videos meet the requirement (612) .
[0110] If the new images or videos do not meet the requirement (output of 612 is “N” ) , then the user may be instructed to make an adjustment (614) and recapture images or videos until the recaptured images or videos meet the requirement. In some embodiments, the adjustment may include adjusting one or more camera parameters, adjusting the position and / or pose of the user, etc.
[0111] If the loaded / newly captured images or videos meet the requirement (output of 608 or 612 is “Y” ) , an aging effect of the user may be predicted based on these images or videos (616) . For example, the aging effect may be predicted utilizing the method proposed by the present disclosure (e.g., method 200, method 300) .
[0112] Next, a result of the prediction may be displayed (618) . For instance, a neutral image (or an expressive image) of the user as well as one or more sub-regions may be displayed on the APP. For each sub-region, a possibility of having an aging effect in that sub-region may be also displayed (e.g., displayed in or near that sub-region) .
[0113] Of course, the above manner of displaying the result of the prediction is only exemplary. In some embodiments, the result of the prediction may be displayed in a different manner. For example, different colors may be used to indicate different ranges of possibilities of having an aging effect. For instance, a greenish color of a sub-region may indicate a relatively low possibility of having an aging effect in the sub-region, and a reddish color of a sub-region may indicate a relatively high possibility of having an aging effect in the sub-region. In this case, a neutral image of the user may be displayed on the APP, and different sub-regions of the neutral image may be shown in corresponding colors.
[0114] Similarly, different grayscale levels may be used to represent different ranges of possibilities of having an aging effect. For example, a lighter grayscale of a sub-region may indicate a relatively lower possibility of having an aging effect in the sub-region, and a darker grayscale of a sub-region may indicate a relatively higher possibility of having an aging effect in the sub-region. In such way, the result of the prediction may be shown more intuitively to the user.
[0115] The above manners of displaying the result of the prediction may be used separately or in combination.
[0116] In some cases, if the user device is quite large in size, the display area for the prediction result is sufficient. In such cases, in addition to the above prediction result, the face dynamic parameter of each sub-region may also be displayed. For example, for a particular sub-region, the magnitude and / or frequency of change of the area of the sub-region between expressive images and neutral image may be displayed. As a result, more detailed information may be displayed to the user.
[0117] In some cases, if the user device is quite small in size, the display area for the prediction result may be limited. In such cases, the face dynamic parameter of each sub-region may not be displayed by default, to save display space. And when the user selects a particular sub-region, the face dynamic parameter of that sub-region will then be displayed.
[0118] In some embodiments, in addition to displaying the prediction result, an anti-aging product and / or treatment may be recommended to the user (620) . For instance, if the prediction result indicates that there is a high possibility that the user will have a wrinkle growing in the corners of the eyes, an anti-aging product and / or treatment for the corners of the eyes may be recommended to the user, such as an eye cream with anti-wrinkle effect. In some embodiments, detailed information for the recommended product and / or treatment may also be accompanied.
[0119] While FIG. 6 shows a specific process 600 for predicting an aging effect based on face dynamics, it should be understood that the present disclosure is not limited thereto. Those of ordinary skill in the art will recognize that processes different than process 600 could also be used to predict an aging effect. For example, the operation of recommending an anti-aging product and / or treatment (i.e., 620) may be omitted. In another example, the user may use a webpage service (rather than the mobile APP) to conduct the aging effect prediction. In such cases, the user may upload images or videos to the webpage, which would then perform the prediction based on the uploaded images or videos.
[0120] FIG. 7 illustrates a block diagram of a system 700 for predicting an aging effect based on face dynamics according to various aspects of the present disclosure.
[0121] As illustrated in FIG. 7, the system 700 includes an image acquisition unit 702, a face mesh generation unit 704, a parameter determination unit 706, an aging effect prediction unit 708, and an optional post-processing unit 710. One or more of these units may be directly or indirectly connected or in communication with each other.
[0122] In some embodiments, the image acquisition unit 702 may be configured to acquire a set of face images of a user, wherein the set of face images comprises a neutral image showing no expression and one or more expressive images, each expressive image showing a specific expression.
[0123] In some embodiments, the face mesh generation unit 704 may be configured to generate a face mesh based on a detected face of the user in the set of face images, wherein the face mesh comprises a set of landmark points indicating one or more key face features of the user.
[0124] In some embodiments, the parameter determination unit 706 may be configured to: select one or more sub-regions from the face mesh; determine a face dynamic parameter of each sub-region, wherein the face dynamic parameter of each sub-region indicates dynamic change of the sub-region.
[0125] In some embodiments, the parameter determination unit 706 may be further configured to determine a face dynamic parameter of each sub-region by: determining a geometric parameter of the sub-region, wherein the geometric parameter comprises at least one of an area, a perimeter, or a centroid of the sub-region; and determining the face dynamic parameter based on a variation of the geometric parameter of the sub-region.
[0126] In some embodiments, the aging effect prediction unit 708 may be configured to predict an aging effect of the user based on the face dynamic parameters of the one or more sub-regions.
[0127] In some embodiments, the post-processing unit 710 may be configured to: display one image of the set of face images and a result of the prediction associated with the one or more sub-regions on a display device; and / or recommend an anti-aging product and / or treatment based on the result of the prediction.
[0128] Although FIG. 7 illustrates particular units of the system 700, it should be understood that, these units are only exemplary and not limiting. In various embodiments, one or more of these units may be combined, split, removed, or additional units may be added. For example, in some embodiments, the system 700 may include a control unit for controlling the operations of other units within the system.
[0129] While the embodiments have been illustrated and described herein, it will be understood by those skilled in the art that various changes and modifications may be made, and equivalents may be substituted for elements thereof without departing from the true scope of the present technology. In addition, many modifications may be made to adapt to a particular situation and the teaching herein without departing from its central scope. Therefore, it is intended that the present embodiments are not limited to the particular embodiment disclosed as the best mode contemplated for carrying out the present technology, but that the present embodiments include all embodiments falling within the scope of the appended claims.
Claims
1.A methsd for predicting an aging effect based on face dynamics, comprising:- acquiring a set of face images of a user, wherein the set of face images comprises a neutral image showing no expression and one or more expressive images, each expressive image showing a specific expression;- generating a face mesh based on a detected face of the user in the set of face images, wherein the face mesh comprises a set of landmark points indicating one or more key face features of the user;- selecting one or more sub-regions from the face mesh;- determining a face dynamic parameter of each sub-region, wherein the face dynamic parameter of each sub-region indicates dynamic change of the sub-region; and- predicting an aging effect of the user based on the face dynamic parameters of the one or more sub-regions.2.The method according to claim 1, wherein the aging effect comprises at least one of wrinkle growing, sagging, losing radiance, losing smoothness, or losing plainness.3.The method according to claim 1, wherein acquiring the set of face images further comprises:capturing multiple images or one or more videos of the user that show a variety of expressions as well as no expression; andextracting the set of face images based on the captured images or videos.4.The method according to claim 1, wherein the key face features comprise eyes, eyebrows, mouth, and nose of the user.5.The method according to claim 1, wherein selecting one or more sub-regions from the face mesh further comprises:selecting a subset of landmark points from the set of landmark points; andobtaining the one or more sub-regions based on the subset of landmark points.6.The method according to claim 1, wherein determining a face dynamic parameter of each sub-region further comprises:determining a geometric parameter of the sub-region, wherein the geometric parameter comprises at least one of an area, a perimeter, or a centroid of the sub-region; anddetermining the face dynamic parameter based on a variation of the geometric parameter of the sub-region.7.The method according to claim 6, wherein the variation of the geometric parameter of the sub-region is determined by comparing values of the geometric parameter of the sub-region from the one or more expressive images with a value of the geometric parameter of the sub-region from the neutral image.8.The method according to claim 7, wherein a result of the prediction indicates a possibility of having an aging effect in each sub-region, wherein a higher value of the face dynamic parameter of a sub-region indicates a higher possibility of having an aging effect in the sub-region, and a lower value of the face dynamic parameter of a sub-region indicates a lower possibility of having an aging effect in the sub-region.9.The method according to claim 1, further comprising:displaying one image of the set of face images and a result of the prediction associated with the one or more sub-regions on a display device.10.The method according to claim 1, further comprising:recommending an anti-aging product and / or treatment based on a result of the prediction.11.The method according to claim 1, further comprising:inputting the set of face images to an aging scoring model to obtain a predicted score, wherein the predicted score describes one or more aging signs on the user’s face, and wherein the aging effect of the user is further predicted based on the predicted score.12.A system for predicting an aging effect based on face dynamics, comprising:- an image acquisition unit configured to acquire a set of face images of a user, wherein the set of face images comprises a neutral image showing no expression and one or more expressive images, each expressive image showing a specific expression;- a face mesh generation unit configured to generate a face mesh based on a detected face of the user in the set of face images, wherein the face mesh comprises a set of landmark points indicating one or more key face features of the user;- a parameter determination unit configured to:- select one or more sub-regions from the face mesh;- determine a face dynamic parameter of each sub-region, wherein the face dynamic parameter of each sub-region indicates dynamic change of the sub-region; and- an aging effect prediction unit configured to predict an aging effect of the user based on the face dynamic parameters of the one or more sub-regions.13.The system according to claim 12, wherein the parameter determination unit is further configured to determine a face dynamic parameter of each sub-region by:determining a geometric parameter of the sub-region, wherein the geometric parameter comprises at least one of an area, a perimeter, or a centroid of the sub-region; anddetermining the face dynamic parameter based on a variation of the geometric parameter of the sub-region.14.The system according to claim 12, further comprising a post-processing unit configured to:display one image of the set of face images and a result of the prediction associated with the one or more sub-regions on a display device; and / orrecommend an anti-aging product and / or treatment based on a result of the prediction.15.A non-transitory computer-readable medium having computer-executable instructions stored thereon that, in response to execution by one or more processors of a computer system, cause the computer system to perform operations comprising:- acquiring a set of face images of a user, wherein the set of face images comprises a neutral image showing no expression and one or more expressive images, each expressive image showing a specific expression;- generating a face mesh based on a detected face of the user in the set of face images, wherein the face mesh comprises a set of landmark points indicating one or more key face features of the user;- selecting one or more sub-regions from the face mesh;- determining a face dynamic parameter of each sub-region, wherein the face dynamic parameter of each sub-region indicates dynamic change of the sub-region; and- predicting an aging effect of the user based on the face dynamic parameters of the one or more sub-regions.