METHODS, APPARATUS FOR OBJECT DETECTION AND STABILIZED RENDERING

The system improves object localization and stabilized rendering in augmented reality applications by using a face tracking engine and training image generator to enhance occluded face detection, addressing the challenge of motion-induced inaccuracies in existing technologies.

FR3147419B1Active Publication Date: 2025-06-20LOREAL SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
FR2023003057
Authority / Receiving Office
FR · FR
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-03-30
Publication Date
2025-06-20
Estimated Expiration
2043-03-30

AI Technical Summary

Technical Problem

Existing image processing techniques using deep neural networks struggle with accurate object localization across video frames due to motion, leading to undesirable effects when augmented reality applications are rendered.

Method used

A system comprising a face tracking engine with a deep neural network for locating faces and a training image generator producing occluded and unoccluded training images to enhance occluded face detection DNN training, thereby improving object localization and stabilized rendering.

Benefits of technology

The proposed solution effectively stabilizes object locations between video frames, enhancing the accuracy of augmented reality applications such as virtual try-on experiences by minimizing jitter and ensuring consistent rendering of effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000029_0000
    Figure 00000029_0000
  • Figure 00000030_0000
    Figure 00000030_0000
  • Figure 00000030_0001
    Figure 00000030_0001
Patent Text Reader

Abstract

METHODS, APPARATUS FOR OBJECT DETECTION AND STABILIZED RENDERING Systems, methods and devices for object detection and systems, methods and devices for stabilizing the rendering of effects such as a makeup effect applied to a facial image are provided. Figure for abstract: none
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: METHODS, APPARATUS FOR OBJECT DETECTION AND STABILIZED RENDERING FIELD OF THE INVENTION

[0001] The present disclosure relates to image processing, for example using deep neural networks, and more particularly to methods and apparatus for object detection and stabilized rendering. CONTEXT

[0002] Deep learning techniques are useful for processing images, including a series of video frames, to locate one or more objects within the images. In one example, the objects are facial features including portions of a user's face. Image processing techniques are also useful for rendering effects associated with objects such as augmenting reality for the user. One example of such augmented reality is providing a virtual try-on (VTO) that simulates the application of a product to an object. Product simulation in the beauty industry includes simulating makeup, hair, and nail effects. Other examples may include locating the iris and simulating a change in color thereof, such as through a colored contact lens. These and other objects and simulations will be apparent.

[0003] Within frames of a video, the location of an object in a current frame may be different from its location in a previous frame due to motion, for example. Localizing a particular object in two or more frames using a deep neural network may lead to undesirable results when effects are applied due to inaccurate localization between frames.

[0004] Improved techniques are desired for determining object information from images to facilitate the provision of augmented realities, including VTO experiences. SUMMARY

[0005] The present invention relates to a system comprising:

[0006] a face tracking engine comprising a deep neural network (DNN) for locating a face in a face input image; and

[0007] a training image generator for producing a training image comprising the face as located, the training image comprising either an occluded training image where an occlusive object is rendered on the face, or a unoccluded training image showing the face without the occluding object, the training image being produced for DNN training of the occluded face.

[0008] According to embodiments of the invention, the system comprises one or more of the following characteristics, taken in isolation or in all technically possible combinations: - the system further comprises a training component for training an occluded face detection DNN with the training image; - the system further comprises a component for configuring a face tracking engine with the trained occluded face detection DNN such that the face tracking engine classifies and localizes facial features and classifies and / or localizes face occluding objects; - the system further comprises a data store storing a plurality of face input images for use with the face tracking engine and the training image generator to produce a plurality of training images for training the occluded face DNN; - the training image generator is configured to randomly produce the occluded training image, instead of the non-occluded training image, according to a probability chosen to maximize training of the face mask DNN; - the probability of producing the occluded training image, instead of the unoccluded training image, is a 55% chance; - the system further comprises a face cropping component configured to crop the face from the face input image in response to the localization, and wherein the training image generator generates the training image using the face as cropped; - the occluding object for rendering includes an isolated occluding object image with a transparent background; - before rendering the occlusive object on the face as located, the training image generator is configured to perform one or more of:

[0009] a. resizing the occlusive object on the face as located; and

[0010] b. increasing the occlusive object to maximize face mask DNN training; and - the occlusive object covers at least a portion of the face and wherein the occlusive object comprises any of a face mask, occlusive glasses, a hat, a scarf, a hand or one or more fingers, a smartphone, hair or a portion of any of these.

[0011] Systems, methods and devices are provided for object detection and systems, methods and devices for stabilizing the rendering of effects such than a makeup effect applied to a facial image.

[0012] In one embodiment, there is provided a computer-implemented method comprising performing by one or more processors the steps of: localizing a face in a face input image using a face tracker comprising one or more deep neural networks (DNNs) trained to localize facial features; and producing a training image comprising the face as localized, the training image comprising either an occluded training image where an occlusive object is rendered on the face or an unoccluded training image showing the face without the face mask, the training image being produced for DNN training of the occluded face.

[0013] In one embodiment, there is provided a system comprising: a face tracking engine comprising a deep neural network (DNN) for locating a face in a face input image; and a training image generator for producing a training image comprising the face as located, the training image comprising either an occluded training image where an occlusive object is rendered on the face or an unoccluded training image showing the face without the occlusive object, the training image being produced for occluded face DNN training.

[0014] In one embodiment, there is provided a computer-implemented method comprising performing by one or more processors the steps of: locating a facial feature in a current frame of a set of frames of a video stream using a face tracking engine having one or more deep neural networks (DNNs) configured to process the current frame to predict a tracking organ location of the facial feature; generating a current stabilized location for the facial feature in the current frame, the generation being responsive to the tracking organ location and prior stabilized locations of the facial feature in prior frames of the video stream; and rendering an effect on the current frame associated with the facial feature in response to the current stabilized location, the effect simulating a try-on product as a component of a virtual try-on experience.

[0015] In one embodiment, there is provided a system comprising: a face tracking engine having computational circuitry configured to locate a facial feature in a current frame of a set of frames of a video stream using one or more deep neural networks (DNNs) configured to process the current frame to predict the location by the tracker of a facial feature; a stabilizer component having computational circuitry configured to generate a current stabilized location for the facial feature in the current frame, the generation being responsive to the tracking organ location and previous stabilized locations of the facial feature in prior frames of the video stream; and a rendering component having computational circuitry configured to render an effect to the current frame associated with the facial feature in response to the current stabilized location, the effect simulating a product for try-on as a component of a virtual try-on experience.

[0016] These and other embodiments will be apparent to a person of ordinary skill in the art. Brief description of the drawings

[0017] [Fig-1] [Fig. 1] is a block diagram showing a training pipeline including components for generating synthetic training images, in accordance with one embodiment.

[0018] [Fig.2A][Fig.2B][Fig.2C][Fig.2D][Fig.2E] Figures 2A, 2B, 2C, 2D and 2E are illustrations of representative facial images including, respectively, a facial image, a cropped facial image, two synthetic training images and a facial image with facial points in accordance with embodiments.

[0019] [Fig.3] [Fig.3] is a flowchart of operations such as for a method im computer-implemented in accordance with one embodiment.

[0020] [Fig.4] [Fig.4] is an illustration of a computer environment, in accordance with one embodiment, such as for performing a virtual fitting.

[0021] [Fig.5] [Fig.5] is a flowchart of operations such as for a method im computer-implemented in accordance with one embodiment.

[0022] [Fig.6] [Fig.6] is an illustration of a computer environment, in accordance with one embodiment, such as for performing a virtual fitting.

[0023] [Fig.7] [Fig.7] is a flowchart of operations such as for a method im computer-implemented in accordance with one embodiment. DETAILED DESCRIPTION

[0024] In accordance with the present embodiments, one or more methods, apparatuses, and techniques are provided for object localization and object stabilization. They may be used independently or together. With respect to object localization, methods, apparatuses, and techniques are described for generating synthetic data and training a network model using this synthetic data.

[0025] In embodiments described herein, the object classes, such as for classification or classification and localization, are various facial features. In one embodiment, such features include a facial contour, a nose, an inner mouth, an outer mouth, a left eye, a right eye, a left eyebrow and a right eyebrow. In one embodiment, an additional class of objects includes a face mask such as a mask worn to reduce the transmission of airborne particles such as aerosols. A face mask may occlude, in whole or in part, one or more facial features from a facial image. Examples of occluded facial features, or portions thereof, include a nose, an inner mouth, an outer mouth, and a facial contour (e.g., portions of a jaw, a chin, or both).

[0026] In embodiments herein, one or more deep neural networks, e.g., as a component of a face tracking engine, classify and localize an input image to detect objects therein and determine the respective locations of at least some of the detected objects. In one embodiment, a first deep neural network determines whether a face is present and provides a bounding box, e.g., with which to crop the input image to localize the face therein. In one embodiment, a second deep neural network classifies and localizes a plurality of facial features (e.g., each of the detected objects) in a facial image comprising the cropped input image. Localizing, in one embodiment of such a network, comprises identifying facial points defining general contours for at least some of the detected facial features.In one embodiment, for example, working in parallel and processing the same cropped facial image, an additional deep neural network classifies whether or not a face mask is present on the cropped face. In one embodiment, localization of the face mask itself is not performed. In one embodiment, the face mask is localized.

[0027] In one embodiment, for example, working in parallel and processing the same cropped facial image, but instead of classifying, an additional deep neural network performs segmentation and provides a segmentation mask showing where a facial mask is present or not on the cropped face.

[0028] Classifying, localizing, and / or segmenting facial masks is useful, for example, in a virtual makeup try-on application. An effect to be tried may be rendered relative to the classification and / or localization, e.g., a non-rendering of part or all of the face to be rendered is occluded. The results of the facial mask classification may be used to non-render lipstick. The rendering may be relative to a segmentation such as a segmentation mask for an occlusive object. Thus, in one embodiment, a face tracking engine may produce one or more of classification results, localization results, and segmentation results for at least some of the objects. The localization results may include facial points providing contours, relative to a processed image. [Fig. 2E] described below illustrates facial points.

[0029] In one embodiment, as noted, rendering operations or a rendering component, as applicable to the methods and apparatuses, render an output image with an effect applied to the face in response to at least some of the detected objects in the input image. In embodiments, the applied effect is a makeup effect, such as, but not limited to, an effect of an eye makeup product, eyebrows, or lips. In embodiments, the rendering is a component of a VTO experience for a user. In some examples, an effect is rendered on one or more detected objects. One example is rendering an eyebrow effect on each eyebrow and another is rendering a lip effect on each lip. Typically, the makeup appearances are rendered symmetrically, without needing to.In some examples, the effect is rendered on a region adjacent to or otherwise located relative to one or more detected objects. One example is rendering an eye makeup effect on each eyelid where each region of the eye is generally located at least partially between a detected eye and a detected pair of eyebrows. Another example is rendering blush or other cheek makeup products on a cheek located, for example, relative to the contour of the eyes and face.

[0030] In one embodiment, such as for providing a VTO experience from a "live" stream of video (e.g., a selfie video), each frame (e.g., each image) of the video is processed to detect and locate objects, and to render at least one effect in response to a location of the detected objects based on a product or service to be virtually tried on. Thus, the input may include a selfie image, which may include a selfie video frame.

[0031] In embodiments, the respective locations for at least some detected objects are derived from contours (e.g., facial points) generated by the face tracking engine (e.g., one or more deep neural networks thereof). In some embodiments, such as when the input image is a current frame of a series of frames of a video, a stabilization operation or component, as applicable to the methods and apparatus, stabilizes the respective locations of at least some of the detected objects before rendering the effect. Object stabilization is described in more detail below. Object detection

[0032] An example of a deep neural network for object detection is described in Sandler et al., “MobileNetV2: Inverted residuals and Linear Bottlenecks” (2018), published January 13, 2018, Computer Science, 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition. A copy of this publication is available at the time of filing at arxiv.org / abs / 1801.04381. A deep neural network configured in accordance with this publication is referred to herein as a network MobileNetV2 deep neural network. MobileNetV2 deep neural networks are themselves derived from a prior single-shot detector, an example of a deep neural network for object detection in images as described in Liu, Wei, et al. "SSD: Single shot multibox detector." In European Conference on Computer Vision, pp. 21-37. Springer, Cham, 2016. An earlier version was published on December 8, 2015, on arxiv.org, available at the time of deposit at arxiv.org / abs / 1512.02325.

[0033] In accordance with a present embodiment, one or more deep neural networks are configured, for example, through training, to detect (in an image) facial features, including a face wearing a face mask or a face otherwise occluded by an occluding object, where the deep neural network is adapted from a MobilenetV2 deep neural network. In one embodiment, the one or more deep neural networks are configured to classify, localize, or segment at least one occluding object (e.g., a face mask or other objects as described herein).

[0034] [Fig.l] is a training pipeline 100 according to one embodiment for defining a new deep neural network, where the pipeline 100 includes components for generating training images for training the new deep neural network. In one embodiment, an image comprising a face (e.g., a facial image) 102 is provided, such as from a data store 103, to a face tracker 104A including one or more deep neural networks 106A (each may have a MobileNetV2 deep neural network backbone). The data store 103 may include a database or other storage configuration that stores a plurality of different facial images to be used for training and / or testing deep neural networks, for example.The data store 103 may include a publicly available open data store, a closed proprietary data store, and / or both open and closed data stores.

[0035] The face tracker 104A is adapted to track (i.e., locate) classes of objects relating to a face, including a face object itself. The output of such a face tracker 104 includes a bounding box, mask, or other structure to obtain a cropped facial image 108. In the cropped facial image 108, for example, any background in the facial image 102 is minimized. [Fig. 2A] shows a representative facial image 102, including a face 202 and a background 204. The background may include other portions of the subject as well as non-subject portions. A bounding box 206 displayed as a dotted box shows the coordinates for defining a cropped image (e.g., 108), including the face 202 and reduced background content (part of the background 204). [Fig.2B] shows a cropped image re- presentation 108.

[0036] A training image generator 110 receives cropped facial images, such as cropped image 108, and, in accordance with its configuration, generates training images, such as image 114. The training image generator 110, in one embodiment, generates a plurality of training images for each cropped facial image it receives. Although not shown, the training images, in one embodiment, are stored in the data store 103 as for later use. In one embodiment, training images are provided to train a deep neural network as described in more detail. Although described as training images, they may include test images in one embodiment.

[0037] In one embodiment, the training image generator 110 generates training images such as by rendering an effect on at least some of the cropped facial images it receives. In one embodiment, the effect is the application of a face mask such as by applying an isolated face mask image 112 (e.g., from the data store 103) via a rendering operation. The data store 103 may store a plurality of isolated face mask images. The isolated face mask images, such as the image 112, may include a portable network graphics image (e.g., ".png") of a face mask or other usable image format, when the background is transparent.That is, when the face mask is applied to a cropped facial image, only the face mask occludes the portion of the cropped image on which it is rendered, and the transparent background of the facial image allows the other portion of the cropped facial image to remain visible.

[0038] In one embodiment, the training image generator 110 operates to: use the predicted / labeled point coordinates to add the face mask image onto the cropped face images to generate synthetic data; resize the face mask image before application to the cropped face; apply augmentations such as rotation and translation to the mask in at least some of the generated training images to ensure diversity of the training data; and generate some training images without masks, e.g., such that a target percentage of images include masks, e.g., 55% of those training images include masks.

[0039] In one embodiment, the face tracker 104A produces facial feature location data, e.g., facial points for a face contour. An example is illustrated in [Fig. 2E] described further herein. The facial points and / or the facial bounding box may be useful for resizing the face mask image for the detected face.

[0040] In one embodiment, the training image generator defines a plurality of training images from a single cropped facial image, where the plurality of training images includes one or more of a cropped facial image without a face mask added, a plurality of cropped facial images each with a different face mask added, and / or each with a face mask added in a different manner (e.g., after rotation or translation of the face mask).

[0041] [Fig.2C] shows a training image 114A including a mask image 112A applied to the cropped image of [Fig.2B]. [Fig.2D] shows a training image 114B including a mask image 112A after rotation (e.g., horizontally flipped) and translation. [Fig.2B] may also define a training image where no face mask is included in the image.

[0042] [Fig.2E] shows an exemplary output 200 of the face tracker 104A including a cropped facial image 202 and groups 204 of facial points such as face contour facial points 204A, eyebrow facial points 204B, and nose facial points 204C, etc. The schematic is of an annotated cropped facial image 202 with the groups of facial points for illustration purposes. The output of the face tracker need not necessarily include an annotated image and the output may be separate data. The facial points in a group of individuals are numbered (e.g. 0, 1, 2...) and help define the contour of the detected object. The face tracker assigns each point so that it is placed in consistent locations relative to the outline of the object it represents. For example, a particular point might always be at the right corner of the mouth.In one example, the facial points are X,Y pixel coordinates relative to the cropped facial image 202 and are associated with respective detected objects from (e.g., one of) the arrays 106B of the tracker 104A.

[0043] The components above the dotted line 116 such as the face tracker 104A and the training image generator 110 are useful for defining training images, including face mask training images and non-face mask training images from the faces as located.

[0044] The training images, such as image 114, are applied to train a deep neural network such as a face tracker component 104B with one or more deep neural networks 106B (each may have a Mo-bileNetV2 backbone as in 106A). The components 104B and 106B are similar to the components 104A and 106A, but include a configuration applicable to classify a face mask object in the facial images. The components 104B and 106B are also configured in an applicable training configuration, for example, as described in the aforementioned publication “MobileNetV2: Inverted Residuals and Linear Bottlenecks”. The deep neural network to be trained for classification face masks may be structured in the same manner as a deep neural network component of the face tracker 104A that classifies (detects) other objects. In one embodiment, the deep neural network for face mask detection does not need to localize the face mask. In one embodiment, the deep neural network for face mask detection localizes and / or segments the face mask and is therefore configured with appropriate structures for training and producing results as appropriate. In other words, the deep neural network may be configured for training to segment and produce a mask.

[0045] After training, the resulting face tracker 104B with its one or more deep neural networks 106B, as trained, is tested, for example using real images of faces with face masks. After the testing and training cycles, as needed, the resulting face tracker 104B with its one or more deep neural networks 106B as may be configured from real-time or live application use (not a training configuration) is useful for classifying facial images to determine whether a mask is present or not and / or for locating / segmenting the face mask. This may also include segmenting. Preferably, the resulting face tracker 104B with its one or more deep neural networks 106B also classifies and localizes for other facial features.In one embodiment, the resulting face tracker 104B with its one or more deep neural networks 106B provides an engine for locating facial features such as for use in an application providing a VTO experience, described later herein.

[0046] Although the embodiments of [Fig. 1] are described with reference to one or more networks 106A each having a Mo-bileNetV2 deep neural network backbone, other neural network backbones defined for image localization tasks may form the backbone of the face tracker and be similarly adapted, for example, through training with synthetic images to detect the presence of facial masks (e.g., classify, localize, and / or segment for an occluding object).

[0047] Although described with reference to detecting a face mask that occludes at least a portion of a face, the methods, apparatus, and techniques presented herein may be adapted, for example, to define synthetic data and to train an occluded face detection network to detect other types of facial occlusion where at least a portion of the face is occluded by another object. For example, generating and training synthetic data may be performed for occlusion by face masks, occlusion by sunglasses / night glasses, or other eyewear. occlusions that occlude a portion of the face, occlusion by hair, occlusion by a scarf, occlusion by a hat, occlusion by a hand / one or more fingers, occlusion by a smartphone (e.g., when a selfie taken in a mirror has a portion of the smartphone covering the face (reflected) in the image), etc.

[0048] For greater diversity of training data, it may be preferable to include a diversity of examples of objects that are applied to faces - for example: different examples of types of occluding glasses, hats, hands, (number of) fingers, smartphones, etc. In one embodiment, the occluded face detection network is configured to detect more than one class of occluding object and is trained with training images including occluding objects for each class. Such occluded face detection may also be configured and trained to localize these objects, including segmentation.

[0049] [Fig. 3] is a flowchart of operations 300 as for a computer-implemented method. The method may comprise one or more processors performing the steps shown in [Fig. 3], for example. In an embodiment 1: the method comprises step 302 which shows localizing a face in a face input image using a face tracker comprising at least one deep neural network (DNN) trained to localize facial features; and step 304 which shows producing a training image comprising the face as localized, the training image comprising either an occluded training image where an occlusive object is rendered on the face or an unoccluded training image showing a face without the occlusive object, the training image being produced for DNN training of the occluded face.In an embodiment 2: Embodiment 1 may include operations such as in step 306 that demonstrate training an occluded face detection DNN with the training image. The training may include training in classification, localization, and / or segmentation of the occlusive object.

[0050] In an embodiment 3: The occlusive object in embodiment 1 or embodiment 2 covers at least a portion of the face and includes any of a face mask, occlusive glasses, a hat, a scarf, a hand or fingers, hair, a smartphone or a portion thereof.

[0051] In an embodiment 4: for any one of embodiments 1 to 3, the occluded face detection DNN comprises a pre-trained DNN for classifying and localizing facial features such that, when trained, the occluded face detection DNN detects the presence of at least one face-occluding object and classifies and localizes facial features.

[0052] In an embodiment 5: for any one of embodiments 1 to 4, steps 302 and 304 may be repeated (not shown) with a plurality of face input images of different faces to produce a plurality of training images for training occluded face DNNs.

[0053] In an embodiment 6: for any one of embodiments 1 to 5, producing the training image randomly produces the occluded training image, instead of the non-occluded training image, according to a probability chosen to maximize the occluded face DNN training. In an embodiment 7: for embodiment 6, the probability of producing the occluded training image, instead of the non-occluded training image, is a 55% chance.

[0054] In an embodiment 8: for any of embodiments 1 to 7, the method comprises (e.g. between steps 203 and 304 but not shown) cropping the face from the input face image in response to locating and producing the training image using the face as cropped.

[0055] In an embodiment 9: for any of embodiments 1 to 8, the occluding object for rendering comprises an isolated occluding object image with a transparent background. In an embodiment 10: for embodiment 9, the method comprises, before rendering the occluding object on the face as located, performing one or more of: resizing the occluding object on the face as located; and augmenting the occluding object to maximize DNN training of the face mask.

[0056] It will be understood that corresponding system embodiments are disclosed for each of embodiments 1-10, for example where the system includes respective components having computational circuitry configured to perform the operations of the computer-implemented method embodiments.

[0057] [Fig.l], by way of example, illustrates a system such as one or more processors and / or computational circuitry providing components capable of performing any of the method embodiments 1 to 10. For example, there is provided a system comprising: a face tracking engine comprising a deep neural network (DNN) for locating a face in a face input image; and a training image generator for producing a training image comprising the face as located, the training image comprising either an occluded face training image where an occluded object is rendered on the face or an unoccluded training image showing the face without the occluded object, the training image produced for occluded face DNN training. VTO ​​Application

[0058] [Fig.4] is an illustration of a computing environment 400, in accordance with an embodiment, such as for practicing one or more method aspects. The computing environment 400 has a user computing device 402, such as a smartphone, a communications network 404, a server 406, and a server 408. The communications network 404 includes wired and / or wireless networks, which may be public or private and may include, for example, the Internet. The server 406 includes a server computing device such as for providing a website. The server 408 includes a server computing device such as for providing e-commerce transaction services. Although shown separately, the servers 406 and 408 may comprise a single server device. The computing environment is simplified. For example, payment transaction gateways and other components such as for completing an e-commerce transaction are not shown.

[0059] The computing device 402 includes a storage device 410 (e.g., a non-transitory device such as memory and / or a solid-state drive (SSD), etc.) for storing instructions that, when executed by a processor (not shown), cause the computing device 402 to perform operations such as a computer-implemented method. The storage device 410 stores a virtual try-on application 412 including components such as software modules providing a user interface 414, a face tracker 104B with one or more deep neural networks 106B as trained in accordance with [Fig.l], a VTO rendering pipeline component 416, a product recommendation component 418 with product data 420, and a shopping component 422 with a shopping cart 424 (e.g., shopping data).

[0060] In one embodiment, the VTO application is a web application as obtained from the server 406. Although not shown, the user device 402 may store a web browser for running the VTO web application 412. In one embodiment, the VTO application is a native application in accordance with an operating system (also not shown) and software development requirements that may be imposed by a hardware manufacturer, for example, of the user device 402. The native application may be configured for web-based or similar communications to the servers 406 and 408, as is known.

[0061] [Fig. 4] shows various input and output data or information associated with the use of the VTO application 412, for example. This includes an input image 426 of the user to be processed for a VTO experience, an output image 428 to which product effects are simulated providing a VTO experience, a VTO product selection 430 comprising a user input selecting one or multiple product effects to be simulated, VTO product options 432 including options for products to be virtually tried on, e.g., for selection by a user of the device 402, and purchase transaction information 434 including purchase information provided to and / or received from a user to purchase a product.

[0062] In one embodiment, via one or more user interfaces 414, VTO product options 432 are presented for selection for virtual try-on by simulating effects on an input image 426. In one embodiment, the VTO product options 432 are derived from or associated with the product data 420. In one embodiment, the product data may be obtained from the server 406 and provided by the product recommendation component 418. Although not shown, user or other input may be received for use in determining product recommendations. The user may be prompted, for example via one of the interfaces 414, to provide input to determine product recommendations. In one embodiment, the product recommendation component 418 communicates with the server 406.The server 406, in one embodiment, determines the recommendation based on the input received via the component 418 and provides product data accordingly. The user interface 414 may present the VTO product choices, for example, by updating the display thereof in response to the received data as the user navigates or otherwise interacts with the user interface.

[0063] In one embodiment, the one or more user interfaces provide instructions and commands for obtaining the input image 426, and a VTO product selection input 430 such as an identification of one or more recommended VTO products to try. In one embodiment, the input image 426 is an image of a user's face, which may be a still image or a frame of a video. In one embodiment, the input image 426 may be received from a camera (not shown) of the device 402 or from a stored image (not shown). The input image 426 is provided to the face tracker 104B, for example, for processing to detect objects in the facial image using one or more deep neural networks 106B as trained. In one example, the network classifies, localizes, or segments a face mask (or other occluding object) in the image.For example, classifying the presence of a face mask is useful for producing a request (e.g., an instruction to a user, e.g., via 414 user interfaces), to lower or remove a face mask. This can apply to any occluding object for which the face tracking engine is trained.

[0064] In one embodiment, the output (not shown) of the face tracker 104B, such as classification results, localization results, or segmentation results for one or more detected objects, is provided to the VTO rendering pipeline component 416. In one example, the output may include a bounding box and, as shown in [Fig. 2E], facial points for detected objects. The input image 426 is also provided (e.g., made available) to the component 416. The VTO product selection 430 is also provided to the component 416 to determine which effects are to be rendered. In an embodiment relating to makeup simulation, one or more effects may be indicated such as for one or more of the product categories including: lips, eyeshadow, eyeliner, blush, etc.

[0065] The VTO rendering pipeline component 416, in one embodiment, determines whether to render one or more product effects on the input image 426 to simulate a try-on. For example, in response to the face mask classification output, the VTO rendering pipeline component 416 may determine not to render a product effect, for example, because a mask is detected. When a face mask is detected, for example, the VTO rendering pipeline component 416 may trigger the user interface 414 to request the user to remove the face mask. A new image may be received and processed by the face tracker 104B. In one embodiment, images are continuously received as a component of a live stream (e.g., a selfie video).

[0066] If the VTO rendering pipeline component 416 determines to render the one or more product effects, in one embodiment, the VTO rendering pipeline component 416 renders effects onto the input image 426, e.g., by drawing (rendering) effects in layers, one layer for each product effect, to produce an output image 428. Portions of the operations of the VTO rendering pipeline component 416 (e.g., such as drawing the layers) may be performed by a graphics processing unit, in one embodiment. The rendering is in accordance with the product data 420 as selected by the VTO product selection 430 and is responsive to the location of detected objects. For example, a VTO product selection of a lipstick, lip gloss, or other lip-related product invokes the application of an effect to one or more detected mouth- or lip-related objects at respective locations.Similarly, a selection of eyebrow-related products invokes the application of a selected product effect to the detected eyebrow objects. Typically, for symmetrical appearances, the same eyebrow effects are applied to each eyebrow, the same lip effect, or the same eye effect to each eye region, but this is not required. In one example, rendering is applied to a region that is relative to the detected objects, such as one or more adjacent detected objects. Some VTO product selections include a selection of more than one product, such as coordinated products for the . eyebrows and eyes or other combinations of detected objects. The VTO 416 rendering pipeline component can render each effect, for example, one at a time until all effects are applied. The order of application can be defined by rules or in the product selection, e.g., a lipstick before a lip gloss.

[0067] In an embodiment where an occluding object is detected and the location is determined, for example, as represented in a segmentation mask, rendering may be responsive to such a segmentation mask. Effect rendering may be applied to portions of the face that are not occluded. A segmentation mask may indicate which pixels of the face are available to (e.g., can) receive an effect such as a makeup effect and which pixels are not available to receive an effect.

[0068] The user interfaces 414 provide the output image 428. The output image 428, in one embodiment, is presented as a portion of a live stream of successive output images (each such as example 428) such as when a selfie video is augmented to present an augmented reality experience. In one embodiment, the output image 428 is presented in conjunction with the input image 426, such as in a side-by-side display for comparison purposes. In one embodiment, the output image 428 may be saved (not shown) such as on the storage device 410 and / or shared (not shown) with another computing device.

[0069] In one embodiment, (not shown), the input images comprise input images of a video conference session and the output images comprise a video that is shared with one (or more) other participant(s) of a video conference session. In one embodiment, the VTO application is a component or plug-in of a video conference application (not shown) allowing the user of the device 402 to wear makeup during a video conference with one or more other conference participants.

[0070] In one embodiment, as described in more detail below, the VTO rendering pipeline component 416 is configured to apply object stabilization to stabilize respective locations of detected objects between, for example, successive frames of a video.

[0071] [Fig. 5] is a flowchart of operations 500 as for a computer-implemented method. The method may comprise executing by one or more processors the steps shown in [Fig. 5] for example. In an embodiment 11: the method comprises: A step 502 which shows processing an input image using a face tracking engine having at least one deep neural network to determine i) facial features from the image input to render an effect and ii) a presence of an occluding object occluding at least a portion of the face; and a step 504 which shows the avoidance of rendering at least a portion of the effect relating to at least one of the detected facial features in response to the presence of the occluding object as detected.

[0072] Embodiment 12: In embodiment 11, the method comprises at least one of: i) providing a recommendation interface for recommending one or more makeup products to be virtually tried on, each of the products being associated with one or more effects to be rendered in association with one or more facial features; and ii) providing a purchase transaction interface for facilitating the purchase of makeup products.

[0073] Embodiment 13: In embodiment 11 or embodiment 12, the processing of the input image by the face tracking engine provides a segmentation of the occluding object and the step of avoiding rendering at least a portion of the effect is responsive to the segmentation such that at least a portion of the effect occluded by the occluding object is not rendered. Embodiment 14: In embodiment 13, the method includes providing an instruction via a user interface to remove the occluding object to facilitate full rendering of the effect.

[0074] Embodiment 15: In any one of embodiments 11 to 14, the method comprises (e.g.: after the step of avoiding rendering, e.g. by rendering no effect) providing an instruction via a user interface to remove the occluding object to facilitate rendering.

[0075] Embodiment 16: In any one of embodiments 11 to 15, the method comprises receiving and processing an additional image using the face tracking engine for facial feature detection and occluding object detection; and rendering the effect after the presence of the occluding object is no longer detected.

[0076] Embodiment 17: In any of embodiments 11 to 16, the effect is a makeup effect and the method is performed in the context of computer operations providing a virtual try-on experience.

[0077] It will be understood that corresponding system embodiments are disclosed for each of embodiments 11-17, for example where the system includes respective components having computational circuitry configured to perform the operations of the computer-implemented method embodiments. Object stabilization

[0078] Localizing objects using deep neural network processing may result in jitter or other instability between images. In other words, the predicted location of an object in a first image by a DNN may be per- noticeably different from the DNN's predicted location of the same object in a second frame. This is particularly noticeable when the first and second frames are two successive frames of a video and an effect is applied in response to the predicted locations. The effect moves with jitter. Tracking the object between successive frames and rendering an effect on the input frames can result in jitter or motion that does not appear to match the underlying input frames when displayed together.

[0079] In one embodiment, stabilization is applied to the localization of a detected object produced by DNN processing of a current frame. In one embodiment, such as for providing a VTO experience from a "live" stream of video (e.g., a selfie video), each frame (e.g., as successive images) of the video is processed to detect and localize the objects, and to render an effect consistent with a product or service to be virtually tried on. The effect is applied at one or more locations or regions relative to at least one of the detected objects. Before rendering, the locations of the detected objects (e.g., at least those associated with the effects) are stabilized to smooth tracking. These stabilized locations are used to render the effect. The effect may be applied to a stabilized location for a detected object (e.g.,a stabilized eyebrow location), or a region adjacent to one or more detected objects such as an eyelid region adjacent to a stabilized location of a detected eye. In some images, for example when a face mask is worn, not all objects are located.

[0080] The stabilization processing is resource intensive. In one embodiment, detected objects are grouped by importance to the task: i.e., by importance to the VTO experience. In one embodiment, detected object locations relative to the mouth and eyes are stabilized using a mixture of a tracking organ prediction from a current frame and an optical flow prediction for the current frame that is responsive to stabilized locations in a prior frame; and detected object locations relative to the eyebrows, nose, and face contour are stabilized using an exponential moving average filter responsive to a net velocity of an object's facial points over the prior n frames.

[0081] In one embodiment, the stabilization operations comprise: 1. Obtain a facial point prediction from a face tracking engine (e.g., 104A or 104B) as trackerPt. The tracker's prediction trackerPt relates to a current frame at time z of a video. The earlier frame is at time t- 1. The tracker's prediction includes locations of variously detected objects, e.g., one set of facial points per detected object, as shown in Figure 2E. The stabi The aim of the stabilization operation is to produce stabilized facial points Pt per detected object for a current frame. The stabilized facial points (stabilized location) per detected object for a previous frame produced by the stabilization operations are denoted by Pt.\. 2. In one embodiment, facial points received from the face tracker, such as those representing an object outline as shown in Figure 2E, are grouped by object into groups of left eye, right eye, left eyebrow, right eyebrow, nose, outer mouth, inner mouth, and face outline points (e.g., a subset of trackerP for each object). In one embodiment, objects are assigned an importance rating, which in one embodiment is one of two ratings (e.g., higher / lower importance). In one embodiment, stabilization of an object is performed in response to the importance rating using one set of operations for objects of higher importance and another set of operations for objects of lower importance.In one embodiment, stabilization operations performed for higher importance objects are more accurate but also more resource and / or processing intensive than operations performed for lower importance objects. Thus, objects are assigned an importance rating that balances accuracy with device performance criteria (e.g., processing time / memory usage, etc.). In one embodiment, the left eye, right eye, outer mouth, and inner mouth objects are assigned the highest importance rating, and the left eyebrow, right eyebrow, nose, and face contour objects are assigned the lowest importance rating. In one embodiment, eyes and lips are prioritized, for example, because many effects relate to eyes and lips. 3. For higher importance objects: Apply an optical flow function to Pt-i only for higher importance objects to obtain OptFlowP. The optFlow function calculates an optical flow (e.g., a frame rate) for a fixed sparse feature using the iterative Lucas-Kanade method with pyramids (prior frame pyramid and current frame pyramid). (See Bouguet, J.-Y. (1999). Pyramidal implementation of the Lucas Kanade feature tracker. At the time of filing, available at semanticscholar.org). OptFlowP will be understood only for a particular object to represent the predicted facial points for the object for the current frame in response to the stabilized facial points Pt-i (locations) produced for the object in the previous frame. In one embodiment implemented with an optical flow function by OpenCV, the points for all objects of greatest importance are provided together, for example, rather than processing each object separately. 4. For larger objects: Blend using blendingfactor and correct if the distance between the tracker location and the optoflow location is above a threshold: At regular intervals, set blendingFactor = 0.4 (a startValue). Regular intervals can be based on time or a frame count (e.g., approximate time equivalent), e.g., every 1.3 seconds or every 40 frames. Time may be preferred for consistency because the frame count vs. approximate time depends on the processing speed. On each frame, run blendingFactor* — 0.080 (a decayValue). The blending factor is used to control the blending between trackerP and optFlowPt and is reset to the startValue (e.g., 0.4) to avoid having optFlowP drift too far from trackerPt. Over time, optflow points and face tracking organ points may drift.If the mixture produces an abrupt change, then the mixture could lead to a discordant result.

[0082] For each group of mixing points (left eye, right eye, inner mouth, outer mouth):

[0083] 4.a. Blending based on blendingFactor;

[0084] = blendingFactor^monitoring bodyPt + (1 - blendingFactor) * optFlowP t

[0085] 4.b. Blend based on distance - compare pixel distances between corresponding face points of trackerP t and optFlowP f. For the mouth object, for example, compare the corner of the mouth face point of trackerP t to the same face point of optFlowP if trackerP and optFlowP are too far apart, blend to trackerPt:

[0086] / ||oprFLwP- organ of stdviP || \ quantity = min^..................., h0)

[0087] p - quantity* tracking organP t + (1.0- quantity) *p

[0088] where distanceBlendingN orm is a normalization factor for the distance to the point and ( ) 6 is used to make small values ​​even smaller. For example, distanceBlendingNorm is 5 pixels in one embodiment.

[0089] 5. For less important objects: Apply a moving average filter ex potential on the points of the left eyebrow, right eyebrow, nose and contour of the

[0090]

[0091]

[0092]

[0093]

[0094]

[0095]

[0096]

[0097] n He face. 5.a. For each group, the net velocity v is calculated and averaged over the previous n frames. In one embodiment, the velocity calculation uses the tracking organ points for the previous frame and the current frame trackerPt and trackerPet does not use the stabilized points P(.\ for the previous frame. These stabilization points are ultimately used when applying the mixture determined using the result of the velocity calculation. This is because using the tracking organ points would allow operations to capture changes in velocity more quickly, rather than using the stabilized points. In . .oreane trackingP-oreane trackingP, velocity = -- v= velocitysrmv = 5.b. The updated facial point coordinates are calculated using a blending factor a: ( l 1 \ \2 a — max min ----■— *r 1 . 0.1 \ \ transitions peearactor )■ j pt = a* monitoring bodyPt + (1.0- a) ] where transitionspeedFactor is a constant that controls the impact of v, and is 1.5 by default. Thus, in connection with the blending operations 4.a and 4.b, a form of linear interpolation is performed for each of the eye and mouth groups respectively (i.e., for respective facial features of the largest group of facial features). In particular, the two locations (tracker location and optoflow location (second location)) for a respective facial point in the current image are blended according to a blending factor. The blending factor weights the contribution of each of the tracker location and the optoflow location to produce a first blended result.A second blending operation produces the current stabilized location and is responsive to the distance between the two locations (e.g., a distance between pixel coordinates of respective facial points of the tracking organ location and the corresponding facial points of the second tracking organ location) and a distance normalization factor to shift the first blending result toward the tracking organ location. Thus, the blending factor initially blends the tracking organ location and the optiflow location in favor of the optiflow location, itself based on the previous stabilized locations; and applies a correction if the two locations are sufficiently far apart, and outputs the current stabilized location at . from the first mixed result as moved to the tracking organ location.

[0098] In one embodiment, the blending factor for the first blending result varies (decreases) by a max amount, over a period (e.g., a series of frames or for a defined duration), and then the blending factor is reset to the max amount. As the blending factor decreases, the optiflow location becomes increasingly preferred in the blend. The reset serves to realign the blend if the locations have drifted. For the distance-based blending threshold, in one embodiment, the distance normalization factor is 5 pixels.

[0099] Thus, in connection with operations 5.a and 5.b performed for each of the less important groups (archs, nose and face contours), an exponential moving average filter is applied. In the exponential moving average filter, the operations only use points from the previous and current frames. The previous frame points implicitly contain information from older frames due to the iterative application of stabilization across frames. In an alternative approach, not shown, a window of past location values ​​is determined and averaged. For example, points from the current frame and the previous N frames (e.g., N = 3), for a total of N + 1 frames, can be used. The resulting point is calculated as an average of the points across the N + 1 frames. The average can be a weighted average, for example with a higher weight on more recent frames.Weight can further be influenced by speed. For example, a higher speed could place even more weight on the most recent frame.

[0100] However, any method for smoothing time series data could be used as an alternative. Another example could be a Kalman filter, which attempts to estimate the current state by modeling the dynamics of the system (such as predicting the current point using the past velocity) and combining this prediction with the current measurement (the tracking organ point).

[0101] [Fig. 6] is an illustration of a computing environment 600, according to one embodiment. The computing environment 600 is similar to the environment 400, however the VTO application 602 differs from the VTO application 412 in that the VTO application 602 includes a stabilization component 604. Although shown as a component included in the VTO rendering pipeline component 606, the stabilization component 604 may be a separate component. The VTO rendering pipeline component 606 is similar to the component 416, but includes stabilization of detected objection locations for rendering effects relating to the stabilized locations.

[0102] In one embodiment, the component stabilization operations 604 are configured as described with reference to steps 1-5 above.

[0103] The VTO application 602 includes the face tracker 104B with its one or more deep neural networks 106B that is configured for face mask classification, localization, or segmentation, so as to detect the presence of a face mask (or other occluding object) in a facial image. In one embodiment, the VTO application could include a face tracker with one or more deep neural networks that localize facial features but without detecting the presence of a face mask (or other occluding object), for example, similar to the face tracker 104A.

[0104] [Fig.7] is a flowchart of operations 700 as for an im method computer-implemented. The method may comprise performing by one or more processors the steps shown in [Fig.7], for example.In an embodiment 18: the method may comprise step 702 of locating a facial feature in a current frame of a set of frames of a video stream using a face tracking engine having one or more DNNs configured to process the current frame to predict the location of a tracking organ of the facial feature; step 704 showing generating a current stabilized location for the facial feature in the current frame, the generation being responsive to the tracking organ location and prior stabilized locations of the facial feature in prior frames of the video stream; and step 706 showing rendering an effect on the current frame associated with the facial feature in response to the current stabilized location, the effect simulating a product to be tried on as a component of a virtual try-on experience.Although not shown, operations may include providing the current frame and effect as rendered (e.g., as an output image) for presentation.

[0105] Embodiment 19: In embodiment 18, the method comprises at least one of the following: i) providing a recommendation interface for recommending one or more makeup products to be virtually tried on, each of the products associated with one or more effects to be rendered in association with one or more facial features; and ii) providing a purchase transaction interface for facilitating the purchase of makeup products.

[0106] Embodiment 20: In embodiment 18 or 19, the method locates a plurality of facial features and the plurality of facial features are grouped by an importance rating associated with the virtual fitting experience to define a larger group of facial features and a smaller group of facial features and wherein respective current stabilized locations for the plurality of facial features are determined in response to the importance rating to choose between different stabilization operations to balance accuracy with device performance criteria. Embodiment 21: In embodiment 20, the plurality of facial features includes a left eye object, a right eye object, and at least one mouth object grouped as more important facial features and a left eyebrow object, a right eyebrow object, a nose object, and a face contour object grouped as less important facial features.

[0107] Embodiment 22: In any one of embodiments 18 to 21, generating the current stabilized location comprises one of the following operations: operation (a): blending the tracking organ location and a second location for the facial element in the current frame using linear interpolation, the predicted second location being responsive to an optical flow determined for the facial feature using the tracking organ location and a previous stabilized location for the facial feature in an immediately prior structure; or operation (b): applying an averaging to the tracking organ location and the previous stabilized location of the facial feature, the averaging being responsive to an average velocity determined from the tracking organ location and the respective prior tracking organ locations for the facial feature over a set of prior frames.

[0108] Embodiment 23: In embodiment 22: the method locates a plurality of facial features and the plurality of facial features are grouped by an importance rating associated with the virtual try-on experience to define a larger group of facial features and a smaller group of facial features; for an individual facial feature of the larger group, the current stabilized location is generated according to operation (a); for an individual facial feature of the smaller group, the current stabilized location is generated according to operation (b); and rendering renders one or more effects associated with at least a portion of the plurality of facial features using respective current stabilized locations.

[0109] Embodiment 24: In embodiment 21 or 22: operation (a) comprises, with respect to a particular facial feature to be stabilized over a set of frames including the current frame and the immediately preceding frame: blending respective facial points of the tracking member location with corresponding respective facial points of the second location according to a blending factor that weights the contribution of each of the tracking member location and the second location to produce a first blended result; and further blending the first blending result and the respective facial points of the tracking member location to produce the current stabilized position in based on a distance between pixel coordinates of respective facial points of the tracking organ location and corresponding facial points of the second tracking organ location, the further blending moving the first blended result toward the tracking organ location in response to a distance normalization factor.

[0110] Embodiment 25: In embodiment 24, the method comprises initializing the blending factor to a maximum amount, at each frame to be processed, decreasing the blending factor, using the blending factor as degraded during blending, and periodically resetting the blending factor to the maximum amount.

[0111] It will be understood that corresponding system embodiments are disclosed for each of embodiments 18-25, for example where the system includes respective components having computational circuitry configured to perform the operations of the computer-implemented method embodiments.

[0112] In addition to the computing device and method aspects, those skilled in the art will understand that computer program product aspects are disclosed, where instructions are stored in a non-transitory storage device (e.g., memory, CD-ROM, DVD-ROM, disk, etc.), which, when executed, cause a computing device to perform any of the method aspects stored therein.

[0113] Although the computing devices are described with reference to processors and instructions that, when executed, cause the computing devices to perform operations, it is understood that other types of circuitry than programmable processors may be configured. Hardware components including specifically designed circuitry may be used, such as, but not limited to, an application-specific integrated circuit (ASIC) or other hardware designed to perform specific functions, which may be more efficient than a general-purpose central processing unit (CPU) programmed using software.Thus, an apparatus aspect herein generally refers to a system or device having circuitry (sometimes references to computational circuitry) that is configured to perform certain operations described herein, such as, but not limited to, those of a method aspect herein, whether the circuitry is configured via programming or via its hardware design.

[0114] The practical implementation may include some or all of the features described herein. These and other aspects, features, and various combinations may be expressed as methods, apparatus, systems, means for performing functions, program products, and other ways, combining the features described herein. A number of embodiments have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the methods and techniques described herein. In addition, other steps may be provided, or steps may be deleted, from the described process, and other components may be added to or removed from the described systems. Accordingly, other embodiments fall within the scope of the following claims.

[0115] Throughout the description and claims of this specification, the terms "include" and "contain" and variations thereof mean "including but not limited to" and are not intended to (and do not) exclude other components, wholes, or steps.

[0116] Any function, feature, integer, compound, chemical unit or group described in conjunction with a particular aspect, embodiment or example of the invention shall be understood to apply to any other aspect, embodiment or example unless inconsistent therewith. Any features disclosed herein (including the claims, abstract and accompanying drawings), and / or any steps of any method or process so disclosed, may be combined in any combination except combinations where at least some of such features and / or steps are mutually exclusive. The invention is not limited to the details of the preceding examples or embodiments.The invention extends to any new feature, or any new combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings) or to any new feature, or any new combination, of the steps of any disclosed method or process.

Claims

Claims

1. A system comprising: a face tracking engine comprising a deep neural network (DNN) for locating a face in a face input image; a training image generator for producing a training image comprising the face as located, the training image comprising either an occluded training image where an occlusive object is rendered on the face or an unoccluded training image showing the face without the occlusive object, the training image being produced for training the occluded face DNN; a training component for training an occluded face detection DNN with the training image; and a component for configuring a face tracking engine with the occluded face detection DNN as trained such that the face tracking engine classifies and localizes facial features and classifies and / or localizes face occlusive objects.

2. The system of claim 1, further comprising a data store storing a plurality of face input images for use with the face tracking engine and the training image generator to produce a plurality of training images for training the occluded face DNN.

3. The system of claim 1, wherein the training image generator is configured to randomly produce the occluded training image, instead of the unoccluded training image, according to a probability chosen to maximize training of the face mask DNN.

4. The system of claim 3, wherein the probability of producing the occluded training image, instead of the unoccluded training image, is a 55% chance.

5. The system of claim 1, further comprising a face cropping component configured to crop the face from the face input image in response to the localization, and wherein the training image generator generates the training image using the face as cropped.

6. The system of claim 1, wherein the occlusive object for the render includes an isolated occluding object image with a transparent background.

7. The system of claim 6, wherein, before rendering the occluding object on the face as located, the training image generator is configured to perform one or more of: a. resizing the occluding object on the face as located; and b. augmenting the occluding object to maximize face mask DNN training.

8. The system of claim 1, wherein the occlusive object covers at least a portion of the face and wherein the occlusive object comprises any of a face mask, occlusive glasses, a hat, a scarf, a hand or one or more fingers, a smartphone, hair, or a portion of any of these.