SYSTEMS AND METHODS FOR AN EFFICIENT RENDERING PIPELINE FOR MAKEUP, INCLUDING CONFIGURATION / SELECTION AND OPTIONAL APPEARANCES IN VTO UI
The enhanced augmented reality makeup pipeline addresses VTO challenges by parallel processing and stabilization techniques, achieving smoother and more realistic makeup application in video conferencing and teleconsultation.
Patent Information
- Application Number
- FR2024006152
- Authority / Receiving Office
- FR · FR
- Patent Type
- Utility models
- Current Assignee / Owner
- Filing Date
- 2024-06-11
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2034-06-11
AI Technical Summary
Existing virtual try-on (VTO) technologies face challenges in handling facial movement, occlusion, varying face sizes, and high frame rates with low latency in applications like teleconsultation and videoconferencing, leading to imperfections that break the realism of makeup effects.
An enhanced augmented reality makeup pipeline is developed with parallel processing of effect rendering and face tracking, stabilization of facial landmarks using optical flow and exponential moving average filters, and occlusion management to maintain smooth video performance.
The solution improves frame rates by up to 2x, stabilizes facial features, and ensures realistic makeup application even with occlusions, providing a seamless VTO experience in video conferencing and teleconsultation.
Smart Images

Figure 00000039_0000 
Figure 00000039_0001 
Figure 00000040_0000
Abstract
Description
Title of the invention: SYSTEMS AND METHODS FOR AN EFFICIENT RENDERING PIPELINE FOR MAKEUP, INCLUDING OPTIONAL CONFIGURATION / SELECTION AND APPEARANCES IN VTO UI FIELD OF INVENTION
[0001] This disclosure relates to image processing and applications using image processing, such as virtual try-on (VTO) applications, augmented reality applications, virtual reality applications, and applications incorporating features of these, such as video chat, teleconsultation, virtual reality (VR) chat, augmented reality (AR) chat, or videoconferencing applications. More specifically, the application relates to devices, systems, and methods for an effects rendering pipeline, including makeup effects. BACKGROUND
[0002] Thanks to VTO technology, makeup is applied virtually to a photo or live video of the user's face. An example of a VTO application is an e-commerce application that allows customers to virtually try on makeup products to help them make purchasing decisions.
[0003] Another application of this technology includes the digital application of virtual makeup or other visual effects to the user's virtual image during a videoconference event. This provides users with a convenient way to appear in meetings with effects such as applied makeup.
[0004] However, applying this technology to applications such as teleconsultation, VR chat, or videoconferencing faces several challenges. These challenges include facial movement, lip movement (e.g., while speaking), the relative size of the face within the overall image, (partial) facial occlusion, and a sufficient video frame rate with low latency to provide smooth video, among others. For example, users frequently talk or turn their heads during a call. This means that facial tracking and makeup rendering must be more robust to accommodate varying lip movements and large facial rotations.Users may also be further away from the camera, causing their face to appear smaller in the captured video image(s), so the face tracking system must be able to detect small faces (relative to the overall video dimensions). Furthermore, it is not uncommon for part of the face to be obscured during a call. For example... The user may have their hand covering their mouth, or hold a cup that covers part of their face.
[0005] In these scenarios, makeup must not appear in areas where there is occlusion. Furthermore, videoconferencing requires that the video be executed at a sufficiently high frame rate and with low latency to provide a smooth experience, which imposes requirements on image processing time.
[0006] It is desirable to minimize imperfections in the VTO so as not to break the realism of the effect, which would distract other participants in the meeting. SUMMARY
[0007] To meet these challenges, in accordance with the embodiments, an enhanced efficiency augmented reality makeup pipeline has been developed, as well as various techniques to meet the above challenges.
[0008] Device, system, and method embodiments are proposed for streamlining the application of an effect (e.g., a virtual try-on effect (“VTO”)) to an object appearing in a sequence of video frames, such as for video chat, conferencing, or teleconsultation applications. Embodiments and / or features of a user interface, such as those for video chat / videoconferencing or teleconsultation, are also proposed. Features of the user interface (UI) may include live and preview modes, recommendations using predefined appearances, appearance refinement, appearance ordering, and access to an e-commerce purchasing interface.In one embodiment, operations of i) effect rendering and ii) object marker determination are performed in parallel where the effect rendering applies an effect in association with markers determined for the object in order to define a sequence of output video frames with the effect applied.
[0009] In one embodiment, a computer device is proposed comprising a processor and a non-transient storage device storing computer-executable instructions for execution by the processor to cause the computer device to: provide a video chat, video conferencing or teleconsultation application to broadcast video frames to at least one other computer device, the application integrated into a virtual try-on pipeline (VTO) to apply one or more effects to the frames to be broadcast;in which the application is configured to present a graphical user interface (GUI) via a display screen, the GUI having: a preview mode in which to receive input to select one or more effects from a plurality of effect options, the preview mode configured to present the one or more effects applied to a sequence of preview frames via the display screen and without broadcasting the; preview frames to at least one other computing device; and a live mode configured to apply one or more effects and stream the frames to which the effects are applied in live mode.
[0010] In one embodiment, subsets of the plurality of effects are grouped into respective appearance groups, and the GUI provides an appearance selection interface to select one of the appearances in order to select the one or more effects grouped for the appearance.
[0011] In one embodiment, each of the one or more effects is associated with a product and, in response to the selection of one of the appearances, the GUI is configured to present a selected appearance interface having respective sub-interfaces for each respective effect of the one or more effects of the appearance, each respective sub-interface providing i) product information for the respective product associated with the respective effect and ii) a refinement control to refine a selection of the respective effect.
[0012] In one embodiment, the refinement control is configured to select a different product for the respective effect so that the respective effect provides a virtual try-on of the different product. In one embodiment, the respective refinement control is configured to disable the respective effect so that the respective effect is not applied to the frames.
[0013] In one embodiment, each of the one or more effects is associated with a product and, in response to the selection of one of the appearances, the GUI is configured to present a selected appearance interface having respective sub-interfaces for each respective effect of the one or more effects of the appearance, each respective sub-interface providing i) product information for the respective product associated with the respective effect and ii) an immediate purchase command to initiate an e-commerce interface for the purchase of products.
[0014] In one embodiment, each of the one or more effects is associated with a product and the GUI is configured to provide access to initiate an electronic commerce interface for the purchase of products.
[0015] In one embodiment, the plurality of effects is any one of the makeup effects, hair effects or nail effects.
[0016] In one embodiment, the appearance interface commands a presentation of the respective appearances to be selected in order to make recommendations to the user.
[0017] In one embodiment, the order of the respective appearances is determined according to the processing of user information, the user information comprising one or more of the following: a user's existing makeup, a user's face shape, a user's skin tone, or any respective appearance as recently selected. In one embodiment, One or more of the user's existing makeup, the user's face shape, or the user's skin tone are determined by processing the frame sequence.
[0018] In one embodiment, the order is prioritized in response to the user's favorite selection of a respective appearance, such that the preferred appearances are ordered before the other appearances, and the other appearances are ordered according to the processing of user information.
[0019] In one embodiment, the preview mode can be invoked before joining a particular instance of a video chat or video conference, integrated into a video chat or video conference user interface to configure video and audio controls for the instance.
[0020] In one embodiment, the preview mode can be invoked during participation in an instance of a video chat or video conference in order to stop all video frame streaming for the instance.
[0021] A computer program product is proposed comprising a non-transient storage device storing computer-readable instructions which, when executed by a processor, cause a computer device to run a video chat, video conferencing or teleconsultation application to broadcast video frames to at least one other computer device, the application being integrated into a virtual try-on pipeline (VTO) to apply one or more effects to the frames to be broadcast;in which the application is configured to present a graphical user interface (GUI) via a display screen, the GUI having: a preview mode in which to receive input to select one or more effects from a plurality of effect options, the preview mode being configured to present the one or more effects applied to a sequence of preview frames via the display screen and without broadcasting the preview frames to at least one other computing device; and a live mode configured to apply the one or more effects and broadcast the frames to which the effects are applied in live mode.
[0022] A computer-implemented method is proposed comprising: running a video chat, videoconferencing, or teleconsultation application to stream video frames to at least one other computer device, the application being integrated into a virtual try-on pipeline (VTO) to apply one or more effects to the frames to be streamed; wherein the application is configured to present a graphical user interface (GUI) via a display screen, the GUI having: a preview mode in which input is received to select one or more effects from a plurality of effect options, the preview mode being configured to present the one or more effects applied to a sequence of preview frames via the display screen and without streaming the preview frames to at least one other computer device; and a live mode configured to apply one or more effects and broadcast the frames to which the effects are applied in live mode. Brief description of the drawings
[0023] [Fig. 1] The [Fig. 1] is a schematic diagram of the operation of a single-wire effects pipeline according to an embodiment of the prior art.
[0024] [Fig.2A] The [Fig.2A] is an illustration showing a representative frame (an example of a face image) including a face and a background.
[0025] [Fig.2B] The [Fig.2B] is an output illustration of a face tracking organ of an effects pipeline showing a cropped face image and facial point groups according to one embodiment.
[0026] [Fig.3] The [Fig.3] is a schematic diagram of the operation of a double-wire effects pipeline according to one embodiment.
[0027] [Fig.4] The [Fig.4] is a flowchart of face tracking operations according to an embodiment.
[0028] [Fig.5] The [Fig.5] is an illustration of a computer environment, according to one embodiment, such as for performing a virtual fitting.
[0029] [Fig.6] The [Fig.6] is an illustration of a computer environment, according to one embodiment, such as for conducting a video chat or videoconference including an integrated virtual fitting.
[0030] [Fig.7] [Fig.8] [Fig.9] [Fig. 10] [Fig. 11] [Fig. 12] Figures 7 to 12 are illustrations of user interfaces of a video chat / video conferencing application with integrated VTO, according to embodiments. Detailed description Pipeline optimization
[0031] According to prior art embodiments, an effects rendering pipeline (for example, a flow of operations of a computing device), such as one for applying a makeup effect to an image, is generally sequential in nature and involves a single thread of operations. Figure 1 is a schematic diagram of the operations of a single-thread effects pipeline according to a prior art embodiment. The pipeline 100 comprises a single thread 102. A current frame to be processed is designated frame t. An immediately preceding frame that has been processed is designated frame t - 1, and an immediately following frame to be processed is designated frame t + 1. Frame t is received during operations 104 for sequential processing during operations 106, such as by a marker detection component (for example, of a face tracking engine). After that, an effects rendering component processes the effects during operations 108 to produce the frame t for output during operations 110, where in output, the frame t has one or more effects applied, for example.
[0032] The face-tracking device includes, in one embodiment, one or more deep neural networks (DNNs) for this purpose. In one embodiment, each DNN includes a MobileNetV2 backbone, according to another embodiment. In one embodiment, the face-tracking device is adapted to track (i.e., locate) classes of objects related to a face, including a face object itself. An output of such a face-tracking device includes a bounding box, a mask, or other structure for deriving a cropped frame (e.g., a cropped facial image). In a cropped frame, for example, any background in the frame, such as the t-frame, is minimized. Figure [Fig. 2A] shows a representative t frame (an example of a facial image 202), including a face 204 and a background 206. The background may include other portions of the subject as well as non-subject portions.A bounding area 208, represented as a dotted area, shows coordinates for defining a cropped image, containing a face 204 and reduced background content (part of the background 206).
[0033] In one embodiment, the landmark detection component or face tracking organ is configured to determine (e.g., predict) a face's location within a frame and locations within the face of one or more facial landmarks (e.g., semantic facial features). In one embodiment, such features include a face contour (e.g., portions of a jaw, a chin, or both), a nose, an inner mouth, an outer mouth, a left eye, a right eye, a left eyebrow, and a right eyebrow, as shown in [Fig. 2B].
[0034] Figure 2B is an illustration of the output of a face-tracking component (not shown) of an effects pipeline showing a cropped face image 210 of a face 204, as described, and facial points 212 according to one embodiment. Respective groups of facial points 212 comprise, in one embodiment, face contour facial points 212A, eyebrow facial points 212B, and nose facial points 212C, etc., for each object for which facial points are determined. The schematic representation is of an annotated cropped face image 210 with the facial point groups for illustrative purposes. The output of the face-tracking component need not necessarily include an annotated image, and the output may be separate data. The facial points of an individual group are numbered or indexed (e.g., 0, 1, 2...) and help to define the outline of the detected object.The face-tracking organ assigns each point so that it is placed in locations consistent with the outline of the object it represents. By. For example, a particular point can always be located at the right corner of the mouth. In one example, face points are X,Y pixel coordinates, relative to the cropped face image 210 and are associated with respective objects detected from (e.g., one of) the arrays of a face tracking organ.
[0035] Figure 3 is a schematic diagram of the operation of a dual-wire (or multi-wire) effects pipeline according to one embodiment. A first wire 303A relates to rendering and a second wire 303B relates to the detection of facial landmarks.
[0036] Wire 303B is shown during operations 304 receiving frame t. The operations for receiving frames t-1 and t+1 are not shown for brevity. During operations 306A, facial features and optionally an occlusion are detected for frame t. During operations 306B, facial features and optionally an occlusion are detected for frame t+1. Wire 303A shows a rendering of the effects on frame t-1 during operations 308A and an output of such a frame t-1 during operations 310A, as well as a rendering of the effects on frame t during operations 308B and an output of such a frame t during operations 310B. The operational output of operations 306A is provided to operations 308B. Figure 3 shows operations that overlap in time, to achieve parallel processing, so that the effects are rendered on the t-1 frame while the detection of facial landmarks is performed on the t frame.Similarly, facial landmark detection is performed on frame t + 1 because effects are rendered on frame t.
[0037] Pipeline 300 is optimized by having a computing device perform effect rendering (e.g., makeup) and face tracking operations (e.g., object) in parallel in order to increase the frame rate. In theory, this can lead to a frame rate increase of up to 2x, but the actual improvement depends on various factors such as the relative timing for rendering versus tracking, as well as the system time resulting from multithreading. The speedup is modeled by the equation:
[0038] $ = ^rendu^xuivi ^^(^remlu^smn'^time xfxtëme thread processing
[0039] Tracking and rendering components similar to those described in relation to [Fig.1] may be used in an embodiment.
[0040] In one embodiment, the landmark detection component or the face tracking engine further detects facial occlusions, determining, for example, which part of a face is missing. In one embodiment, one or more DNNs are configured, for example by training, to at least classify, locate, or segment an occlusion object (for example, a face mask or other objects). In one embodiment, as noted previously, one or more DNNs are configured as MobileNetV2 DNNs.
[0041] In one embodiment, the resulting face-tracking organ with its deep neural network(s) provides a motor for locating facial features such as for use in an application providing a VTO experience, described further below. Improvements to face tracking; Optimization of facial landmark detection
[0042] Object localization using deep neural network processing can lead to jitter or other instability between frames. In other words, the DNN's predicted location of an object in a first frame may be perceptibly different from the DNN's predicted location of the same object in a second frame. This is particularly noticeable when the first and second frames are two successive frames of a video and an effect is applied in response to the predicted locations. The effect moves with the jitter. Tracking the object between successive frames and rendering an effect on the input frames can result in jitter or movement that does not appear to match the underlying input frames when displayed together.
[0043] In one embodiment, stabilization is applied to the localization of a detected object produced by DNN processing of a current frame. In one embodiment, such as for providing a VTO experience from a "live" video stream (for example, a selfie video, a video conference, or a video chat), each frame (for example, as successive images) of the video is processed to detect and localize objects, and to render an effect consistent with a product or service to be virtually tried. The effect is applied in one or more locations or regions relative to at least one of the detected objects.
[0044] In one embodiment, prior to rendering, the locations of detected objects (for example, at least those associated with effects) are stabilized to smooth tracking. These stabilized locations are used to render the effect. The effect can be applied to a stabilized location for a detected object (for example, a stabilized eyebrow location or a stabilized lip location), or to a region adjacent to one or more detected objects, such as an eyelid region adjacent to a stabilized location of a detected eye. In some images, such as where a face mask is worn, not all objects are localized.
[0045] In one embodiment, stabilization is performed using an optical flow technique (e.g., "optical flow tracking") that predicts the location of an object in a current frame. In one embodiment, facial point locations, such as those output by a face-tracking device as described above, can be stabilized.
[0046] In one implementation, optical flow tracking operations are performed on each frame for the purpose of temporal stabilization of the facial landmarks. Given the image and landmarks of the previous frame and the image of the current frame, the optical flow can predict the location of the landmarks in the current frame. The prediction is combined with the output of the facial landmark model using a stabilization algorithm as described below.
[0047] The stabilization process is resource-intensive. In one embodiment, detected objects are grouped by importance to the task: that is, by importance to the VTO experience. In one embodiment, detected object locations related to the mouth and eyes are stabilized using a combination of a tracking organ prediction from a current image and an optical flow prediction for the current image that is reactive to the stabilized locations in a previous frame; and detected object locations related to the eyebrows, nose, and facial contour are stabilized using an exponential moving average filter in response to a net velocity of the facial points of an object over the previous n frames.
[0048] The following is an embodiment of the stabilization operations Stab-1 to Stab 5a / 5b, which include:
[0049] Stab-1. Obtaining a facial point prediction from a face-tracking engine (e.g., 104A or 104B) such as trackerPt. The trackerPt organ prediction relates to a current frame at time t of a video. The previous frame is at time t-7. The organ-tracking prediction includes locations of variously detected objects, e.g., a set of facial points per detected object, as shown in [Fig. 2B]. Stabilization aims to produce stabilized facial points pt per detected object for a current frame. The stabilized facial points (stabilized location) per detected object for a previous frame produced by the stabilization operations are denoted by P
[0050] Stab-2. In one embodiment, the facial points received from the face-tracking organ, such as those representing an object contour as shown in Figure 2E, are grouped by object into groups of points for left eye, right eye, left eyebrow, right eyebrow, nose, outer mouth, inner mouth, and face contour (e.g., a subset of trackerPt for each object). In one embodiment, objects are assigned an importance rating, which, in one embodiment, is one of two ratings (e.g., higher / lower importance). In one embodiment, object stabilization is performed in response to the importance rating using one set of operations for higher-importance objects and another set of operations for lower-importance objects. In one embodiment, the stabilization operations performed Operations for objects of greater importance are more precise but also more resource-intensive and / or processing-intensive than those performed for objects of lesser importance. Therefore, objects are assigned an importance rating that balances precision with device performance criteria (e.g., processing time / memory usage, etc.). In one embodiment, the left eye, right eye, outer mouth, and inner mouth objects are assigned the highest importance rating, while the left eyebrow, right eyebrow, nose, and face outline objects are assigned the lowest importance rating. In one embodiment, the eyes and lips are prioritized, for example, because many effects relate to the eyes and lips.
[0051] Stab-3. For objects of greater importance: Apply an optical flow function to pU only for objects of greater importance to obtain optFlowPt. The optFlow function calculates an optical flow (e.g., a frame rate) for a fixed sparse feature using the iterative Lucas-Kanade method with pyramids (backward frame pyramid and current frame pyramid). (See Bouguet, J.-Y. (1999). Pyramidal implementation of the Lucas Kanade feature tracker. At the time of filing, available at semanticscholar.org). It will be understood that optFlowPt is such that a particular object represents predicted facial points for the object for the current frame in response to the stabilized facial points (locations) pt_i produced for the object in the back frame.In one embodiment implemented with an optical flow function by OpenCV, the points for all objects of greatest importance are provided together, for example, rather than processing each object separately.
[0052] Stab-4. For objects of greater importance: Blend using blendingfactor and correct if the distance between the tracking organ location and the optoflow location is above a threshold: At regular intervals, set blendingFactor = 0.4 (a startValue). The regular intervals can be based on time or a frame count (e.g., a rough time equivalent), for example, every 1.3 seconds or every 40 frames. Time may be preferred for consistency because the frame count versus the rough time depends on the processing speed. On each frame, execute blendingFactor *= 0.080 (a decayValue). The mixing factor is used to control the mixing between trackerPt and optFlowPt and is reset to the starting value (startValue) (e.g. 0.4) to avoid having a drift of optFlowPt too far from trackerPt.Over time, OptFlow points and face tracking organ points can drift. If the blending produces a sudden change, then the blending could lead to a discordant result.
[0053]
[0054]
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068] For each group of blending points (left eye, right eye, inner mouth, outer mouth): Stab-4.a. Blending based on blendingFactor: pt = Mixing Factor * tracking elementPt + (1 - blendingFactor) * optFlowPt Stab-4.b. Blend based on distance - compare the pixel distances between corresponding facial points in trackerPt and optFlowPt. For the mouth object, as an example, compare the facial point at the corner of the mouth from trackerPt to the same facial point from optFlowPt. If trackerPt and optFlowPt are too far apart, blend towards trackerPt. . . the\iptFlowPf monitoring organP^ \ quantity — nuny distance Blending Norm ' Pt = quantity * monitoring organ Pt + (1.0 - quantity) * Pt where distanceBlendingNorm is a normalization factor for the distance to the point and 0#6 is used to make small values even smaller. For example, distanceBlendingNorm is 5 pixels in one implementation. Stab-5. For less important objects: Apply an exponential moving average filter to the points of the left eyebrow, the right eyebrow, the nose and the contour of the face. Stab-5.a. For each group, the net velocity v is calculated and averaged over the preceding frames. In one embodiment, the velocity calculation uses the tracking unit points for both the preceding and current frames (trackerPt and trackerPt-1), and does not use the stabilized points ptJ for the preceding frame. These stabilization points are ultimately used when applying the mixture determined using the velocity calculation result. This is because using the tracking unit points would allow operations to detect velocity changes more quickly than using the stabilized points. Z_, monitoring organPf- monitoring organPpl speed = —— --------n------------: V — Group speed — n[ / tsgmulK Stab-5.b. The updated facial point coordinates are calculated using a mixing factor a: d — max( min^ transitionSpeedFactor ' 1), 0.1 j Pt = a * monitoring organPt + (1.0 - a) * PtA where transitionSpeedFactor is a constant that controls the impact of v, and is 1.5 by default. Thus, in relation to the Stab-4.a and Stab-4.b mixing operations, a form of linear interpolation is performed for each of the eye and mouth groups respectively (that is, for respective facial features from the largest group of facial features). In particular, the two locations (tracking organ location and optoflow location (second location)) for a respective facial point in the current image are blended according to a blending factor. The blending factor weights the contribution of each of the tracking organ location and the optoflow location to produce a first blended result. A second blending operation produces the current stabilized location and is responsive to the distance between the two locations (e.g., a distance between the pixel coordinates of respective facial points in the tracking organ location and the corresponding facial points in the second tracking organ location) and a distance normalization factor to shift the first blended result toward the tracking organ location.Thus, the mixing factor initially mixes the tracking organ location and the optiflow location in favor of the optiflow location, itself based on previous stabilized locations; and applies a correction if the two locations are sufficiently far apart, and generates the current stabilized location from the first mixed result as shifted towards the tracking organ location.
[0069] In one embodiment, the mixing factor for the first mixing result varies (decreases) by a maximum amount over a period (for example, a series of frames or for a defined duration), and then the mixing factor is reset to the maximum amount. As the mixing factor decreases, the Optiflow location becomes increasingly preferred in the mixing. The reset serves to realign the mixing if the locations have drifted. For the distance-based mixing threshold, in one embodiment, the distance normalization factor is 5 pixels.
[0070] Thus, in relation to the Stab-5.a and Stab-5.b operations performed for each of the less important groups (eyebrows, nose contour, and face), an exponential moving average filter is applied. In the exponential moving average filter, the operations use only the points from the previous and current frames. The points from the previous frame implicitly contain information from earlier frames due to the iterative application of stabilization across frames. In an alternative approach, not shown, a window of previous location values is determined and averaged. For example, points from the current frame and N previous frames (e.g., N = 3), for a total of N + 1 frames, can be used. The resulting point is calculated as an average of the points across the N + 1 frames.The average can be a weighted average, for example with a higher weight on more recent frames. Weight can also be influenced by speed. For example, a higher speed could place even more weight on the most recent frame.
[0071] However, any method for smoothing time series data can be used as an alternative. Another example could be a Kalman filter, which attempts to estimate the current state by modeling the system dynamics (such as predicting the current point using past velocity) and combining this prediction with the current measurement (the tracking organ point).
[0072] In one embodiment, to accelerate face tracking, an optimization is applied whereby operations of the facial landmark model (for example, a DNN of a face-tracking organ) that predict facial points for landmarks of a current frame are skipped if the optical flow tracking results are determined to be sufficiently accurate. Since the facial landmark model operations are the slowest step in the face-tracking operations, this can result in a significant time saving.
[0073] In one embodiment and with respect to the Stab-1 to Stab-5a / 5B operations thus described, some of their operations may be performed in a different order or omitted as follows. In one embodiment, 1) the optical flow operations of the Stab-3 step are performed before the facial point detection step of Stab-1; and 2) optical flow operations are applied to all objects rather than only high-importance objects because all objects must be tracked if the facial point pattern were to be skipped. This means that Stab-5 with respect to low-importance objects does not need to be performed and that objects are no longer classified as high or low-importance.
[0074] According to one embodiment, instead of running the facial landmark model, the optical flow result can be used directly as the facial landmark result. However, optical flow tracking is not always reliable and errors can accumulate over time, so the facial landmark model is skipped only if the optical flow error is below a certain threshold or if a certain amount of time has elapsed since the last time the facial landmark model operations were performed.
[0075] Optical flow algorithms such as Lucas-Kanade, referenced above, are capable of providing an estimate of the tracking error.
[0076] Figure 4 is a flowchart of face-tracking operations 400 according to one embodiment. In 402, operations detect or track a face area, for example by running a deep neural network model trained for this purpose. In 404, operations execute an optical flow function such as optflow as described herein. In 406, a measure of an optical flow error is evaluated. If the error is considered small (for example, if it is at a threshold level or in below), then via the Yes branch to 408, operations use the optical flow points as the output for face tracking (e.g., for facial points) as described here. If the optical flow error is not low (e.g., relative to the threshold), via the No branch to 410, operations perform facial landmark detection. Such detection can be done by running a deep neural network model as described. In 412, operations perform stabilization of the detected points using the optical flow results as described here.
[0077] The acceleration provided by this optimization depends strongly on the amount of facial motion. The less facial motion there is, the greater the acceleration. The acceleration provided is given by the equation:
[0078] s = _
[0079] Where t other is the time for all steps except the facial landmark pattern, psauter is the probability that the facial landmark pattern is skipped, with psauter approaching 1.0 when there is little facial movement.
[0080] Face-scale invariant optical flow stabilization
[0081] In one embodiment, stabilization parameters are scaled according to the image size. However, in an alternative embodiment, stabilization parameters are scaled according to the face size. In one embodiment, distance or speed-related parameters, such as the "transitionSpeedFactor" or "distanceBlendingNorm" mentioned above, are then defined relative to a reference size. For example, the reference size is a specific image width (in pixels) when it is a function of the image size. The parameters are then scaled according to the ratio between the effective size and the reference size to ensure that the stabilization behaves consistently across different sizes.Rather than scaling the image size, in one embodiment, scaling the face size is preferred, as in a video conferencing application, to handle a wider range of distances between the face and the camera. This allows stabilization to work more consistently for faces both far away and close to the camera.
[0082] Enhanced face area detection for small faces
[0083] Here, a face is considered small if the face area relative to the image area is less than a certain threshold, such as a percentage, but could be a ratio or another measure. In one embodiment, the threshold is < 10%. The face area detector may have difficulty detecting small faces (faces far from the camera). This is because the entire image is reduced to the input size of the face area model before detection, which makes small faces they become even smaller to the point that detection can become inaccurate or be completely missed.
[0084] To improve this, the following algorithm is used:
[0085] FBD1: If a face has been detected previously and the size of the face area was small compared to the image, so crop the input image down to the last known face area with some fill applied to account for movement;
[0086] FBD2: Reduce the input image to the model input size and run the face area model; FBD3: If a face has been detected, save it as the last known face area.
[0087] FBD4: If a face has not been detected previously, but is detected in an uncropped image and the face area is small (i.e., it is a first detection and the face area is small), then rerun the detection with cropping (for example, using FBD1 to FBD3). This is to handle the case where the first frame contains a small face.
[0088] FB5: If no face was detected and a crop was applied, rerun the detection without the crop. This is to handle the case where rapid facial movement could cause the face to move outside the cropped area.
[0089] Heuristic for redetecting face area using face angles
[0090] In one embodiment, face area detection can be omitted for successive frames to speed up processing, assuming that the face (or the camera) does not move significantly. Face area detector operations are thus skipped for most frames, and the face area is estimated for the new frame based on the facial landmarks from the previous frame. In one embodiment, there are several heuristics used to determine when the face area detector is re-executed, such as: i) whether a face has been detected previously; ii) the amount of time elapsed since it was last executed; or iii) whether the face has been lost.
[0091] Predicting the face area of the current frame using the results of the previous frame works well if the face angle does not change significantly. However, the prediction can be poor if the face angle changes significantly. In one embodiment, an additional heuristic is incorporated where the face area detector is re-executed if there has been a significant change in the face angle (e.g., greater than a threshold) since the last time the detector was executed. In one embodiment, a 3D face rotation (e.g., consisting of yaw, roll, and pitch angles (6 degrees of freedom (6DOF))) is estimated from the 2D facial points. For example, a static reference 3D face model is used and pairs of Points in 2D / 3D are formed and passed to a Perspective-n-Point (PnP) solver. The PnP solver can estimate translation and rotation in 3D using known techniques. Once the face rotation is obtained, the angles are converted into direction vectors. The angle between the current and previous direction vectors is calculated and compared to a threshold (for example, set at 10 degrees). Improvements to rendering Occlusion management
[0092] Occlusion handling renders the effect in response to occlusion detection, for example, to avoid rendering an effect where at least a portion of the object to which the effect is to be applied is not present in the frame because another object appears in front. An example is avoiding rendering an eye effect when the hair has fallen out and occludes at least a portion of the eye. A mask representing the occlusion can be used to render portions that are not occluded. In one embodiment, partial occlusion may not trigger rendering on any portion of the occluded object. Mask occlusion and binary or Boolean occlusion are discussed further below.
[0093] To manage occlusion, the rendering engine is configured to selectively hide different parts of the makeup. In one embodiment, the approach to hiding the makeup depends on the type of occlusion information available. The occlusion information is comparatively coarser or more granular in the embodiments below. Binary region occlusion
[0094] In this approach, the face is divided into regions (such as left eye, right eye, mouth, etc.) and a Boolean value is assigned to each region to indicate whether that region is occluded.
[0095] The rendering engine is modified to divide the makeup rendering into different occlusion regions and render only unoccluded regions. Mask occlusion
[0096] In this approach, a mask is provided that indicates which pixels of the face are occluded. The mask region could cover certain areas of the face (such as the lips or eyes), or it could cover the entire face. In one embodiment, the mask region is where the occlusion mask should be placed and is specified as a rectangular region of the face (for example, it could be a rectangle covering the entire face or covering only the lips / eyes). The occlusion mask is an image where each pixel indicates the probability of occlusion, and this mask is placed and fitted onto the mask region.
[0097] After the makeup rendering, the rendering would blend back into the original image using the occlusion mask as the blending quantity.
[0098] Improved error tolerance in lipstick rendering
[0099] Different facial tracking organ markers on the lips may have different expected levels of accuracy. For example, lip points near the teeth may have lower average accuracy than points on the outer edge of the lips. In one embodiment, inaccuracies are made less noticeable in the rendering by blurring. For example, when rendering lipstick, the edges of the lipstick may be blurred to hide imperfections in the lip markers. Furthermore, different amounts of blurring may be applied to different areas of the lipstick, according to one embodiment, and the amount of blurring may depend on the expected accuracy of the markers in that area. This allows the rendering to better handle less precise areas without making the lipstick appear blurrier overall.Two approaches are proposed to determine accuracy: One approach consists of measuring the average error for each landmark after training the facial landmark model, comparing the predicted landmarks with the ground truth landmarks from the training data. Another approach consists of testing the face tracking device on various individuals and providing a qualitative rating of the accuracy of the different landmark determinations. VTO application.
[0100] Figure 5 illustrates a computing environment 500, according to one embodiment, such as for practicing one or more aspects of a method, for example, including, but not limited to, VTO operations. The computing environment 500 shows a user computing device 502, such as a smartphone, a communication network 504, a server 506, and a server 508. The communication network 504 includes wired and / or wireless networks, which may be public or private and may include, for example, the Internet. The server 506 includes a server computing device such as for providing a website. The server 508 includes a server computing device such as for providing e-commerce transaction services. Although shown separately, the servers 506 and 508 may comprise a single server device. The computing environment is simplified.For example, payment transaction gateways and other components such as those used to perform an e-commerce transaction are not represented.
[0101] The computing device 502 includes a storage device 510 (for example, a non-transient device such as memory and / or an integrated circuit disk (SSD), etc.) for storing instructions which, when executed by a processor (not shown), cause the computing device 502 to perform operations such as a computer-implemented method. The storage device 510 stores a virtual fitting application 512 comprising components such as software modules providing a user interface 514, a face tracking organ 516 with one or more deep neural networks 518 configured for face detection including the determination of facial points, a VTO rendering pipeline component 520 with a stabilization component 522, a product recommendation component 524 with product data 526, and a purchase component 528 with a basket 530 (e.g. purchase data).
[0102] In one embodiment, face tracing and effects rendering operations (for example by components 516 (face tracking organ) and 520 (VTO rendering pipeline) are carried out in parallel as shown and described here (see [Fig.3]).
[0103] In one embodiment, the VTO application is a web application as obtained from the 506 server. Although not shown, the 502 user device may store a web browser for running the web-based VTO 512 application. In one embodiment, the 512 VTO application is a native application conforming to an operating system (also not shown) and to the software development requirements that may be imposed by a hardware manufacturer, for example, of the 502 user device. The native application may be configured for web-based or similar communication to the 506 and 508 servers, as is known.
[0104] Figure 5 shows various input and output data or information associated with the use of the VTO application 512, for example. This includes an input image 540 of the user to be processed for a VTO experience, an output image 542 to which product effects are simulated providing a VTO experience, a VTO product selection 550 comprising user input selecting one or more product effects to be simulated, VTO product options 552 comprising options for products to be virtually tried, for example for selection by a user of the device 502, and purchase transaction information 560 comprising purchase information provided to and / or received from a user to purchase a product.
[0105] In one embodiment, via one or more of the user interfaces 514, VTO product options 552 are presented for selection to be virtually tried by simulating effects on an input image 540. In one embodiment, the VTO product options 552 are derived from, or associated with, product data 526. In one embodiment, the product data 526 may be obtained from the server 506 and provided by the product recommendation component 524. Although not shown, user or other input may be received for use in determining recommendations. of product. The user may be prompted, for example via one of the 514 interfaces, to provide input to determine product recommendations. In one embodiment, the product recommendation component 524 communicates with the server 506. The server 506, in one embodiment, determines the recommendation based on the input received via the 514 component (for example, 524) and provides product data accordingly. The 514 user interface may present the VTO product choices, for example, by updating its display in response to the data received when the user navigates or otherwise interacts with the 514 user interface.
[0106] In one embodiment, one or more user interfaces 514 provide instructions and controls for obtaining the input image 540 and the VTO product selection input 550, such as identifying one or more recommended VTO product options 552 to try. In one embodiment, the input image 540 is a user's face image, which may be a still image or a frame from a video. In one embodiment, the input image 540 may be received from a camera (not shown) of the device 502 or from a stored image (not shown). The input image 540 is provided to the face-tracking organ 516, such as for processing to detect objects in the facial image using one or more trained deep neural networks 518. In one example, the network classifies, locates, or segments a face mask (or other occlusive object) in the image.In one embodiment, a face mask presence classification example is useful for outputting a request (e.g., an instruction to a user, for example via 514 user interfaces) to lower or remove a face mask. This can be applied to any occlusive object for which the face tracking engine is trained. In one embodiment, the occlusion can be manipulated during rendering, as described here, to avoid rendering on an inclusion.
[0107] In one embodiment, an output (specifically not shown) from the face-tracking organ 516, such as classification results, localization results, or segmentation results for one or more detected objects, is provided to the VTO rendering pipeline component 520. In one example, the output may include a bounding box, for example 208 of [Fig. 2A], and, as shown in [Fig. 2B], facial points 212 (for example, groups thereof) for detected objects, etc. The input image 540 is also provided (for example, made available) to the VTO rendering pipeline component 520. The VTO product selection 530 is also provided to the VTO rendering pipeline component 520 to determine the effects to be rendered. In a makeup simulation embodiment, one or more effects may be specified, such as for one or several product categories including: lips, eyeshadow, eyeliner, eyeshadow, etc.
[0108] In one embodiment, the VTO rendering pipeline component 520 determines whether to render one or more product effects on the input image 540 to simulate a try-on. For example, in response to the face mask classification output, the VTO rendering pipeline component 520 may decide not to render a product effect, for example, because a mask (an occlusion) is detected. When a face mask is detected, for example, the VTO rendering pipeline component 520 may trigger the user interface 414 to prompt the user to remove the face mask. A new image (a new instance of the image 540) may be received and processed by the face-tracking component 516. In one embodiment, images are received continuously as a component of a live stream (for example, a selfie video).In one embodiment, occlusions are handled during rendering in such a way as to avoid rendering on an inclusion, as described here.
[0109] If the VTO rendering pipeline component 520 determines to render one or more product effects, in one embodiment, the VTO rendering pipeline component 520 renders effects on the input image 540, such as by drawing (rendering) layered effects, one layer for each product effect, to produce an output image 542. Portions of the operations of the VTO rendering pipeline component 520 (for example, drawing the layers) may be performed by a graphics processing unit, in one embodiment. The rendering conforms to the product data 526 as selected by the VTO product selection 550 and is responsive to the location of detected objects. For example, a VTO product selection of a lipstick, lip gloss, or other lip-related product calls for the application of an effect to one or more detected mouth or lip-related objects at their respective locations.Similarly, a selection of eyebrow-related products invokes the application of a selected product effect to the detected eyebrow objects. Generally, for symmetrical appearances, the same eyebrow effects are applied to each eyebrow, the same lip effect to each lip, or the same eye effect to each eye region, but this is not always the case. In one example, the rendering is applied to a region that is relative to the detected objects, such as one or more adjacent detected objects. Some VTO product selections include a selection of more than one product (for example, defining an "appearance"), such as coordinated products for eyebrows and eyes, or other combinations of detected objects, including the entire face.Product data can define respective "appearances" grouping related products, for example, and associating the appearance with a name to be displayed via the user interface, such as displayed associated with a command enabling. The user selects an appearance from a group of appearances presented in a list, table, or other presentation format. The VTO 520 rendering pipeline component can render each effect, for example, one at a time until all effects are applied. The order of application can be defined by rules or in the product selection, for example, a lipstick before a lip gloss.
[0110] In an embodiment where an occlusive object is detected and its location is determined, for example, as represented in a segmentation mask, the rendering can be responsive to such a segmentation mask. An effect rendering can be applied to portions of the face that are not occluded. A segmentation mask can indicate which pixels of the face are available to receive an effect, such as a makeup effect, and which pixels are not available to receive an effect.
[0111] User interface 514 provides output image 542. Output image 542, in In one embodiment, output image 542 is presented as a portion of a live stream of successive output images (each like Example 542), such as where a selfie video is augmented to present an augmented reality experience. In one embodiment, output image 542 is presented together with input image 540, as in a side-by-side display for comparison. In one embodiment, output image 542 may be saved (not shown), for example on storage device 510, and / or shared (not shown) with another computing device.
[0112] In one embodiment, the input images (not shown) include input images from a videoconference session, and the output images include a video shared with another (or several other) participant(s) in a videoconference session. In one embodiment, the VTO application is a component or extension module of a teleconsultation application or a videoconferencing application (each not shown) that allows the user of device 502 to wear makeup during a teleconsultation or videoconference (respectively) with one or more other conference participants.
[0113] In one embodiment, the VTO 520 rendering pipeline component is configured to apply object stabilization (for example, using a 522 stabilization component) to stabilize respective locations of detected objects between, for example, successive frames of a video.
[0114] Although the 522 stabilization component is shown as a component included within the 520 VTO rendering pipeline component, the 522 stabilization component can be a separate component. In one embodiment, the operations of the 522 stabilization component are configured as described with reference to Stab-1 through Stab-5b operations above. Stabilization and tracking operations of The faces of the bitters, in one embodiment, are further in accordance with the operations of [Fig.4].
[0115] In one embodiment, the face-tracking component 516 locates facial features but without detecting the presence of a face mask (or other occlusive object). Consequently, in such an embodiment, the operations of the VTO rendering pipeline component 520 are configured without taking occlusions into account.
[0116] Figure 6 illustrates a computer environment 600, according to one embodiment, for conducting a teleconsultation, video chat, or videoconference with integrated virtual fitting. Environment 600 is similar to environment 500. In environment 600, a user device 602 provides a teleconsultation or videoconferencing application 604 having integrated VTO features. The application 604 is stored in the storage device 606 and is shown in a simplified manner. Integrated VTO features are provided, for example, by the VTO application components 512 as described further below.
[0117] Device 602 is configured to also communicate with server 608, which provides videoconferencing services, in order to communicate with one or more other user devices, such as, but not limited to, device 610, device 612, or both devices. Examples of platforms providing videoconferencing services, which are not exhaustive, include Microsoft Teams™ available from Microsoft Corporation of Redmond, WA; Zoom One™ available from Zoom Video Communications, Inc. of San Jose, CA; and Google Meet™ available from Google LLC of Mountain View Parkway, among others. In one embodiment, server 608 can be configured to provide the functionality of one or both servers 506 and 508.
[0118] In short, teleconsultation or videoconferencing services allow the sharing of live video between two or more user devices communicating via an intermediary device, namely a server. A first user device (for example, 602) obtains a video stream from a camera (either an internal or external camera coupled to it) and provides it to the server 608 for communication with other participating devices such as device 610, device 612, or both devices (for example, and their respective users (for example, conference members in a videoconference, a clinician, or a beauty consultant in a teleconsultation) who are participating in the conference as maintained by the server 608. The server 608 provides the respective video streams received from device 610, device 612, or both devices to device 602.It is understood that server 608 can process (for example, perform video processing of) any of the video streams it receives and retransmits for a conference or teleconsultation.
[0119] Respective user teleconsultation or videoconferencing applications running on respective devices present the received video streams as in accordance with a layout or view selected in a user interface on a display device. A layout or view may display a member who is the active speaker or a pinned conference member or all conference members, etc., as is known.
[0120] In one embodiment, the application 604 is configured to apply at least one effect to the images from the device 602, allowing a virtual fitting during the teleconsultation or videoconference meeting, so that other members receive the output images as rendered using the integrated VTO application with the at least one effect applied.
[0121] In [Fig. 6], the input image 540 represents a frame of an input video stream from a camera local to the device 602, while the output image 542 represents a frame of an output video stream determined from one or more frames (e.g., 542) of the input video stream. Each output image 542 is presented according to the user interface or other commands of the application 604. Thus, sometimes during a teleconsultation or conference, the output image 542 may not be displayed by the device 602, such as when another member has a focus of attention and only that member's stream is presented. However, the output image 542 is communicated to the server 608 for retransmission for display by either the devices 6100, 612, or both, according to the respective commands of their local teleconsultation or videoconferencing applications.It is understood that no VTO effect is applied if the camera command is "disabled" and no camera image is shared with the 608 server.
[0122] In one embodiment, the application 604 is configured with user interfaces having controls to allow a user to select whether to apply a VTO effect. In one embodiment, the VTO effect is selectable from a plurality of appearances. In one embodiment, each appearance offers at least one makeup effect and preferably a plurality of effects for different makeup products. In one embodiment, appearances are associated by product brand, and each appearance has a name. In one embodiment, the user interface is activated to receive user input to select a preview of an appearance, invoking the VTO components to process the input video stream and render an output video stream with the rendered appearance effect(s) for display by the device 602.In one embodiment, during the preview, the output video stream is not shared with the 608 server and is therefore not provided to other devices during the preview period. In one embodiment, the user interface is enabled to display detailed information. on each of the products' appearance and is further activated to allow the purchase of products.
[0123] Figure 7 illustrates a portion of a user interface (UI) of a videoconferencing application with integrated VTO, according to one embodiment. Figure 7 illustrates a portion of a graphical user interface, such as a portion of a window or other interface construct. As is known, graphical elements present various information and can delimit or identify regions for a user to control within the interface the regions that define respective portions of the interface (for example, dividing the space of a display). A graphical element in a particular region can define a user interface control (or command) associated with a software function.When a command is triggered by user input, the command (i.e., the trigger) can invoke associated user interface code, including a portion of an underlying application (software), to perform a function associated with the command. Examples of commands include fields for entering information such as alphanumeric text, icons to invoke an associated action (e.g., muting a microphone), menus (including a drop-down list or click overlay, etc.) to provide additional commands or submenus, and slider or radial controls to select a data value (e.g., adjusting the speaker volume).In one embodiment, the interfaces shown here, such as in any of Figures 7 to 12, are gesture-activated to receive gesture-based user input via one or more applicable hardware input / output (I / O) devices. In one embodiment, touch-type gestures (touching, tapping, swiping, etc.) can be received via a touchscreen or via a camera receiving user input images. In one embodiment, a mouse or other pointing device (e.g., a pen) provides input. Key-based input, for example from a physical or virtual keyboard, can also be enabled.
[0124] In one embodiment, when users start the video conferencing experience, they first enter a window where they can adjust their video and audio settings. Here, they may have the option to open a video effects menu and navigate to a region of the interface where they can select VTO makeup appearance effects.
[0125] Interface 700 allows user input for selecting audio and video settings, including video effects. Interface 700 can be presented, for example, before joining a conference. Interface 700 includes region 702 for displaying a video image of the user when the camera is on.
[0126] In region 708, a plurality of effect options are presented, including 708A filters, none of which are currently being selected, as noted by command 708B in a subregion of 708. A plurality of 708C commands provide a selection of associated filter groups such as VTO, Background, Blur.
[0127] Figure 7 shows, in the embodiment, a plurality of VTO 708E filter commands that are displayed (e.g., as a menu) in a subregion of region 708 (e.g., below the 708C commands). The 708E filters are selected from the filter options associated with the 708C commands using command 708F. The selection invokes their display in the subregion of region 708. In the embodiment, command 708F is highlighted in bold to show that it is active. Bold type is used here to facilitate patent line drawings; however, color, shading, etc., are useful for highlighting an actual user interface.Region 708 can be scrolled to bring more filters into the display, bringing filters theoretically below the background of region 708 into the user interface and moving thumbnails theoretically out of the user interface to the top near the 708C controls.
[0128] Each of the respective filter controls 708E comprises a selectable portion of region 708 (e.g., a sub-subregion). In the embodiment, each is shown as a card or thumbnail (e.g., 708F) according to common user interface techniques. Another style of control is an icon (not shown). The illustration in [Fig. 7] is simplified. In one embodiment, each label (e.g., "Visual" in thumbnail 708F) is replaced by an image showing a human face and / or one or more additional body parts to which at least one effect is applied to suggest the results of selecting that particular filter. In one embodiment, the effect is a makeup effect, a hair effect, a nail effect, or any combination thereof.These effects are associated with their respective VTO filter effects, which can be invoked by selecting the thumbnail to activate the selected VTO filter for video conferencing or VR chat. In one implementation, the VTO filter can be previewed (for example, in a "preview mode") so that the results (for example, a video stream with the VTO filter effect applied to modify the user's appearance) are displayed in region 702, but the results are not provided to the 608 server. A command can be provided to allow the user to exit preview mode and use the selected VTO filter in a "live mode" where the results are provided to the 608 server.
[0129] In one embodiment, the thumbnails (filter commands 708E) are associated with respective "appearances," each comprising one or more makeup effects. An individual appearance may include an effect of Makeup, hair, or nail effects, or any combination thereof. Each appearance is associated with a name and source of the product(s) on which the VTO effect is based. Generally, the products share a common source represented by a brand (e.g., a trademarked brand). One example of an appearance and brand, which is not exhaustive, is Glossy by Maybelline™ (available from L'Oréal, specifically L'Oréal USA Creative, Inc. of New York, USA). Selecting (e.g., by clicking, etc.) one of the commands, such as 708G from the 708E commands, invokes the application of the associated VTO filters for the appearance in order to apply the appearance to the user's video in preview mode. In a real-world scenario, joining a conference stops preview mode and enters live mode.
[0130] Figure 8 is an enlarged portion 800 (within the dotted area) of the interface 700, although this is in response to user input to that effect. Thumbnail 708G is highlighted in bold as a selected thumbnail 708G from among the commands 708E. Figure 8 shows information 802 associated with thumbnail 708G. The information 802 provides textual information, for example, about the filter associated with thumbnail 708G. In one embodiment, the information provides a name for the appearance and brand, for example, Glossy by Maybelline, and indicates the preview mode. Selection invokes an overlay of the thumbnail to display ellipses (for example, three dots with an associated command) 804 prompting further action (for example, further user engagement, such as by tapping, etc.).The symbol 806 reflects a filter state, namely whether the filter has been downloaded to the local computing device 602 from a server such as server 608, or from another server. In one embodiment, the filter includes various effects data to be used to simulate the product, as well as additional images, text, etc., for user interfaces that provide further product information, as described later. The filter can be defined in a dataset for downloading and unpacking for inclusion in the user interface and for VTO operations. In one embodiment, as shown in [Fig. 7], newly made-available filters are sorted to appear at the top of a command table (e.g., at 708E) displaying the filters. Other scheduling operations are possible, as described later below.
[0131] A further selection of the thumbnail (via the ellipse) modifies the content of the user interface displayed in region 708 to show additional appearance information (e.g., in preview mode). [Fig. 9] shows the portion of user 700 in which region 708 presents a menu 900 comprising expanded thumbnails (e.g., 902A), each relating to the thumbnails 708E, according to one embodiment. Menu 902 can be scrolled to bring additional expanded thumbnails into the display, while removing others. Command 902A is associated with command 708G in that each represents the same appearance and VTO effects. Command 902A is selected and active on [Fig. 9]. A user interface element containing the text "Learn More" is associated with command 902B, which, when invoked, modifies region 708 to provide more information about the particular appearance associated with command 902A.
[0132] Figure 10 illustrates the portion of user interface 700, according to one embodiment, showing appearance information in a menu 1000 in region 708 for a representative appearance (e.g., a selected appearance interface with sub-interfaces for each product / effect to be applied), for example, by invoking command 902A. Menu 1000 includes tile 902A with the user interface element and associated command 902B removed, as well as a plurality of tiles (e.g., product tile 1000A) providing information on the respective products (e.g., Product 1 to Product 5). Each tile shows the product type (e.g., lipstick, mascara, eyeliner, eyebrows, foundation, blush, eyeshadow, etc.).), the product name (which may be a trademark (e.g., Maybelline Fit Me™ foundation) and a color or other characteristic of the product (which may be proprietary, such as the Molten Rose Gold™ color of Master Chrome Highlighter™ available from Maybelline).
[0133] Although not shown, in one embodiment, the product thumbnails (e.g., 1000A) in menu 1000 are each associated with a "Buy Now" command (e.g., a button command). The user interface is equipped with the button command, so that when it is clicked or otherwise invoked by user input, the user interface receives the user input and offers the user access to a branded product page (e.g., via a web browser) to purchase the actual product. The product can be purchased via an e-commerce transaction. In one embodiment, a single "Buy Now" button forms links from menu 1000 to an appearance-based product page from which the respective products can be purchased.In one embodiment, the command invokes a web-based interface, for example, a link to a web page via the internet for execution and display by a web browser. It would be desirable if the user interface could be configured to provide access to an e-commerce interface in other ways, such as a tab, a drop-down menu, or a right-click menu while hovering over a selection. product / effect, etc. Effects (and products) can include makeup, hair or nail effects (and products), for example.
[0134] Figures 7 to 10 illustrate an example of a UI operation flow for configuring the application of a VTO filter including an appearance, for example, during a pre-conference or pre-VR chat start period (providing an embodiment of a preview mode) to initiate the conference or VR chat with an applied appearance. [Fig. 11] is an illustration of a user interface 1100 of a video conferencing application with integrated VTO, according to one embodiment, showing a drop-down menu 1102 selected for configuring the application of a VTO filter during a video conference (for example, while video is being provided to the server 608 and / or received from the server 608).
[0135] UI 1100 displays a window, in one embodiment, in which, compared to [Fig. 7], regions 702 and 708 are deleted and replaced by region 1104, which displays a live view of the user video stream. Region 1106 is located below region 1104 and provides a menu containing a plurality of menu items, including a plurality of icons, each icon being associated with a command. One such icon includes an ellipse, 1106, defining yet another command to invoke drop-down menu 1102. Menu 102 displays a plurality of menu items, including menu item 1102A. Menu item 1102A is highlighted in bold as the selected item to invoke a filter selection.
[0136] Figure 12 illustrates user interface 1100 according to one embodiment, as invoked from drop-down menu 1102. Figure 12 shows a region 1202 having a VTO filter menu like the one shown in region 708 of Figure 7 for configuring the application of a VTO filter including an appearance during a video conference. Region 1104 has been reduced in size in user interface 1100 compared to Figure 11 to provide space for region 1202. Invoking the video effects selection sub-interface of region 1202 puts the video stream into preview mode. The user video stream (e.g., with VTO effects applied) is not provided to server 608. Region 1202 displays information 1202A notifying of preview mode. Selecting a thumbnail (e.g., 1202B) selects the VTO effect(s) of the appearance for preview in region 1104.Selecting the close command (e.g., "X" 1202C) closes the video effects selection sub-interface of region 1202, returning the video to live mode with the associated VTO filter(s) of the appearance applied to the user's video. The user's video image is expanded within interface 1100, and the video with effects is provided to server 608.
[0137] In one embodiment, the order of the thumbnails for respective appearances as shown in Figures 7-9 and 12 (for example, ordering of the The content for commands 708E is determined by operations to make user-oriented recommendations. In one embodiment, the operations process available user information and prioritize (i.e., order) (e.g., content associated with the respective thumbnails / commands), for example, according to rules or other operations to present appearances, so that the most recommended ones are placed at the top (closer to commands 708C, for example). In one embodiment, user information for determining an appearance order encompasses a variety of factors, such as users' existing makeup, face shape, skin tones, and appearances they have recently selected (i.e., based on the most recent order of previous selections).In one implementation, a makeup purchase history is also used when available.
[0138] In one embodiment, an order for respective appearances, as shown in Figures 7-9 and 12, is responsive to a user's favorite selection. For example, an option is provided allowing the user to mark one or more appearances as favorites. These selected favorite appearances are prioritized from the top. In one example, the favorite selection of an appearance overrides and replaces other operations for processing user information (e.g., no selection).In one embodiment, for a plurality of favorite selections, operations can order such favorite appearances using additional user information processing - favorite selected appearances over unselected appearances, where favorite selections are ordered among themselves based on user information processing, and unselected appearances are ordered based on user information processing.
[0139] In one embodiment, an alternative product selection is activated. For example, in [Fig. 10], menu 1000 is activated to open an expanded area under a selected product thumbnail (not shown) or to open a context or embedded menu. A list of alternative products is provided with associated commands so that a user can select a product to replace the product in the appearance (e.g., in menu 1000 and in the VTO effects to be applied). Selecting an alternative product replaces that product in the appearance of menu 1000. The alternative product selection is saved for use as the associated effect and is saved for reuse, as when the appearance is selected for later use.
[0140] In one embodiment, although not shown, a particular product can be identified (via a checkbox or similar control) to disable the associated filter effect in the conference appearance. For example, a user can Select the eyeshadow (or other product) to disable the associated eyeshadow effect, or re-enable it, etc.
[0141] In addition to the computer device and method aspects, the person skilled in the art will understand that computer program product aspects are disclosed, where instructions are stored in a non-transient storage device (e.g. memory, CD-ROM, DVD-ROM, disk, etc.), which, when executed, cause a computer device to perform any of the method aspects stored there.
[0142] Certain aspects and features will be understood from the following numbered statements:
[0143] Declaration 1: A method for streamlining the application of an effect to an object appearing in a sequence of video frames, the method comprising: a. the parallel performance of i) an effect rendering, and ii) object marker determination, wherein the effect rendering applies an effect in association with markers determined for the object in order to define an output video frame sequence with the effect applied; and b. the provision of the output video frame sequence for display.
[0144] Declaration 2: Method according to Declaration 1, wherein the frame sequence comprises frame t - 1, frame t and frame t - 1 in sequence, and wherein step a. determines object markers for frame t in parallel with the application of the effect to the object in association with object markers determined previously for frame t - 1.
[0145] Declaration 3: Method according to Declaration 2, wherein step a. detects an occlusion of the object and wherein the rendering effect is guided by the occlusion as detected.
[0146] Declaration 4: Method according to Declaration 3, wherein step a. provides object mask information at a pixel level according to the occlusion as detected to guide the rendering of effects.
[0147] Declaration 5: Method according to Declaration 1, wherein object landmark determination provides pixel locations for the object, object landmark determination comprising detecting pixel locations for at least some of the video frames using a deep neural network.
[0148] Declaration 6: Method according to Declaration 5, wherein the frame sequence comprises frame t -1, frame t and frame t + 7 in sequence, and wherein step a. comprises object landmark stabilization for frame t in accordance with a prediction of the location of object landmarks for frame t using an optical flux function.
[0149] Declaration 7: Method according to Declaration 5, wherein the frame sequence comprises frame t-1, frame t and frame t+1 in sequence, and wherein the determination of object landmark comprises the calculation of an optical flux function in relation to frame t to predict locations in frame t in response to locations in frame t-1, the determination of an optical flux error for frame t, the skipping of a detection of pixel locations for frame t in response to the optical flux error and the use of pixel locations in response to the optical flux function.
[0150] Declaration 8: Method according to Declaration 1, wherein, for each of the video frames, the determination of object landmark determines a bounding area within which the object is located, the bounding area comprising a subset of video frame pixels.
[0151] Declaration 9: Method according to Declaration 1, wherein steps a. and b. are carried out by a first computing device and wherein step b. comprises the communication of the output video frame sequence via a communication network for display by at least one other computing device participating in a video chat, videoconference or teleconsultation with the first computing device.
[0152] Declaration 10: A method according to Declaration 1, wherein the method applies respective effects to a plurality of respective objects and step a. performs object detection for each of the plurality of respective objects and effect rendering applies respective effects relative to at least some of the plurality of respective objects.
[0153] Declaration 11: Method according to Declaration 10, wherein the video frame sequence includes a face, the plurality of objects includes respective regions of the face and the respective effects include respective makeup effects.
[0154] Declaration 12: Method according to Declaration 11, wherein the regions include one or more of the following: a left eye, a left eyebrow, a right eye, a right eyebrow, a nose, a mouth, an upper lip or a lower lip.
[0155] Declaration 13: Method according to Declaration 1, wherein the effect includes a makeup effect, a hair effect or a nail effect and wherein the method includes providing a user interface presenting a plurality of makeup, hair or nail effects associated with respective products for selection by user input, and wherein the user interface is configured to provide access to an electronic commerce interface to conduct a product purchase transaction.
[0156] Declaration 14: Method of applying effects to a sequence of input video frames to define a sequence of output video frames for a video chat, video conference, or teleconsultation, the method comprising: receiving the sequence of frames comprising a first frame followed by a second frame; processing the first frame for object markers for at least one object; while applying one or more effects to the first frame to define an output frame, the one or more effects applied with respect to at least some of the object markers, furthermore processing the second frame to determine object markers; and providing the output frame for the video chat, video conference, or teleconsultation.
[0157] Declaration 15: Method according to Declaration 14, wherein the one or more effects include virtual fitting effects (VTO) to simulate one or more products.
[0158] Declaration 16: Method according to Declaration 14, wherein one or more effects include makeup effects, hair effects or nail effects and at least one object includes a part of a user's body.
[0159] Declaration 17: Method according to Declaration 16 comprising the provision of a user interface presenting a plurality of makeup effects, hair effects or nail effects associated with respective products for selection by user input, the user interface being configured to provide access to an electronic commerce interface to conduct a product purchase transaction.
[0160] Declaration 18: Method according to Declaration 14, wherein the method is carried out by a video chat or video conferencing application.
[0161] Declaration 19: Method according to Declaration 14, in which the method is carried out by a teleconsultation application.
[0162] Declaration 20: Method according to Declaration 14, wherein the treatment for object markers detects an occlusion of the object and wherein the application of one or more effects is guided by the occlusion as detected.
[0163] Declaration 21: Method according to Declaration 20, wherein the processing for object markers provides object mask information at a pixel level according to the occlusion as detected to guide the application of one or more effects.
[0164] Declaration 22: Method according to Declaration 14, wherein the processing for object markers provides pixel locations for the object using a deep neural network.
[0165] Declaration 23: Method according to Declaration 22 comprising stabilizing object landmarks for the second frame according to a prediction of the location of object landmarks for the second frame using an optical flux function.
[0166] Declaration 24: Method according to Declaration 22, wherein the processing for object markers includes the calculation of an optical flow function relative to the second frame to predict locations within the second frame in response to locations in the first frame, the determination of an optical flow error for the second frame, the skipping of a detection of pixel locations for the second frame in response to the optical flow error and the use of pixel locations in response to the optical flow function.
[0167] Declaration 25: Method according to Declaration 14, wherein the video frame sequence includes a face, the plurality of objects includes respective regions of the face and the respective effects include respective makeup effects.
[0168] Declaration 26: Computer device comprising a processor and a non-transient storage device storing computer-executable instructions for execution by the processor to cause the computer device to perform the method according to any of the preceding method declarations.
[0169] Declaration 27: Computing device comprising a processor and a non-transient storage device storing computer-executable instructions for execution by the processor to cause the computing device to: provide a video chat, conferencing or teleconsultation application for broadcasting video frames to at least one other computing device; wherein the video chat or conferencing application is integrated with a virtual try-on (VTO) pipeline for applying one or more effects to the frames to be broadcast, the VTO pipeline having an object detection function and an effect rendering function configured to run in parallel per frame to optimize the application of one or more effects to the frames.
[0170] Declaration 28: Computer device according to Declaration 27, wherein computer-executable instructions for execution by the processor cause the computer device to provide recommendations of effects associated with respective products including makeup, hair or nail products, and to provide access to an electronic commerce interface for the purchase of a product.
[0171] Declaration 29: A computing device comprising a processor and a non-transient storage device storing computer-executable instructions for execution by the processor to enable the computing device to: provide a video chat, videoconferencing, or teleconsultation application for broadcasting video frames to at least one other computing device, the application being integrated into a virtual try-on pipeline (VTO) for applying one or more effects to the frames to be broadcast; wherein the application is configured to present a graphical user interface (GUI) via a display screen, the GUI having: a preview mode in which to receive input for selecting one or more effects among a plurality of effect options, preview mode is configured to present one or more effects applied to a sequence of preview frames via the display screen and without streaming the preview images to at least one other computing device; and a live mode is configured to apply one or more effects and stream the frames to which the effects are applied in live mode.
[0172] Declaration 30: Computer device according to Declaration 29, in which subsets of the plurality of effects are grouped into respective appearances, and the GUI provides an appearance selection interface to select one of the appearances to select one or more of the effects grouped for the appearance.
[0173] Declaration 31: Computer device according to Declaration 30, wherein each of the one or more effects is associated with a product and, in response to the selection of one of the appearances, the GUI is configured to present a selected appearance interface having respective sub-interfaces for each respective effect of the one or more effects of the appearance, each respective sub-interface providing i) product information for the respective product associated with the respective effect and ii) a refinement control to refine a selection of the respective effect.
[0174] Declaration 32: Computer device according to Declaration 31, wherein the refinement control is configured to select a different product for the respective effect so that the respective effect provides a virtual try-on of the different product.
[0175] Declaration 33: Computer device according to Declaration 31, wherein the respective refinement command is configured to disable the respective effect so that the respective effect is not applied to the frames.
[0176] Declaration 34: Computer device according to Declaration 30, wherein each of the one or more effects is associated with a product and, in response to the selection of one of the appearances, the GUI is configured to present a selected appearance interface having respective sub-interfaces for each respective effect of the one or more effects of the appearance, each respective sub-interface providing i) product information for the respective product associated with the respective effect and ii) an immediate purchase command to initiate an electronic commerce interface for the purchase of products.
[0177] Declaration 35: Computer device according to Declaration 30, wherein each of the one or more effects is associated with a product and the GUI is configured to provide access to initiate an electronic commerce interface for the purchase of products.
[0178] Declaration 36: Computer device according to Declaration 30, wherein the plurality of effects are any one of the makeup effects, hair effects or nail effects.
[0179] Declaration 37: Computer device according to Declaration 30, in which the appearance interface orders a presentation of respective appearances for selection in order to make recommendations to the user.
[0180] Declaration 38: Computer device according to Declaration 37, in which the order of respective appearances is determined in accordance with a processing of user information, the user information comprising one or more of: an existing makeup of a user, the shape of a user's face, the complexion of a user or any respective appearance as recently selected.
[0181] Declaration 39: Computer device according to Declaration 38, in which one or more of an existing user makeup, the user's face shape or the user's skin tone are determined by processing the frame sequence.
[0182] Declaration 40: Computer device according to Declaration 37, in which ordering is prioritized in response to the user's favorite selection of a respective appearance such that favorite appearances are ordered before other appearances and other appearances are ordered in accordance with the processing of user information.
[0183] Declaration 41: Computer device according to Declaration 29, in which preview mode can be invoked before joining a particular instance of a video chat or conference, integrated into a video chat or conference user interface to configure video and audio controls for the instance.
[0184] Declaration 42: Computer device according to Declaration 29, in which preview mode can be invoked during participation in an instance of a video chat or conference to stop all transmission of video frames for the instance.
[0185] A feature of any method statement has an equivalent device aspect such as a computer device, a computer system or a computer program product and vice versa.
[0186] A practical implementation may include all or part of the features described herein. These features, characteristics, and various combinations thereof, as well as others, may be expressed in the form of processes, apparatus, systems, means for performing functions, program products, and other ways, combining the features described herein. A number of embodiments have been described. Nevertheless, it is understood that various modifications may be made without departing from the spirit and scope of the processes and techniques described herein. Furthermore, other steps may be proposed, or steps may be eliminated, from the described process, and other components may be added to the systems described or to be removed from it. Consequently, other modes of implementation fall within the scope of the following claims.
[0187] Throughout the description and claims of this document, the terms "include" and "contain" and their variations mean "including but not limited to" and are not intended to exclude (and do not exclude) other components, integers or steps.
[0188] The features, integers, characteristics, compounds, chemical fractions, or groups described in conjunction with a particular aspect, embodiment, or example of the invention shall be understood as applicable to any other aspect, embodiment, or example, unless inconsistent with them. All features disclosed herein, and / or all steps of any method or process so disclosed, may be combined in any combination, except combinations in which at least some of these features and / or steps are mutually exclusive. The invention is not limited to the details of the preceding examples or embodiments. The invention extends to any new feature, or any new combination, of the features disclosed in this patent memorandum or to any new step, or any new combination, of the steps of any method or process disclosed.
Claims
Demands
1. A computer device comprising a processor and a non-transient storage device storing computer-executable instructions for execution by the processor to enable the computer device to: a. provide a video chat, video conferencing or teleconsultation application to broadcast video frames to at least one other computer device, the application being integrated into a virtual try-on pipeline (VTO) to apply one or more effects to the frames to be broadcast; b.in which the application is configured to present a graphical user interface (GUI) via a display screen, the GUI having: - a preview mode in which to receive input to select one or more effects from a plurality of effect options, the preview mode being configured to present the one or more effects applied to a sequence of preview frames via the display screen and without broadcasting the preview frames to at least one other computing device; and - a live mode configured to apply the one or more effects and to broadcast the frames to which the effects are applied in live mode.
2. A computer device according to claim 1, wherein subsets of the plurality of effects are grouped into respective appearances, and the graphical user interface provides an appearance selection interface for selecting one of the appearances to select one or more of the grouped effects for the appearance.
3. A computer device according to claim 2, wherein each of the one or more effects is associated with a product and, in response to the selection of one of the appearances, the GUI is configured to present a selected appearance interface having respective sub-interfaces for each respective effect of the one or more effects of the appearance, each respective sub-interface providing i) product information for the respective product associated with the respective effect and ii) a refinement control to refine a selection of the respective effect.
4. Computer device according to claim 3, wherein the respective fine-tuning control is configured to disable the respective effect so that the respective effect is not applied to the frames.
5. A computer device according to claim 2, wherein each of the one or more effects is associated with a product and, in response to the selection of one of the appearances, the GUI is configured to present a selected appearance interface having respective sub-interfaces for each respective effect of the one or more effects of the appearance, each respective sub-interface providing i) product information for the respective product associated with the respective effect and ii) an immediate purchase command to initiate an e-commerce interface for the purchase of products.
6. Computer device according to claim 2, wherein each of the one or more effects is associated with a product and the GUI is configured to provide access to initiate an electronic commerce interface for the purchase of products.
7. Computer device according to claim 2, wherein the plurality of effects is any one of makeup effects, hair effects or nail effects.
8. A computer device according to claim 2, wherein the appearance interface orders a presentation of the respective appearances to be selected in order to make recommendations to the user; and wherein the order of the respective appearances is determined in accordance with a processing of user information, the user information comprising one or more of the following: an existing makeup of a user, the shape of a user's face, the skin tone of a user or any respective appearance as recently selected.
9. A computer device according to claim 1, wherein the preview mode can be invoked before joining a particular instance of a video chat or conference, integrated into a video chat or conference user interface to configure video and audio controls for the instance.
10. Computer device according to claim 1, wherein the preview mode can be invoked during participation in an instance of a video chat or video conference to stop any broadcasting of video frames for the instance.