Virtual makeup trying method and device
By using the sclera as a lighting adaptation reference and pre-generated texture maps, combined with deep neural networks and physically based rendering technology, the challenge of real-time rendering of makeup effects on mobile devices, especially glitter effects, is solved, achieving an efficient and realistic VTO experience.
Patent Information
- Application Number
- CN202380093317.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-11
- Filing Date
- 2023-12-22
- Publication Date
- 2025-09-19
AI Technical Summary
Under uncontrolled lighting conditions, existing technologies have difficulty accurately rendering complex makeup effects, especially glitter effects, in real time on mobile devices, and methods that rely on external references or machine learning models are not suitable for real-time VTO experience.
The sclera is used as a reference for lighting adaptation, combined with deep neural networks to locate facial features, and through pre-generated texture maps and environment maps, combined with physically based rendering technology, makeup effects, especially glitter effects, are rendered in real time, reducing the computational burden.
This enables real-time and accurate rendering of makeup effects, especially glitter effects, on mobile devices, improving the realism and efficiency of the VTO experience and reducing dependence on external references and machine learning models.
Smart Images

Figure CN120676887A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Application No. 63 / 435,603, filed December 28, 2022, the entire contents of which are incorporated herein by reference. This application also claims priority to French Application No. FR2303587, filed April 11, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present disclosure relates to image processing, such as using neural networks, and to generating output images for virtually trying out products or services by applying simulated effects to one or more detected objects in an input image. Background Art
[0004] Deep learning techniques help process images comprising a series of video frames to locate one or more objects in the image. In an example, the objects are facial features comprising portions of a user's face. Image processing techniques also help render effects associated with such objects, such as augmenting reality for the user. One example of such augmented reality is providing a virtual try-on (VTO) that simulates applying a product (or service) to an object. Product simulations in the beauty industry include simulated makeup, hair, and nail effects. Other examples may include iris location and simulating color changes, such as through colored contact lenses. These objects and simulations, as well as others, will be visible.
[0005] In many VTO scenarios, users generate input images using their own camera-equipped computing devices, such as smartphones, tablets, or other computing devices with cameras (e.g., webcams), and do so under varying and uncontrolled lighting conditions.
[0006] Increasingly complex simulation effects are expected, for example to better simulate real product effects or to otherwise augment reality. For example, real product effects may include glitter effects or special lighting effects. Augmented reality simulations may include shaping effects that simulate changes in the shape of facial features (e.g., distortion) or the application of beauty filters. The shape changes may be the result of simulations by professional services (e.g., a beautician or plastic surgeon) or as a result of personal activities (such as self-care or for fun). Examples may include eyebrow shaping, nose shaping (narrowing of the nostrils, narrowing of the nose bridge), facial contour changes, eye or eyelid changes (e.g., vertical or horizontal eye enlargement), etc.
[0007] Improved techniques are expected to process images to facilitate providing augmented reality including VTO experiences. Summary of the Invention
[0008] Methods, systems, and techniques for rendering effects, such as makeup effects, are provided (e.g., in embodiments). In one embodiment, facial reshaping of facial features located by a facial tracking engine is performed by mapping and warping. Rendering renders the facial features as warped. In one embodiment, an input image is processed using a facial tracking engine with one or more deep neural networks to locate facial features of a face; and an output image is rendered using a rendering pipeline, the output image including makeup effects at locations associated with at least some of the facial features to simulate a virtual try-on of a makeup product, the effect having a 2D shape derived from a predefined 2D mask image adjusted using the 3D shape of the locations.
[0009] In one embodiment, a computer-implemented method is provided, comprising performing, by one or more processors, the steps of: processing an input image using a face tracking engine having one or more deep neural networks to locate facial features to respectively generate facial points defining a contour for each of the located facial features; and rendering, using a rendering pipeline, an output image derived from the input image, the output image being derived by applying one or more shape changes to a particular facial feature, the one or more shape changes being determined by: i) mapping a grid of spaced-apart grid points to pixels of the particular facial feature and any associated facial features; and ii) warping at least some of the spaced-apart grid points using a corresponding shape change function, the warping changing positions of at least some of the spaced-apart grid points for changing positions of facial points for the particular facial feature; and wherein the rendering determines output pixels for the particular facial feature and any associated facial features for the output image in response to the warping.
[0010] In one embodiment, a system is provided that includes at least one processor and a memory storing instructions executable by the at least one processor to cause the system to: process an input image using a face tracking engine having one or more deep neural networks to locate facial features to respectively generate facial points defining a contour for each of the located facial features; and render, using a rendering pipeline, an output image derived from the input image, the output image being derived by applying one or more shape changes to a particular facial feature, the one or more shape changes being determined by: i) mapping a grid of spaced-apart grid points to pixels of the particular facial feature and any associated facial features; and ii) warping at least some of the spaced-apart grid points using a corresponding shape change function, the warping changing positions of at least some of the spaced-apart grid points for changing positions of facial points for the particular facial feature; and wherein the rendering determines output pixels for the particular facial feature and any associated facial features for the output image in response to the warping.
[0011] In one embodiment, a computer-implemented method is provided, comprising the steps of: processing an input image using a face tracking engine having one or more deep neural networks to locate facial features of a face; and rendering an output image derived from the input image using a rendering pipeline, the output image including a makeup effect at locations associated with at least some of the facial features to simulate a virtual try-on of a makeup product, the effect having a 2D shape obtained from a predefined 2D mask image adjusted using a 3D shape of the locations.
[0012] These and other method and system aspects and other types of aspects will be apparent. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a diagram showing an input image of a typical scene for a VTO experience.
[0014] Figure 2 is an illustration of an eye providing the sclera as a reference according to an embodiment herein.
[0015] Figure 3 is an illustration of a computing environment such as for practicing one or more methodological aspects according to an embodiment.
[0016] Figure 4 is a screenshot of a user interface (UI) for parameter setting according to one embodiment.
[0017] Figure 5 According to one embodiment, Figure 3 A flowchart of the operation of a computing device.
[0018] Figure 6 An example of a facial image (eg, a cropped image of a face) annotated with multiple sets of facial points is shown.
[0019] Figure 7A and Figure 7B is an illustration of a facial distortion configuration 700 for distorting eyebrows, according to one embodiment.
[0020] Figure 8A and Figure 8B According to one embodiment, Figure 3 A flowchart of the operation of a computing device.
[0021] Figure 9A 、 Figure 9B and Figure 9C is an illustration of 3D polygonal meshes including an eye mesh and a lip mesh for facial features related to eyes and lips, according to an embodiment.
[0022] Figure 10 is an illustration of an eye region 1000 with an eye mask 1002 according to one embodiment.
[0023] Figure 11 According to one embodiment, Figure 3 A flowchart of the operation of a computing device. DETAILED DESCRIPTION
[0024] Lighting Adaptation
[0025] In one embodiment, a VTO application executed by a computing device includes a tracking engine and a rendering engine. The input image provided to the virtual try-on application is considered a scene. Typically, the scene consists of a face, surroundings (such as a background), and some other objects. The scene also includes one or more light sources that will modify the color and brightness of the scene. How these light sources modify the scene depends on their nature: natural (such as direct light from the sun) or artificial (such as light bulbs or neon signs). Such modifications may be a shift in the tint (hue) specifically for artificial light, or a difference in brightness / darkness, and occasionally some artifacts such as shadows. In one embodiment, the VTO application can improve the realism and accuracy of the makeup applied to the user's face by taking these light sources into account.
[0026] Correcting lighting for the purpose of applying makeup or skin effects to an image is described in US Pat. No. 10,892,166 B2, entitled “System and Method for Light Field Correction of Colored Surfaces in an Image,” issued on January 12, 2021, the entire contents of which are incorporated herein by reference (hereinafter referred to as the “'166 patent”). In one embodiment, the VTO application utilizes color correction as described in the '166 patent and further adapted as described herein.
[0027] Being able to infer the shift in hue and brightness of an image without any reference to the lighting source used or without an external reference is very difficult. Humans tend to do this naturally when the image is not within normal conditions, and it seems to be an easy problem to solve. However, for computers, the task seems to be very complex. In one example, Figure 1 A diagram of an input image 100 is shown in which cool natural light 102 may originate from a source on the upper left and directed diagonally toward a central face 104. In the same image, artificial warm light 106 may originate from a source on the lower right and directed diagonally upward toward the face 104. As a result, there is a hue shift and brightness shift in the skin tones, etc., of the face, such as represented by the diagonal dashed line 108. While this shift is represented by a straight line (108), this shift need not be straight nor steep.
[0028] Recoloring algorithms may have difficulty detecting hue shifts and rendering realistic colors—often lacking sufficient data to infer overall hues, while using external references (such as backgrounds) may lead to errors. For example, the background (e.g., area 110 outside the area of the human subject) may be noticeably shifted toward cool colors, while the face 104 may typically be warmly lit.
[0029] While difficult, inferring hue shift is not impossible, and a variety of techniques exist. In one previous example, the technique requires the use of a machine learning model that is trained using a large database of reference images where the lighting conditions are known. The image is then compared to this reference set, allowing the device to approximate the light in the scene and use the output to adapt the makeup rendering. This approach has similarities to how the human brain works, by comparing the scene light to a memory set and adapting the perception. However, these models are expensive to run and are generally not suitable for real-time experience. In one embodiment, the goal for VTO applications is to execute at more than 20 FPS (frames per second) on a mobile device such as a smartphone or tablet, where the VTO application is based on a web browser. This goal is generally incompatible with this technique of using machine learning methods for colorization.
[0030] Other techniques for adapting brightness and hue shifts require the use of an external reference, such as a color grid where the colors are calibrated in advance and known. This is called a color swatch, but it is impractical for many VTO use cases. While such techniques work well and are computationally fast, they require the color grid to be available throughout the entire experience (in case lighting changes), and more importantly, they require such a color swatch to be distributed to each user. Since the VTO experience is run by any user, this type of color grid is not a valid option.
[0031] Instead of relying on color charts and machine learning models, in one embodiment according to the novel teachings herein, color manipulation uses a reference available to any user, namely the sclera. Figure 2 As shown, the sclera 202 is the white portion of the user's eye outside of the iris 204. The sclera is generally consistent across different ethnicities, with a slight yellowish tint for individuals with darker skin tones, such as those of Black African descent.
[0032] The '166 patent describes a lighting adaptation operation that involves sampling different parts of the face to find an average skin color. In one embodiment, the lighting adaptation operation also samples for an average skin color, but further utilizes sclera evaluation, with the goal of improving the results while maintaining a reasonable "cost" associated with computing device resource usage and processing time.
[0033] The input image may include the face of an individual subject with eyes obscured by, for example, hair or sunglasses. In one embodiment, to mitigate this issue, a preliminary step performed by the eye coverage detector includes, for example, detecting whether the eyes are covered. In this case, sclera evaluation may be bypassed, so that a full illumination adaptation operation is not performed. In one embodiment, illumination adaptation, such as that described in the '169 patent, may be performed in response to sampling skin color.
[0034] VTO Application
[0035] Figure 3 3 is a diagram of a computing environment 300, such as for practicing one or more method aspects, according to one embodiment. Computing environment 300 shows a user computing device 302 (such as a smartphone), a communication network 304, a server 306, and a server 308. Communication network 304 includes wired and / or wireless networks, which may be public or private and may include, for example, the Internet. Server 306 includes a server computing device, such as for providing a website. Server 308 includes a server computing device, such as for providing e-commerce transaction services. Although shown separately, servers 306 and 308 may comprise a single server device. The computing environment is simplified. For example, not shown is a payment transaction gateway and other components, such as for completing e-commerce transactions.
[0036] Computing device 302 includes a storage device 310 (e.g., a non-transitory device such as a memory and / or a solid-state drive) for storing instructions that, when executed by a processor (not shown) (e.g., a central processing unit (CPU), a graphics processing unit (GPU), or both), cause computing device 302 to perform operations such as a computer-implemented method. Storage device 310 stores a virtual try-on application 312, which includes components such as software modules that provide: a user interface 314; a face tracker 315 having a deep neural network (DNN) (e.g., a convolutional neural network (CNN)) face detector 315A that provides facial points 315B; a VTO rendering pipeline component 316; a product recommendation component 318 having product data 318A; and a purchase component 322 having a shopping cart 324 (e.g., purchase data). In one embodiment not shown, the VTO application does not have a recommendation component and / or a purchase component, for example, providing a product selection component for selecting products to be visualized as a VTO effect.
[0037] In one embodiment, the VTO application is a web-based application, such as one obtained from server 306. In one embodiment, the VTO application, such as for native applications, is provided by a content delivery network. Product data 318A may be obtained from a content management system 307 associated with server 307A. Content management system 307 includes a data storage device 307B that stores product-related data, such as color palette data and rendering effect data. Color palette data typically includes an image of a product, which can be used to illustrate the product. In one embodiment, the color palette data (e.g., image) may be processed to extract data therefrom. The color palette data in the form of extracted data may include, for example, the color or other light-related properties (e.g., brightness) of the color palette data. Product data (e.g., color, etc.) may be provided to the server in other ways, such as in the form of colors or other parameters. For example, a slider may provide an input that maps to a data value. In one embodiment, the product data may be provided to a user device (e.g., 302) by being included in a native application bundle, or may be provided by server 306, such as in the form of an update to the native application. Thus, product data may be provided from various sources. The rendering effect data may include data for rendering an effect, such as to simulate makeup or other attributes, such as a matte look, a glossy look, a metallic look, a vinyl look, or a glitter look as further described herein for a makeup effect. Figure 4 ) is provided to a provider of a product (such as a brand owner) to upload a color palette image and provide input such as for defining color palette data, product data and / or rendering effect data. In one embodiment, the UI 307B is web-based.
[0038] Although not shown, the user device 302 may store a web browser for executing a web-based VTO application 312. In one embodiment (not shown), the VTO application 312 is a native application according to an operating system (also not shown) and software development requirements that may be imposed, for example, by the hardware manufacturer of the computing device 302. As is known, the native application may be configured for web-based communication or the like with the servers 306, 307A, and / or 308.
[0039] For example, Figure 3Various input data and output data or information associated with the use of the VTO application 312 are shown. Such input data and output data include: an input image 326 of a user to be processed for a VTO experience; an output image 328 on which a product effect is simulated to provide the VTO experience; a VTO product selection 320 including user input selecting one or more product effects to be simulated; a VTO product selection 322 including options for products to be virtually tried (e.g., for selection by a user of the device 302); and purchase transaction information 324 including purchase information provided to and / or received from the user to purchase the product. As mentioned, not all VTO application implementations include e-commerce capabilities.
[0040] In one embodiment, VTO product options 322 are presented via one or more user interfaces 314 for selection and virtual try-on by simulating effects on an input image 326. In one embodiment, VTO product options 322 are derived from or associated with product data 318A. In one embodiment, the product data may be obtained from server 306 and provided by product recommendation component 318, which, in one embodiment, may be a product data parser that does not itself make user-based recommendations. Instead, it makes any available product data available for selection. Although not shown, user input or other input may be received to determine product recommendations. The user may be prompted to provide input for determining product recommendations, such as via one of interfaces 314. In one embodiment, product recommendation component 318 communicates with server 306. In one embodiment, server 306 determines recommendations based on the input received via component 318 and provides product data accordingly. User interface 314 may present VTO product options, for example, updating the display of VTO product options in response to data received as the user navigates or otherwise interacts with the user interface.
[0041] In one embodiment, one or more user interfaces provide instructions and controls to obtain an input image 326 and a VTO product selection input 320, such as the identification of one or more recommended VTO products to try. In one embodiment, the input image 326 is an image of the user's face, which can be a still image or a frame from a video. In one embodiment, the input image 326 can be received from a camera (not shown) of the device 302 or from a stored image (not shown). The input image 326 is provided to the face tracker 315, such as for processing using a deep neural network 315A to locate features (e.g., objects) in the facial image.
[0042] In one embodiment, the position output from the face tracker 315 may include classification results, segmentation masks, or other position data for one or more detected objects, and the position output is provided to the VTO rendering pipeline component 316. In one embodiment, the positioning output includes facial points 315B from the face tracker 315. An input image 326 is also provided to (e.g., made available to) the pipeline component 316. A VTO product selection 320 is also provided to the pipeline component 316 for use in determining which effects to render. In one embodiment related to makeup simulation, one or more effects may be indicated, such as for any one or more product categories including lips, eyeshadow, eyeliner, blush, etc.
[0043] The VTO rendering pipeline component 316 renders effects on the input image 326, such as by drawing (rendering) the effects in layers (one layer per product effect), to generate an output image 328. The rendering is based on the product data 318A selected by the VTO product selection 320 and is responsive to the locations of the detected objects. For example, a VTO product selection of a lipstick, lip gloss, or other lip-related product calls for applying the effect to one or more detected mouth or lip-related objects at the corresponding locations. Similarly, a brow-related product selection calls for applying the selected product effect to the detected brow objects. Typically, for a symmetrical look, the same brow effect is applied to each brow, the same lip effect to each lip, or the same eye effect to each eye area, but this is not necessarily the case. Some VTO product selections include selecting more than one product, such as coordinated eyebrow and eye products or other combinations of detected objects, which can be labeled as "product looks." The VTO rendering pipeline component 316 can render each effect, for example, one product effect per layer at a time, until all effects have been applied. The order of application can be defined by rules or by the choice of products, e.g. lipstick before a top coat of lip gloss).
[0044] The user interface 314 provides an output image 328. In one embodiment, the output image 328 is presented as part of a real-time stream of continuous output images (each of which is an example 328), such as where a selfie video is enhanced to present an augmented reality experience. In one embodiment, the output image 328 can be presented together with the input image 326, such as displayed side by side for comparison. In one embodiment, it can be a before-and-after "in situ" comparison interface, where the user moves a slider to reveal more of the original image or the processed image. In one embodiment, the output image 328 can be saved (not shown) to a device such as the storage device 310 and / or the input image can be shared with another computing device (not shown).
[0045] In one embodiment, (not shown) the input image comprises an input image of a video conference session, and the output image comprises a video shared with another participant (or more than one participant) of the video conference session. In one embodiment, the VTO application (which may have another name) is a component or plug-in of the video conference application (not shown) that allows a user of device 302 to apply makeup during a video conference with one or more other conference participants.
[0046] exist Figure 3 In the illustrated embodiment, the VTO rendering pipeline component 316 is configured to apply lighting adaptation to the effect to be rendered. The components of the VTO rendering pipeline component 316 are shown as a process flow, which, for example, illustrates operations (e.g., steps of a method) performed by one or more processors (e.g., a CPU, a GPU, or both) of the computing device 302. It is understood that corresponding modules or software components, for example, can implement such a process. It is understood that, for the sake of brevity, other components for other adaptations that can be applied prior to rendering are not shown.
[0047] In this embodiment and for at least some purposes, color is modeled using a hue, saturation, and value / brightness (HSV or HSB) model. At 316A, an operation performs a skin average color check to determine an average skin tone color. In one embodiment, the operation is performed in accordance with the '166 patent. For example, the operation evaluates left and right skin (cheeks), facial skin (cheeks and forehead), left and right eye color, left and right sclera, minimum / maximum eye brightness, and lips. At 316B, the operation performs an eye coverage check to determine if the eye is covered such that the sclera is unavailable.
[0048] At 316C, the operation adjusts the brightness and saturation of the product data for applying the effect to the input image in response to the product selection 320. The brightness and saturation are adjusted in response to the average skin color and minimum / maximum eye brightness as detected at 316A. If more than one product effect is to be applied, step 316C is performed for each effect.
[0049] At 316D, the hue of the effect to be applied is adjusted based on the white of the eye detected from the sclera in step 316A (if available, where availability was detected in step 316B). In one embodiment, the operation determines whether the current color is biased towards warm or cool tones. The operation interpolates between different H, S, and V values based on the current color of the sclera. Thus, the hue of the effect is pushed towards warmer or cooler tones based on the detected condition. To determine whether it is a warm or cool toned hue, the HSV value of the sclera is examined. An example of a more detailed operation is described further below.
[0050] At 316E, the adapted product effect is rendered in a layer associated with the input image 326 to define an output image 328. In one embodiment, such as when more than one effect is to be applied, steps 316D and 316E are repeated for each effect to be applied. Once the steps of the pipeline component 316 are completed, the output image is provided, such as via the user interface 314.
[0051] In one embodiment, steps 316A through 316C are performed by a central processing unit or CPU (not shown), while steps 316D and 316E are performed by a graphics processing unit or GPU (not shown), as indicated by the dashed boxes.
[0052] Below is an example of more detailed operation of interpolation in response to the warmness or coolness of the sclera. Using the following input: c skin Refers to the HSV color of the current skin pixel, where the value of each HSV channel ranges from 0 to 1; c sclera refers to the average HSV color of the sclera of the eye (left or right eye, depending on which half of the face the current pixel is located), where each HSV channel ranges from 0 to 1; and c makeup Refers to the HSV color for the target makeup color, where each HSV channel ranges from 0 to 1.
[0053] Throughout the paper, the smoothstep(e0, e1, x) function is used, which is defined as a function that smoothly interpolates between 0 and 1 starting with x=e0 and ending with x=e1. At the time of filing, an exemplary implementation could be found at en.wikipedia.org / wiki / Smoothstep. RotationalMix(h0, h1, t) is also used, which is a function that linearly interpolates between hues h0 and h1 based on t. This function takes into account hue wraparound. For example, interpolating between 0.9 and 0.1 with t=0.5 will produce 0.0 as the output.
[0054] The following calculation is performed to determine whether the skin tone is warm:
[0055] condition skin_h_1 =smoothstep(0.03,0.2,c skin,h )
[0056] condition skin_h_2 =1.0-smoothstep(0.3,0.35,c skin,h )
[0057] condition skin_s=smoothstep(0.3,0.4,c skin,s )
[0058] condition skin_v =smoothstep(0.3,1.5,c skin,v )
[0059] The following calculation is performed to determine if the sclera is warm:
[0060] condition eye_h_1 =smoothstep(0.06,0.7,c sclera,h )
[0061] condition eye_h_2 =1.0-smoothstep(0.4,0.45, csclera,h )
[0062] condition eye_s =smoothstep(0.1,0.15,c sclera,s )
[0063] condition eye_v =smoothstep(0.6,0.7,c sclera,v )
[0064] To determine whether a warming adjustment should be made, all conditions are multiplied together to form condition warm The adjustment is done using the following formula:
[0065] c makeup,h,new =rotationalMix(c makeup,h ,0.2,condition warm )
[0066] A similar process is also applied to check for cool skin tones. The following operation is performed to determine if the skin is cool:
[0067] condition skin_h_1 =smoothstep(0.5,0.8,c skin,h )
[0068] condition skin_h_2 =1.0-smoothstep(0.8,0.85,c skin,h )
[0069] condition skin_s =smoothstep(0.01,0.5,cskin,s )
[0070] condition skin_v_1 =smoothstep(0.6,0.7,c skin,v )
[0071] condition skin_v_2 =1.0-smoothstep(0.9,1.0,c skin,v )
[0072] The following calculation is performed to determine if the sclera is cool:
[0073] condition eye_h_1 =smoothstep(0.5,0.85,c sclera,h )
[0074] conditione ye_h_2 =1.0-smoothstep(0.85,0.9,c sclera,h )
[0075] condition eye_s =smoothstep(0.1,0.15,c sclera,s )
[0076] condition eye_v =smoothstep(0.6,0.7,c sclera,v )
[0077] To determine whether a cooling adjustment should be made, all conditions are multiplied together to form condition cold The adjustment is done using the following formula:
[0078] c makeup,h,new =rotationclMix(c makeup,h ,0.75,condition cold )
[0079] In terms of methods, the following hue-related embodiments are provided: Hue embodiment 1: A computer-implemented method comprising one or more processors performing the following steps: processing the sclera in a face of an input image to determine a reference hue for the face; adjusting the hue of a makeup effect in response to the reference hue to render the adjusted makeup effect to the input image; and presenting an output image defined from the input image and the adjusted makeup effect via a user interface as part of a virtual try-on experience.
[0080] Hue Implementation 2: In Hue Implementation 1, the hue of the makeup effect is pushed toward warmer or cooler tones in response to the reference hue.
[0081] Hue Implementation 3: In Hue Implementation 1 or 2, the processing adjusts the brightness and saturation of the makeup effect in response to the average skin color determined by processing the skin in the face, and in response to the eye brightness (e.g., minimum / maximum) determined by processing one or more eyes in the face.
[0082] Hue Embodiment 4: In any of Hue Embodiments 1 to 3, adjusting the hue is in response to detecting the presence of a sclera for adjusting the hue. For example, if the sclera is not present, the hue is not adjusted. Brightness and saturation may still be adjusted.
[0083] Hue Implementation 5: In any of Hue Implementations 1 to 4, the method includes: processing the image using a deep neural network to locate facial features in the face; and wherein, in response to the located facial features, processing the sclera and rendering the makeup effect.
[0084] Hue Implementation 6: In any of Hue Implementations 1 to 5, the makeup effect is associated with a makeup product, and the method includes: providing recommendations of multiple makeup products to be virtually tried via a user interface; and receiving selection input via the user interface to select a makeup product for the virtual trial experience.
[0085] Hue embodiment 7: In Hue embodiment 6, the selection input selects a plurality of products, and the method includes adjusting the hue of each associated makeup effect in response to the reference hue for rendering the plurality of adjusted makeup effects to the input image.
[0086] Hue Embodiment 8: In any of Hue Embodiments 1-7, the method includes providing a purchasing service via the user interface to conduct a purchase transaction (eg, via e-commerce) to purchase the one or more cosmetic products.
[0087] It will be understood that system aspects and computer program product aspects corresponding to each of Hue Embodiments 1-8 are disclosed.
[0088] For example, in terms of the system, the following hue-related embodiments are provided: Hue embodiment 9: A system comprising: a rendering pipeline, which is configured (for example, via circuitry) to: process the sclera in a face of an input image to determine a reference hue of the face; adjust the hue of a makeup effect in response to the reference hue to render the adjusted makeup effect to the input image; and provide an output image defined from the input image and the adjusted makeup effect for presentation via a user interface as part of a virtual try-on experience.
[0089] Makeup effect with glitter
[0090] Glitter is used in many different makeup products, such as eyeshadow, lipstick, but also blush or eyeliner. Glitter typically requires many input variables and calculations to generate realistic reflections, making it difficult to simulate correctly.
[0091] The core of any glitter effect lies in calculating a set of particles that should be spread across a surface, where they should emit light. Calculating the position and reflectivity (i.e., particle properties) for each particle represented in a texture map can require too much computational power to run in real-time on mobile devices.
[0092] According to one embodiment, texture maps (position and particle properties) are pre-generated before rendering occurs, such as to reduce the computational burden during real-time mode rendering. These texture maps can be reused during rendering of glitter products.
[0093] Based on the techniques and empirical values in this article, pre-generated textures provide improved glitter effects when used in real-time.
[0094] Particle Texture: The first step is to calculate the particle texture, which will be used to place the particles and store the properties of each particle (e.g., size, reflectivity, base color, orientation, etc.). In one embodiment, the texture is calculated using a Voronoi diagram (structure). An example of a Voronoi diagram can be found at en.wikipedia.org / wiki / Voronoi_diagram.
[0095] Voronoi textures have certain properties that make them useful for defining glitter particles for rendering. Voronoi textures are created by generating random 2D points and creating boundaries between the points, where the boundaries have equal distances to the nearest points. These Voronoi regions formed by the boundaries define the locations of the glitter particles. It has been determined that the manner in which random 2D points are generated may have perceptible effects. In one embodiment, the random 2D points are generated according to a Poisson disk sampling technique. Poisson disk sampling produces more evenly distributed samples in the image than uniform sampling. Discussion and examples can be found in the article "Visualizing Algorithms" by Bostock, M., available at bost.ocks.org / mike / algorithms on June 26, 2014, at the time of filing, and incorporated herein by reference.
[0096] In one embodiment, for example, relative to a product swatch image, there are two texture images (maps) for storing particle information. One stores the normals of the particles, while the other stores the centers of the particles. Normals are normalized vectors whose angles to the +z direction are randomly sampled within a hard-coded range determined empirically (e.g., using a Poisson disk). The center of the particle is a coordinate normalized by the dimensions of the image. The texture map defines the glitter particles using positions in 3D space, for example, using the particle center and a normal vector. In one embodiment, the normal vector is a 3D direction and includes an x-value, a y-value, and a z-value (where the x-value, the y-value are different from the (x, y) position relative to the pixel or other grid of the particle center).
[0097] These textures are sampled at specific sizes based on the glitter size parameters in the product, and they are configured to be resampled across boundaries. In one embodiment, the textures are sampled using mipmapping, where textures of different resolutions are pre-generated. Mipmapping is a graphics technique that scales and filters an original high-resolution texture image or map and scales it down to multiple smaller resolution texture maps. During rendering, the operation dynamically selects which texture to use based on various factors, such as the output resolution relative to the glitter texture size. The textures are scaled based on the glitter size.
[0098] Environment Mapping
[0099] A pre-generated environment map is used to define the lighting of the environment according to image-based lighting (IBL) techniques. It has been found that instead of using a real-life environment map, controlling how "sparkly" the glitter is can be achieved by generating an environment map that acts as if there is a spotlight coming from a certain direction and concentric rings of light around the center point. It is important to note that this environment map only needs to be grayscale since only the brightness of the environment is of interest, not the color. In one embodiment, the environment map is stored as a cube map (e.g., 6 textures that can be wrapped / folded to define a cube), where each cube face is stacked vertically into one image. If the light is viewed from a specific face, each cube face will appear to be a spotlight coming from a specific direction.
[0100] Glitter Rendering
[0101] To render the glitter effect, the operation finds the normal of the glitter particles, which can be transformed by the rotation of the face and the surface (if provided). In one embodiment, the face tracking engine can provide object positioning information, such as for the eyes, lips and / or facial contours, for example, facial points. These points (positioning information) can be used to determine the rotation of the face (e.g., relative to a standard). 3D or other models can provide additional surface information, such as the surface of the area where the glitter is to be applied. 3D shape information is further described herein.
[0102] In one embodiment, rendering is applied using physically based rendering (PBR) techniques and image based lighting (IBL) techniques. Physically based rendering material properties (e.g., roughness, F0, etc.) and viewing direction are hard-coded (e.g., pre-calculated and stored), while the brightness of the environment is found from the environment map (cubemap and its texels) using the normal vector. This data is used in the lighting operations of the PBR-based image to find the specular brightness of the particles, such as described in Learn OpenGL - Graphics Programming "Specular-IBL", de Vries, Joey, June 17, 2020, available at learnopengl.com / PBR / IBL / Specular-IBL at the time of filing and incorporated herein by reference. It is noted that in one embodiment, the diffuse term (one of the IBL elements) is ignored because the particles are expected to be primarily reflective.
[0103] As described herein, a makeup effect simulates applying a product to a portion of a face, such as applying lipstick to the lips, applying eyeshadow to the eyelid area around the eyes, and the like. The portion of the face with the effect is overlaid, for example, by a GPU, on an input image. Thus, the portion / effect includes pixel data to be determined for the overlay operation, such as using a texture map, an environment map, and rendering data. The pixels of the effect can then be considered to represent glitter particles and non-glitter particles. Whether a particular pixel (e.g., the current pixel) is a glitter particle pixel can be determined using various parameters, such as the distance to the center of a glitter particle (within a texture map) and the size of the particular glitter particle. The center of the particular glitter particle can serve as a seed for determining particle parameters, such as particle size (e.g., in a number generator that determines a size within a normalized range) or color (e.g., in a number generator that determines a value within a range of a color model).
[0104] In one embodiment, the glitter particles are circular, and the size distance from the center is used to draw the particle's circular shape and glow alpha (only the center portion is fully lit, and the light gradually fades towards the edges of the particle).
[0105] The results of these operations are combined to render the lighting of the particles.
[0106] One limitation of Voronoi textures is that they do not allow for overlap. In one embodiment, to make the glitter effect appear at a higher density and with overlapping particles, the rendering operation is repeated, where a texture map is used to define multiple overlapping texture map layers. The operation is repeated and the glitter rendering is overlaid with particle textures (i.e., texture maps) rotated 90 degrees, 180 degrees, and 270 degrees. In one embodiment, the operation varies the density by assigning a probability of presence for a particular glitter particle in the texture map to the layer, and whether the particle exists depends on a random value greater or less than the probability of presence. That is, if the layer has a probability of presence of 0.8 and a particle has a value of 0.6, it will exist; but if its value is 0.9, it will not exist (e.g., its glow alpha is set to 0).
[0107] Figure 4 4 is a screenshot of a UI 400 for, for example, defining glitter parameter settings for rendering effect data. UI 400 can be made available to brand owners or other entities to define characteristics for a makeup effect (such as for a VTO experience). The UI can be configured to receive product data, such as to define a base makeup effect and any special rendering effects, such as special lighting or glitter effects, makeup effect shapes (e.g., mask images), etc. In one embodiment, UI 400 is an example of UI 307B.
[0108] UI 400 facilitates adjusting a glitter effect (e.g., a glitter look or "glitter") such as provided by glitter / glitter particles of a makeup effect. Exemplary makeup effects are eye effects such as provided by eye shadow, lip effects such as provided by lipstick, cheek effects such as provided by loose powder, etc.
[0109] UI 400 includes an icon strip 402, such as for switching user interfaces for each of the different products, product effects, and / or associated appearances (other user interfaces are not shown). Icon 402A is associated with glitter ("glitter") and invokes UI 400. UI 400 includes a plurality of input controls (e.g., 404 to 416) for receiving input to define corresponding glitter appearance attributes / parameters. Controls 404 to 416 can take various forms, and several controls herein include slider controls. Other types of input controls for entering text or similar values are known and useful. Voice activation controls may be used. Controls 404 to 416 include:
[0110] Color 404: The base color of the glitter (to be modified by reflection, intensity, and color variation); calling color control 404 may present further color section controls (not shown), such as a color wheel and / or red, green, and blue (RGB) input value controls for defining colors using RGB additive color values.
[0111] Reflection 406: The reflectivity of the glitter;
[0112] Color change 408: 0 means no change, 100% means change within the entire color spectrum;
[0113] Density 410: How much flash we produce;
[0114] Intensity 412: the amount of visible flash;
[0115] Size 414: The base size of the flash (can be modified by size changes); and
[0116] Size variation 416: 0 means no variation, 100% means that all glitter particles have a variation in size (eg, within a normalized range).
[0117] Thus, in one embodiment, the rendering operation first renders a base product, for example, in the shape of an area of the face associated with one or more facial features. The rendering operation then renders one or more glitter layers of glitter particles. Preferably, multiple glitter layers of glitter particles are applied on top (e.g., four layers total). Multiple layers are preferably used to allow for particle overlap. For the random size of the glitter particles within the glitter particle layers, the rendering operation generates random sizes between hard-coded minimum and maximum values (which are affected by the glitter size) and linearly interpolates between the base glitter size and the random size based on a parameter called "glitter size variation." If this parameter is zero, there is no randomness, while if it is one, the size is completely random. For the random color of the glitter particles, the operation generates random colors within a certain hue, saturation, and lightness (HSL) range (e.g., in one embodiment, limited to colors with high saturation and high lightness). Similar to the glitter size, the operation linearly interpolates between the base glitter color and the random color based on a parameter called "glitter color variation." If this parameter is zero, there is no randomness, while if it is one, the colors are completely random. The pixel values of the glitter makeup effect are assigned in response to the distance from the center of the glitter particle. If the distance from the current pixel to the center is within the circle size, the current pixel is considered a glitter pixel, otherwise no change is made to the pixel (because the base product has already been rendered). Since each pixel of the pre-generated texture map contains the center of the nearest particle, only one texture lookup is required for the current pixel to determine the nearest particle.
[0118] Figure 5 According to one embodiment, Figure 3 Flowchart of operation 500 of a computing device (e.g., device 302) of FIG. Operation 500 defines a computer-implemented method, which includes steps of the method performed by one or more processors. For example, in terms of the method, the following glitter-related embodiments are provided:
[0119] Glitter Embodiment 1: A computer-implemented method comprising the following steps, performed by one or more processors: processing an input image of a face to determine the location of facial features, the input image being processed by a face tracking engine comprising at least one deep neural network to locate the facial features (step 502); rendering a makeup effect associated with the facial features, the makeup effect comprising a glitter effect defined from a pre-computed texture map and a light environment map for positioning and lighting the glitter effect (step 504); and providing an output image defined from the input image and the makeup effect for presentation via a user interface as part of a virtual try-on experience (step 506). For example, the face tracking engine may include a face tracker 315 as described above. For example, rendering (and providing the output image) may be provided by a rendering pipeline 316 adapted for use with the glitter effect teachings herein.
[0120] Glitter embodiment 2: In the glitter embodiment 1, the makeup effect includes a lip makeup effect, an eye area makeup effect, or a cheek area makeup effect.
[0121] Glitter Implementation 3: In Glitter Implementation 1 or Glitter Implementation 2, the rendering step is to define the glitter effect in response to a physically based rendering technique.
[0122] Glitter Embodiment 4: In any of Glitter Embodiments 1 to 3: for each glitter particle in a plurality of glitter particles, a pre-computed texture map defines a glitter position and a glitter reflection angle; a light environment map simulates a light source (e.g., in three-dimensional space) for determining a specular brightness of a particular glitter particle based on the reflection angle of the particular glitter particle; and rendering a pixel value defining a pixel for a makeup effect, thereby adjusting the illumination of the current pixel based on a distance from the current pixel to the glitter position of the particular particle, the specular brightness of the particular glitter particle, and the size of the particular glitter particle. Glitter Embodiment 5: In Glitter Embodiment 4, the specular sparkle is further determined in response to a three-dimensional shape of the makeup effect and facial rotation.
[0123] Glitter embodiment 6: In glitter embodiment 4 or 5, the sizes of the specific glitter particles are randomly distributed within a normalized range.
[0124] Glitter Embodiment 7: In any of Glitter Embodiments 4 to 6, the lighting is further responsive to a glow alpha of a particular glitter particle, defining the light to dim from the center.
[0125] Glitter Embodiment 8: In any of Glitter Embodiments 4 to 7, rendering glitter colors responsive to specific glitter particles further colors the makeup effect.
[0126] Glitter embodiment 9: In any one of glitter embodiments 4 to 8, multiple glitter particles are randomly spaced apart in a texture map without overlapping, and wherein rendering uses the texture map in multiple overlapping texture map layers to repeat the definition of pixel values to overlap the glitter particles in the glitter effect.
[0127] Glitter Implementation 10: In Glitter Implementation 9, the glitter particle density of the glitter effect is rendered to vary, thereby randomly determining whether a particular particle in the texture map exists in any of the overlapping texture map layers.
[0128] Glitter Embodiment 11: In any one of Glitter Embodiments 4 to 10, the normal to the reflection angle (e.g., defined as a vector) is aligned with the light environment. Figure 1 Used together to determine spectral sparkle.
[0129] Glitter Embodiment 12: In any one of Glitter Embodiments 4 to 11, each of the glitter positions of the plurality of glitter particles is randomly assigned using Poisson disk sampling.
[0130] Glitter Embodiment 13: In any of Glitter Embodiments 1 to 12, the rendering is responsive to any one or more rendering effect parameters, including a color parameter, a reflectance parameter, a color variation parameter, a density parameter, an intensity parameter, a size parameter, and a size variation parameter.
[0131] It will be understood that system aspects and computer program product aspects corresponding to each of Glitter Embodiments 1 to 13 are disclosed.
[0132] Other glitter-related method aspects are disclosed, such as for a computing device configured to precompute maps for use during rendering. For example, glitter embodiment 14 is provided: a computer-implemented method comprising performing the following steps by one or more processors: precompute texture maps and light environment maps for locating and lighting a glitter effect as part of a makeup effect to be rendered, the makeup effect being associated with facial features located by a facial tracking engine of a virtual makeup try-on application; and providing the precomputed texture maps and light environment maps for use by a rendering pipeline of the virtual makeup try-on application for rendering the makeup effect with the glitter effect. Glitter embodiment 15: In glitter embodiment 14, the precomputed texture maps and light environment maps are defined for generating a glitter effect according to a physically based rendering technique that models particles in an illuminated environment. Glitter Embodiment 16: In Glitter Embodiment 14 or 15, a texture map is pre-computed based on a Voronoi diagram technique to provide random positions for a plurality of glitter particles for a glitter effect, and the texture map is further computed to include a reflection angle for each of the plurality of glitter particles, respectively. Glitter Embodiment 17: In Glitter Embodiment 16, a light environment map is defined to model a light source so that, when rendered, a specular sparkle is provided for each of the plurality of glitter particles in the glitter effect, respectively, the specular sparkle being responsive to the reflection angle. Glitter Embodiment 18: In Glitter Embodiment 16 or 17, a Poisson disk sampling distribution is used to distribute the positions of the glitter particles.
[0133] It will be understood that any of the adapted additional method-related embodiments applicable to any of Glitter Embodiments 1 to 13 may also be applicable to Glitter Embodiments 14 to 18. It will be understood that system aspects and computer program product aspects corresponding to or adapted as mentioned in each of Glitter Embodiments 14 to 18 are disclosed. Any of the Glitter embodiments may be combined with any one or more of the Hue embodiment, the Shaping embodiment, and the Grid embodiment, for example, to combine their method aspects, define corresponding system aspects, etc.
[0134] Sculpting facial features
[0135] Typically, when a facial feature is identified by a facial tracking engine (such as face tracker 315), the makeup VTO applies the product effect to that feature. The VTO typically uses only the original image and applies a "smear" color or texture on top to cover the original color of the detected facial feature, just like regular makeup does. However, for the eyebrow category, some beauty products sold include shaping tools and products that allow their users to change the overall or local shape of the eyebrows. It is desirable to simulate reshaping. In addition to detecting shapes and rendering the detected shapes, additional sets of transformations such as those described herein can also be used to reshape, for example to perform eyebrow warping.
[0136] In one embodiment, eyebrow warping is an image manipulation technique that transforms the pixels of the original image. Pixel transformation is preferred over replacing the original eyebrows with synthetic eyebrows. Removing and replacing them with an overlay can be difficult and provide less realistic results.
[0137] Some examples of eyebrow transformations include: global brow arch elevation, inner thickness reduction or increase, top cleanup, and bottom cleanup. The eyebrow shape can be characterized as having an inner portion closer to the nose, a middle portion, and an outer portion farthest from the nose. The inner, middle, and outer portions are positioned along or define the brow arch, which has a top line (closest to the forehead) and a bottom line (closest to the eyes).
[0138] In one embodiment, the values of the eyebrows that can be transformed (e.g., eyebrow parameters) include: outer part: horizontal and vertical alignment, thickness; inner part: horizontal and vertical alignment, thickness; brow arch: local and global increase / decrease, sharpness of the brow arch; middle part: thickness; cleanup: top, bottom, inner part; and global: horizontal and vertical offset.
[0139] In one embodiment, a facial tracker (such as tracker 315) can locate facial features and provide information such as Figure 6 The facial points depicted in . Figure 6 An example of a facial image 600 (e.g., a cropped image of a face 602) annotated with multiple sets 604 of facial points is shown. The depiction of the points on the face is for illustration purposes only. The output of the face tracker 315 need not include the annotated image, and the output of the facial points and the cropped face may be, for example, separate data, or the cropped face itself may not be provided.
[0140] In one embodiment, multiple groups 604 of facial points include a facial outline facial point 604A, corresponding right and left eyebrow facial points 604B and 604C, a left and right eye facial point 604E and 604D, a nose facial point 604F, and two lip facial points 604G and 604H. In one embodiment, the lip facial points include an outer lip group 604G (around the mouth) and an inner lip group 604H (between the lips). Facial points within a single group are numbered (e.g., 0, 1, 2, ...) and help define the outline of a detected object. In one embodiment, the facial tracker assigns each point so that each point is placed at a consistent location relative to the outline of the object it is representing. For example, a particular point may always be at the right corner of the mouth. In one example, a facial point is an X, Y pixel coordinate relative to the cropped facial image 602 and is associated with a corresponding detected object from a network (e.g., one of the networks) of trackers 315.
[0141] In one embodiment, operations such as for a facial tracker and rendering pipeline component for a virtual try-on application perform a computer-implemented method comprising one or more of the following steps:
[0142] 1. Process the input image using a face detector provided with one or more neural networks to locate facial features to obtain eye facial points and eyebrow facial points.
[0143] 2. Figure 7A and Figure 7B is an illustration of a facial distortion configuration 700 for distorting eyebrows (an example of a facial feature), according to one embodiment. Figure 7A As depicted in , a rectangular warping grid (704) is defined over each eyebrow (e.g., the left eyebrow 702) centered around each eyebrow region. The size of this box (grid) is calculated from the sizes and distances of the various facial features. The grid is also rotated (via an affine transformation) to match the rotation of the face. For example, using the facial contour facial points 604A, the rotation of the face relative to a standard position can be determined. A coordinate system is defined within the grid (e.g., as represented by the points in the grid box 704) where the coordinates range from (0,0) to (1,1). This simplifies the warping operations because they no longer need to take into account the rotation of the face or the size of the eyebrows. The eyebrow facial point set and the eye facial point set for the right or left eye are mapped into this coordinate system (mapped into the corresponding grid determined for each of these facial feature pairs) for use during the warping operation.
[0144] 3. Discretize the warped grid into a grid of 2D points (e.g., the points shown in box 704, including points 704A, 704B, 704C, 704D, and 704E, where 704C is located between point 704A and point 704D within the eyebrow 702). By moving a grid point (e.g., 704D), the facial image pixel at or near that grid point and its surrounding pixels (e.g., near grid point 704E) are also moved. For a description of the movement of the grid points and the resulting movement of the eyebrow pixels to reshape the eyebrow 702, see Figure 7B . A linear interpolation of the warp is performed between the points. In this embodiment, the GPU provides this linear interpolation as a built-in function. There is a trade-off, choosing a denser grid will result in smoother deformations, but will result in higher (GPU) processing time. For example, it has been found through experimentation that a 25 by 50 point grid (1250 points total) provides a good balance between quality and speed for GPUs provided by mobile devices (such as mobile phones) that are commonly available at the time of filing the application. Anything higher will result in diminishing returns on quality. If the hardware is more powerful, a denser grid can be used.
[0145] 4. Each point on the grid is then warped, where the warp is computed using the applicable facial points and warp parameters. That is, a predefined grid point warp (movement) may be defined in association with a specific eyebrow transformation (e.g., parameters).
[0146] 5. The facial image (e.g., its pixels) is then warped based on the warping grid. For example, eyebrow shaping shapes two eyebrows, each with a grid. Eyebrow shaping can be performed in conjunction with other beauty filter shaping, such as eye or eyelid shaping, lip shaping, nose bridge shaping, nostril shaping, facial contour shaping (e.g., chin narrowing), etc., where each feature being shaped has a corresponding grid.
[0147] Distortion operation
[0148] While the warping operations used for each eyebrow parameter are different, there are some shared techniques between them. Each parameter has an associated warp function:
[0149] p out =f param (p in ,p brows ,p eyes ,v param )
[0150] Where: p in are the input grid points,
[0151] p brows It is a collection of eyebrow points.
[0152] p eyes is an example of a set of eye points, associated facial features and helps define eye region facial features (e.g., associated facial features) along the upper curve of the eye between the eyebrow point and the eye point; initially, the set of eye points and eyebrow points is provided by the face tracker.
[0153] v param is the value of the parameter,
[0154] p out are the output grid points.
[0155] In one embodiment, the warping is performed by iteratively applying each warp function (e.g., the shape parameter (param) in each warp function represents the change to the eyebrows) on the warp grid. The overall process is:
[0156] For each parameter param:
[0157] By adding f param Applied to each grid point to distort the grid;
[0158] f param Applied to p brows and p eyes In order to obtain the new eyebrow point p' brows (and eye point p' eyes ) set, since the eyebrow points (and possibly the eye points) have changed due to the warping. Use the new points for the next iteration.
[0159] Although not shown, a user interface (e.g., similar to user interface 400) may be provided to receive input of eyebrow parameters. In one embodiment, a slider or other control may be provided to input how much to change the parameter.
[0160] Distortion Technology
[0161] There are some common techniques among the warping functions. This section lists these techniques:
[0162] Curve Matching: The curve of the current eyebrow is calculated (by fitting a curve along the middle of the eyebrow using the eyebrow points of the face tracker), and a target curve is calculated from these parameters. This helps to change the overall curvature of the eyebrow. The eyebrow is deformed (e.g. in a grid) so that the current curve is warped to match the target curve. This is done by matching points on the source curve to the target curve and warping based on the delta. For points in the grid that are far from the curve, a 2D Gaussian falloff function is applied so that points close to the curve are warped more strongly than points further away. The σx and σy of this Gaussian function can be changed to adjust the "area of influence". For example, σy needs to be large enough so that it affects the entire thickness of the eyebrow. In some cases, the Gaussian function can be replaced with a different function, such as an asymmetric version of the Gaussian function or a Gaussian function with a "shorter tail" or a "longer tail".
[0163] Point Matching: This is similar to Curve Matching, but instead of matching curves, you match discrete points. This is useful for things like adjusting the alignment of a specific part of an eyebrow.
[0164] Area Expansion / Compression: You can "compress" or "expand" the surrounding area along a curve or at a point (in a specific direction, such as vertical, or uniformly). This is achieved by making nearby points closer / further away and using a Gaussian function to decay. This is useful for operations such as changing the thickness of eyebrows or sharpening the edges of eyebrows.
[0165] Implementation: Render a warped grid using grid points as vertices, such as to define a polygonal mesh. Triangles are fitted to the vertices, and an image is UV mapped onto the vertices. UV mapping in this implementation is a 2D modeling process that projects the surface of a 2D model onto a 2D image for texture mapping. The letters "U" and "V" represent the axes of a 2D texture.
[0166] Beauty Filter: Similar to the eyebrow distortion operation, the beauty filter operation is a set of transformations that change the shape of other facial features. In one embodiment, the transformations include: eye enlargement (vertical and horizontal); jaw narrowing; nostril narrowing; and nose bridge narrowing. The implementation for the beauty filter is the same as for the eyebrow filter, but the affected area is different, depending on the situation.
[0167] Figure 8A and Figure 8B According to one embodiment, Figure 3 Flowchart of operations 800 and 804 of a computing device (e.g., device 302) of FIG. Operation 800 defines a computer-implemented method, including steps performed by one or more processors. In one embodiment, operation 804 is further defined as steps 804A to 804C.
[0168] For example, in terms of methods, the following shaping-related embodiments are provided: Shaping embodiment 1: A computer-implemented method comprising performing the following steps by one or more processors: processing an input image using a face tracking engine having one or more deep neural networks to locate facial features to respectively generate facial points defining the outline of each of the located facial features (step 802); and rendering an output image derived from the input image using a rendering pipeline, the output image being derived by applying one or more shape changes to a specific facial feature (step 804), the one or more shape changes being determined by: i) mapping a grid of spaced-apart grid points to pixels of the specific facial feature and any associated facial features (804A); and ii) warping at least some of the spaced-apart grid points using corresponding shape change functions, the warped change positions of at least some of the spaced-apart grid points being used to change positions of facial points of the specific facial feature (804B); and wherein the rendering determines output pixels for the specific facial feature and any associated facial features of the output image in response to the warping (804C).
[0169] For example, the facial tracking engine may include a facial tracker 315 as previously described. Rendering (and providing an output image) may be provided by a rendering pipeline 316 adapted for use with the glitter effect teachings herein, for example. Shaping Implementation 2: In Shaping Implementation 1, the facial tracking engine and rendering pipeline are components of a VTO application for simulating the effect of a makeup product applied to a facial feature. Shaping Implementation 3: In Shaping Implementation 2, the rendering pipeline renders the makeup effect to the specific facial feature that has been reshaped, such that the output image includes the specific facial feature that has been reshaped and has the makeup effect.
[0170] Shaping embodiment 4: In any of shaping embodiments 1 to 3, the method includes providing a user interface to receive input for defining shape parameters for one or more shape changes, and wherein the rendering is responsive to the user input.
[0171] Reshaping embodiment 5: In reshaping embodiment 4, at least some of the shape changing functions perform one or more of the following: curve matching to match an intermediate curve along the middle of the contour of a specific facial feature to a target curve defined by one or more shape parameters of the shape change, thereby attenuating the position changes of spaced grid points in response to the distance to the target curve; point matching to match discrete points of the contour of a specific facial feature to target points defined by one or more shape parameters of the shape change; and region expansion or region compression to expand or compress the region along the facial point curve or around a specific pixel in response to the shape parameters, the region expansion or region compression being attenuated by an attenuation function in response to the distance from the facial point curve or the specific pixel.
[0172] Sculpting embodiment 6: In any of sculpting embodiments 1 to 5, determining pixels includes fitting triangles to vertices defined by the warped spaced-apart grid points, and UV mapping pixels of the particular facial feature and any associated features onto the vertices.
[0173] Reshaping embodiment 7: In any of Reshaping embodiments 1 to 6, the particular facial feature defines a first facial feature, and the rendering step is repeated with respect to a second facial feature to produce an output image having at least two shape-changed facial features.
[0174] Contouring embodiment 8: In any of contouring embodiments 1 to 7, the shape change applied to a specific facial feature includes any of the following: eyebrow contouring; nose contouring, such as narrowing of the nostrils or narrowing of the nose bridge; facial contour changes, such as narrowing of the mandible; eye or eyelid changes, such as vertical or horizontal eye enlargement; or lip changes, such as lip fulling.
[0175] It will be understood that system aspects and computer program product aspects are disclosed, for example, corresponding to each of the shaping embodiments 1 to 7. Any of the shaping embodiments may be combined with any one or more of the hue embodiments, glitter embodiments, and grid embodiments.
[0176] 3D mesh based on 2D landmarks
[0177] In one embodiment, to facilitate more realistic rendering of facial features, whether applying sculpting effects or applying cosmetic effects (such as makeup looks), rendering operations (e.g., pipeline 316) can be configured to utilize 3D polygon meshes. In one embodiment, the 3D polygon meshes are used to render special lighting effects (e.g., metal, vinyl) to the cosmetic effects.
[0178] A mesh is a collection of vertices, edges, and faces that defines (e.g., models) the shape of a 3D object. In addition to enabling complex lighting effects, there are many advantages to using polygonal meshes: polygonal meshes can be rendered efficiently because many commercially available GPUs used in consumer smartphones, tablets, etc. at the time of filing are optimized to process polygons, resulting in faster rendering speeds.
[0179] In a VTO application, user input can indicate which product or products to apply as one or more makeup effects associated with at least some facial features located by a facial tracker. The effects can be applied to an image using a mask (e.g., a 2D mask image) to provide a shape for the makeup effect relative to at least some of the facial features. For example, to provide a realistic effect that responds to the 3D shape of the facial features, the mask for the effect can be warped into a 3D mesh model that realistically fits the facial features. The warped 2D mask image can be UV mapped for use in shaping the effect during rendering.
[0180] Generate 3D mesh
[0181] In one embodiment, the face tracker 315 detects 2D facial points. The rendering operation uses the 2D facial points to generate a 3D mesh. In one embodiment, the operations for generating the 3D mesh include:
[0182] 1. Estimate 3D facial points using the detected 2D facial points and the existing 3D facial model - project the 2D facial points onto 3D space to obtain 3D facial points. Typically, cameras and human eyes observe in perspective projection. When 3D points are projected into 2D space, a perspective projection matrix is applied. In one embodiment, in order to go from 2D points to 3D points, the inverse matrix of the projection matrix can be applied. However, the depth is unknown, so the result is actually a 3D ray, not a 3D point. To obtain 3D points, the operation approximates the face using a 3D plane, and intersects the 3D ray with the plane to obtain 3D points. A hard-coded offset is added to each point to account for the fact that the face is not actually flat. For example, for lip points, points near the corners of the mouth will be farther from the camera in depth than points in the center;
[0183] 2. Use 3D facial points to interpolate spline curves and use these smoothed curves as the outlines of facial features;
[0184] 3. Generate a 3D grid (e.g., of spaced grid points) and deform the grid points to fit the curvature of the 3D contour line; and
[0185] 4. Extract 3D surfels (with triangular shapes) from the deformed grid and group them together to form a 3D facial mesh.
[0186] Figure 9A 、 Figure 9B and Figure 9C is an illustration of a 3D polygonal mesh 900 including an eye mesh 902 and a lip mesh 904 for facial features related to eyes and lips, according to an embodiment. Figure 9A shows a front view, and Figure 9B and Figure 9C Partial left-turn view and partial right-turn view are shown respectively. Figure 9B and Figure 9C , Figure 9A It is apparent that the eye-related mesh in this embodiment includes the area around the corresponding eyes, including the associated eyebrows.
[0187] Mapping a 2D mask image to a 3D mesh
[0188] In one embodiment, known techniques for mapping can be used. For example, to "warp" a 2D mask image onto the surface of a 3D mesh, the operation utilizes a technique known as UV mapping. In this embodiment, the mask image references a shape for a makeup effect, such as the shape of eyeshadow applied to the eye area, which can be positioned adjacent to the top outline of the eye and extending toward the lower outline of the eyebrow and beyond the outside of the eye. Figure 10 1 is an illustration of an eye region 1000 with an eye mask 1002 according to one embodiment. The mapping process can be interpreted as a paper 3D model of an object (e.g., a sphere) being placed flat on a table, and each of the object's 3D coordinates being mapped to a 2D coordinate on a flat image. The purpose of this unfolding of 3D coordinates is to map these 3D coordinates to an image / picture so that the 3D image can have a realistic-looking surface, where textures are derived from these images.
[0189] Using an eyeshadow effect as an example, the general process for mapping a mask image onto a 3D mesh is as follows: use an eye template image (e.g., mask image 1002 for creating an eye makeup effect) as a texture image; unfold the 3D mesh into 2D; and adjust the 2D UV vertices so that the texture appears correctly on the 3D mesh. In one embodiment, the 3D mesh can be unfolded, such as by using available third-party software. The 2D vertices can be aligned with the mask image (e.g., giving the shape of the makeup effect) and the reference point.
[0190] Figure 11 According to one embodiment, Figure 31. Operation 1100 of a computing device (eg, device 302) is a flowchart of operation 1100 of a computing device (eg, device 302) of FIG. Operation 1100 defines a computer-implemented method including steps of the method performed by one or more processors.
[0191] For example, in terms of methods, the following mesh-related embodiments are provided: Mesh embodiment 1: A computer-implemented method comprising one or more processors performing the following steps: processing an input image using a face tracking engine having one or more deep neural networks to locate facial features of a face (step 1102); and rendering an output image derived from the input image using a rendering pipeline, the output image including makeup effects at locations associated with at least some of the facial features to simulate a virtual try-on of a makeup product, the effect having a 2D shape obtained from a predefined 2D mask image adjusted using the 3D shape of the locations (step 1104).
[0192] Mesh implementation 2: In mesh implementation 1, the face tracker separately generates facial points defining a shape outline of each of the located facial features; and the pipeline uses the 3D model and the facial points of at least some of the facial features to generate a 3D mesh to define the 3D shape of the location; warps a predefined 2D mask image to the 3D mesh to adjust the 2D mask image; and maps (e.g., unfolds) the warped 2D mask image to provide a 2D shape of the makeup effect at the location.
[0193] Mesh Implementation 3: In Mesh Implementation 2, the pipeline performs UV mapping to map the distorted 2D mask image using a 3D mesh.
[0194] Grid implementation 4: In any one of grid implementations 1 to 3, the makeup effect is associated with a makeup product, and the method includes: providing recommendations of multiple makeup products to be virtually tried via a user interface; and receiving selection input via the user interface to select a makeup product for the virtual trial experience.
[0195] Grid embodiment 5: In any of grid embodiments 1 to 4, the selection input selects one or more products for rendering one or more makeup effects at two or more locations associated with at least some facial features; and wherein the rendering uses a first predefined 2D mask image to render the first makeup effect at the first location, the first predefined 2D mask image adjusted using the 3D shape of the first location; and wherein the rendering uses a second predefined 2D mask image to render the second makeup effect at the second location, the second predefined 2D mask image adjusted using the 3D shape of the second location. In one example, the one or more products include eye shadow, the one or more makeup effects include an eye shadow effect, the first location and the second location are respective left and right eye regions associated with eye facial features and eyebrow facial features, and the first predefined 2D mask image and the second predefined 2D mask image are a left eye mask image and a right eye mask image, which can be an eye image and its mirror image.
[0196] Grid embodiment 6: In any of grid embodiments 1 to 5, the method includes providing a purchasing service via the user interface to conduct a purchase transaction (eg, via e-commerce) to purchase the cosmetic product.
[0197] It will be understood that system aspects and computer program product aspects are disclosed, for example, corresponding to each of grid embodiments 1 to 6. Any one or more of the grid embodiments may be combined with any one or more of the hue embodiments, glitter embodiments, and shaping embodiments as described.
[0198] In addition to the computing device aspects and method aspects, one of ordinary skill in the art will also understand that a computer program product aspect is disclosed wherein instructions are stored in a non-transitory storage device (e.g., memory, CD-ROM, DVD-ROM, optical disk, etc.) and when executed, the instructions cause a computing device to perform any of the method aspects stored therein.
[0199] Actual implementations may include any or all of the features described herein. These and other aspects, features, and various combinations may be expressed as methods, devices, systems, means for performing functions, program products, and in other ways combining the features described herein. A variety of embodiments have been described. However, it will be understood that various modifications may be made without departing from the spirit and scope of the processes and techniques described herein. In addition, other steps may be provided or steps may be removed from the described processes, and other components may be added to or removed from the described systems. Accordingly, other embodiments are within the scope of the appended claims.
[0200] Throughout the description and claims of this specification, the words "comprise" and "include" and their variations mean "including but not limited to", and they are not intended to (and do not) exclude other components, integers or steps. Throughout this specification, unless the context requires otherwise, the singular encompasses the plural. In particular, where the indefinite article is used, this specification will be understood to contemplate plurality as well as singularity, unless the context requires otherwise.
[0201] Unless incompatible therewith, features, integer properties, compounds, chemical moieties or groups described in conjunction with a particular aspect, embodiment or example of the invention are to be understood to be applicable to any other aspect, embodiment or example. All features disclosed herein (including any accompanying claims, abstract and drawings) and / or all steps of any disclosed method or process may be combined in any combination, except for at least some mutually exclusive combinations of such features and / or steps. The invention is not limited to the details of any foregoing examples or embodiments. The invention extends to any novel feature or any novel combination of features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel step or any novel combination of steps of any disclosed method or process.
Claims
1. A computer-implemented method comprising the steps of: processing an input image using a face tracking engine having one or more deep neural networks to locate facial features to respectively generate facial points defining an outline of each of the located facial features; as well as Rendering, using a rendering pipeline, an output image derived from the input image, the output image being derived by applying one or more shape changes to a particular facial feature, the one or more shape changes being determined by: mapping a grid of spaced-apart grid points to pixels of the particular facial feature and any associated facial features; as well as warping at least some of the spaced-apart grid points using corresponding shape changing functions, the warping changing positions of at least some of the spaced-apart grid points for changing positions of facial points of the specific facial feature; as well as wherein the rendering determines, in response to the warping, output pixels for the particular facial feature and any associated facial features for the output image.
2. The method according to claim 1, wherein The facial tracking engine and the rendering pipeline are components of a VTO application for simulating the effect of makeup products applied to facial features.
3. The method according to claim 2, wherein: The rendering pipeline renders the makeup effect to the specific facial feature that has been changed in shape, so that the output image includes the specific facial feature that has been changed in shape and has the makeup effect.
4. The method according to any one of claims 1 to 3, wherein The method includes providing a user interface to receive input of shape parameters defining the one or more shape changes, and wherein the rendering is responsive to the user input.
5. The method according to claim 4, wherein At least some of the shape changing functions perform one or more of the following: curve matching to match an intermediate curve along the middle of the contour of the particular facial feature to a target curve defined by the shape parameters of the one or more shape changes, thereby attenuating position changes to spaced-apart grid points in response to distance from the target curve; or point matching to match discrete points of the contour of the specific facial feature to target points defined by the shape parameters of the one or more shape changes; or Region expansion or region compression, to expand or compress a region along a facial point curve or around a specific pixel in response to the shape parameter, wherein the region expansion or the region compression is attenuated by a attenuation function in response to a distance from the facial point curve or the specific pixel.
6. The method according to any one of claims 1 to 5, wherein Determining the pixels includes fitting triangles to vertices defined by the warped spaced-apart grid points and UV mapping the pixels of the particular facial feature and any associated features onto the vertices.
7. The method according to any one of claims 1 to 6, wherein The particular facial feature defines a first facial feature, and the rendering step is repeated for a second facial feature to produce an output image having at least two facial features with changed shapes.
8. The method according to any one of claims 1 to 7, wherein The shape changes applied to the specific facial feature include any of the following: eyebrow shaping; nose shaping, such as narrowing of the nostrils or narrowing of the nose bridge; facial contour changes, such as narrowing of the chin; eye or eyelid changes, such as vertical or horizontal eye enlargement; or lip changes, such as lip filler.
9. A system comprising at least one processor and a memory storing instructions, the instructions being executable by the at least one processor to cause the system to: processing an input image using a face tracking engine having one or more deep neural networks to locate facial features to respectively generate facial points defining an outline of each of the located facial features; as well as Rendering, using a rendering pipeline, an output image derived from the input image, the output image being derived by applying one or more shape changes to a particular facial feature, the one or more shape changes being determined by: mapping a grid of spaced-apart grid points to pixels of the particular facial feature and any associated facial features; as well as warping at least some of the spaced-apart grid points using corresponding shape changing functions, the warping changing positions of at least some of the spaced-apart grid points for changing positions of facial points of the specific facial feature; as well as wherein rendering determines, in response to the warping, output pixels for the particular facial feature and any associated facial features for the output image.
10. A computer-implemented method comprising the steps of: processing an input image using a face tracking engine having one or more deep neural networks to locate facial features of a face; and An output image derived from the input image is rendered using a rendering pipeline, the output image including a makeup effect at locations associated with at least some of the facial features to simulate a virtual try-on of a makeup product, the effect having a 2D shape obtained from a predefined 2D mask image adjusted using a 3D shape of the locations.
11. The method according to claim 10, wherein: The facial tracker respectively generates facial points defining a shape outline of each of the located facial features; and the pipeline: generating a 3D mesh using a 3D model and the facial points of at least some of the facial features to define the 3D shape of the location; warping the predefined 2D mask image to the 3D grid to adjust the 2D mask image; and Mapping is performed by unfolding the distorted 2D mask image to provide the 2D shape of the makeup effect at the location.
12. The method according to claim 11, wherein The pipeline performs UV mapping to map the distorted 2D mask image using the 3D mesh.
13. The method according to any one of claims 10 to 12, wherein The makeup effects are associated with makeup products, and the method includes providing recommendations of a plurality of makeup products to be virtually tried on via a user interface; and receiving selection input via the user interface to select the makeup product for a virtual try-on experience.
14. The method according to any one of claims 10 to 13, wherein The selection input selects one or more products for rendering one or more makeup effects at two or more locations associated with the at least some facial features; and wherein the rendering renders a first makeup effect at a first location using a first predefined 2D mask image, the first predefined 2D mask image being adjusted using a 3D shape of the first location; and wherein the rendering renders a second makeup effect at a second location using a second predefined 2D mask image, the second predefined 2D mask image being adjusted using the 3D shape of the second location.
15. The method according to claim 14, wherein The one or more products include eye shadow, the one or more makeup effects include an eye shadow effect, the first position and the second position are corresponding left eye areas and right eye areas associated with eye facial features and eyebrow facial features, and the first predefined 2D mask image and the second predefined 2D mask image are left eye mask image and right eye mask image.
16. The method according to any one of claims 10 to 15, wherein The method includes providing a purchasing service to conduct a purchase transaction via the user interface to purchase a cosmetic product.
17. A computing device or computer program product comprising a memory storing instructions for execution by a processor to cause the processor to perform the method according to any preceding claim.
Citation Information
Patent Citations
Procede de reaction integre, notamment pour la production de clinker de ciment portland
FR2303587A1
System and method for light field correction of colored surfaces in an image
US10892166B2