METHOD AND SYSTEM FOR EFFECTIVE CONTOURING IN AN AUGMENTED REALITY EXPERIENCE
The described system addresses the challenge of rendering realistic contours in augmented reality experiences by using a combination of deep neural networks and shape models to accurately locate and reshape nail objects, resulting in an enhanced virtual try-on experience.
Patent Information
- Application Number
- FR2023003263
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-04-03
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-04-03
AI Technical Summary
Existing techniques for augmented reality experiences, particularly in virtual try-on (VTO), struggle to render realistic shapes or contours effectively.
A system comprising a nail localization engine, a nail shaping engine, and a rendering component, utilizing deep neural networks and shape models to accurately locate and reshape nail objects, and then render them with realistic effects.
The system achieves a more realistic and stable virtual try-on experience by accurately reshaping and rendering nail objects, enhancing user interaction and product simulation.
Smart Images

Figure 00000032_0000 
Figure 00000033_0000 
Figure 00000034_0000
Abstract
Description
Title of the invention: METHOD AND SYSTEM FOR EFFECTIVE CONTOURING IN AN AUGMENTED REALITY EXPERIENCE FIELD OF THE INVENTION
[0001] This application relates to image processing so as to locate an object in one or more images to render an effect associated with the object, and more particularly to effect contouring in an augmented reality experience such as a virtual try-on (VTO). CONTEXT
[0002] Deep learning techniques are useful for processing images, including a series of video frames to locate one or more objects within the images. In one example, the objects are portions of a user's body such as a face or a hand, particularly fingernails. Image processing techniques are also useful for rendering effects in association with these objects such as augmenting reality for the user. An example of such augmented reality is the provision of a VTO that simulates the application of a product to the object. Product simulation in the beauty industry includes simulating makeup, hair, and nail effects. Other examples may include locating the iris and simulating a change in color thereof, for example, by a colored contact lens. These and other VTO effects will be apparent.
[0003] Improved techniques are desired for rendering an effect to provide realistic shapes or contours, for example when providing augmented realities, including VTO experiences. SUMMARY
[0004] Embodiments of methods, apparatus, and techniques shape localized objects using a shape model for rendering a reshaped object with an effect such as for a virtual try-on experience (VTO). A VTO may simulate the finger / toe nail effects of a nail product or service on nail objects. An example system includes a nail localization engine having computational circuitry for locating one or more nail objects in an input image of a hand or foot via one or more deep neural networks; a nail shaping engine having computational circuitry for reshaping, via a trained shape model, the one or more nail objects as localized by the nail localization engine; and a A rendering component comprising computational circuitry for rendering an output image simulating a nail product or nail service applied to the one or more nail objects in accordance with the reshaping by the nail shaping engine to provide a virtual try-on experience.
[0005] A computer-implemented method is provided comprising executing on a processor one or more steps comprising: rendering an effect on an object having an original shape as detected in an input image, wherein the effect is applied to the input image in accordance with an updated shape obtained from a trained shape model that constrains the original shape in response to important shape features identified by the trained shape model.
[0006] A system is provided comprising: a nail localization engine having computational circuitry for locating one or more nail objects in an input image of a hand or foot via one or more deep neural networks; a nail shaping engine having computational circuitry for reshaping, via a trained shape model, the one or more nail objects as located by the nail localization engine; and a rendering component having computational circuitry for rendering an output image simulating a nail product or nail service applied to the one or more nail objects in accordance with the reshaping by the nail shaping engine to provide a virtual try-on (VTO) experience.
[0007] These and other aspects will be apparent to one skilled in the art. Brief description of the drawings
[0008] [Fig. 1] [Fig. 1] is a graph representing an extracted contour comprising a plurality of numbered points for an instance of an object represented in a class map.
[0009] [Fig.2] [Fig.2] is a graph representing a standardized contour converted from the contour of [Fig.l].
[0010] [Fig.3] [Fig.3] shows a graph representing a normalized contour constructed from the standardized contour of [Fig.2].
[0011] [Fig.4A] [Fig.4A] is a graph representing two contours, namely an input shape comprising the contour points of [Fig.3] and an output shape determined from the output points of a principal component analysis (PCA) model after applying dimensionality reduction using PCA. [Fig.4B] is an enlargement representing a portion of [Fig.4A].
[0012] [Fig.5] [Fig.5] is a graph representing two contours comprising an input shape with added noise and an output shape comprising the output points of the PCA model after applying dimensionality reduction using PCA.
[0013] [Fig.6A][Fig.6B] Figures 6A and 6B respectively show an image rendered without object shaping, and an image rendered with object shaping in accordance with an embodiment herein.
[0014] [Fig.7A][Fig.7B] Figures 7A and 7B respectively show an image rendered without object shaping, and an image rendered with object shaping in accordance with an embodiment herein.
[0015] [Fig.8A] [Fig.8B] Figures 8A and 8B respectively show a rendered image without object formatting, and an image rendered with object formatting in accordance with one embodiment herein.
[0016] [Fig.9A][Fig.9B] Figures 9A and 9B respectively show an image rendered without object shaping, and an image rendered with object shaping in accordance with an embodiment herein.
[0017] [Fig. 10] [Fig. 10] is an illustration of a computer network providing an environment for various aspects according to embodiments herein.
[0018] [Fig. 11] [Fig. 11] is a functional diagram of a network computing device computers of [Fig. 10].
[0019] [Fig. 12] [Fig. 12] is a flowchart of operations of a computing device in accordance with one embodiment.
[0020] [Fig. 13] [Fig. 13] is a graphical illustration of a CNN processing an image according to an example.
[0021] [Fig. 14] [Fig. 14] is a graphical illustration of a portion of the CNN of [Fig.13].
[0022] [Fig.l5A] [Fig.l5A] is a graphical illustration of a BlurZoom operation in accordance with one embodiment; and Figures 15B, 15C, 15D, 15E, 15F, and 15G are examples of tensor data related to BlurZoom operations in accordance with embodiments.
[0023] [Fig. 16] [Fig. 16] is a flowchart of operations of a computing device in accordance with one embodiment. DETAILED DESCRIPTION
[0024] In accordance with embodiments herein, one or more object localization and object shaping techniques are described. In embodiments described herein, the object class is a nail for nail coloring such as one or more fingernails of a hand or toenails of a foot. An image may include multiple instances of these objects. In an example of a hand or foot for nail localization, the number of objects is typically five. In embodiments herein, a deep neural network classifies and segments the nail objects.
[0025] In one embodiment, such as to provide a VTO experience from a "live" stream of video (e.g., a selfie video), each frame (e.g., each image) of the video is processed to locate the objects, generate a processed shape of the objects, and render the objects with one or more effects in accordance with the processed shape.
[0026] An example of a deep neural network for object localization is described in US2020 / 0349711A1 published on November 5, 2020, now US 11410314B2, and entitled “Image Processing Using a Convolutional Neural Network to Track a Plurality of Objects. An example of a deep neural network for object localization, as adapted from US2020 / 0349711A1, is described in more detail below with reference to Figures 13 and 14.
[0027] In one embodiment, object localization includes determining the location of one or more objects in an image. In one example, a deep neural network is trained as a semantic classifier to detect, for example, a particular object class. The deep neural network may perform semantic segmentation to process an image to provide a mask classifying each pixel to indicate whether or not a pixel is an object pixel, i.e., whether a pixel is part of the object class or is a background pixel. In one embodiment, a mask is a one-channel output of the network indicating a probability for each pixel of an input image, as processed by the network, of whether the pixel is a fingernail. The probability values are typically between 0 and 1.A threshold may be applied to the probability value of a pixel to determine which pixel is a nail pixel and which pixel is a background pixel, for example, to produce a threshold mask of binary values. In one embodiment, each pixel is assigned a 1 or a 0 in response to the threshold applied to the probability in the mask from the network. A threshold may be 0.5. The resulting thresholded mask is therefore a two-dimensional shape representation of the three-dimensional class object or objects in the image. Some deep neural networks can classify multiple classes of objects.
[0028] In one embodiment, the deep neural network further produces a class map. In a nail embodiment that locates five nails per hand, for example, the class map is a five-channel output, consisting of the probabilities for each finger class (pinky, ring, middle, index, and thumb). However, these probabilities are only valid within the nail regions defined by the nail mask. Therefore, the class map does not actually contain information indicating whether the current pixel is a nail or not.
[0029] For some tasks, a class map is useful instead of a segmentation mask because the mask tends to be noisier while the class map has cleaner edges. An augmented class map for the objects can be constructed to contain the background class, using the threshold mask. In one embodiment, to facilitate working with the class map and to define the augmented class map, the class map is processed by flattening the 5-channel class map into a 1-channel map where each pixel contains an integer representing the class, where the class is determined by the finger with the highest probability. The "background" class is also added using the data from the mask. The augmented class map can be used, in one embodiment, to find contours of the objects (fingernails), for example. It is noted that it is possible to use either the threshold mask or the augmented class map for a contour extraction step.
[0030] Additionally, in one embodiment, the deep neural network further produces a direction map indicating a direction of each object such as from a first end to a second end of the object (e.g., from a base to a tip or vice versa for a nail object).
[0031] Object localization is described in more detail below, after a description of object shaping. Object shaping refines the object shape from object shape information presented in the class map.
[0032] Improving an object shape
[0033] Object shape information in class maps may be inaccurate, and detection differences between frames of a video may highlight shape differences, particularly inaccuracies. One goal of the object shaping task is to improve the shape and stability of an object (e.g., when rendered into an output image).
[0034] From the class map, an outline of each object can be determined, where the outline includes a margin or border around the object following the outer contour of the object's shape as represented on the class map. In one embodiment, an augmented class map is defined and used for outline extraction. In one embodiment, a shape model is used to refine the shape of each object, in response to the outlines. The refined shape is useful for updating the mask.
[0035] In embodiments, rendering, the process of applying an effect to the object to construct an output image such as for a VTO experiment, is done in response to the processed object contours instead of the original class map contours for each detected object. A revised mask may be determined in accordance with the refined shapes as determined by applying the contours to a shape model. During rendering, the mask is used because the probability values make the rendered result appear smoother. In one embodiment, during rendering, the mask is used because the probability values make the rendered result appear smoother. rendering, the processing renders one nail at a time in the output image, so the class map is also useful for determining which mask pixels belong to the current class.
[0036] According to one embodiment, the object shaping steps comprise: 1. Extracting object contours from the class map; 2. For each outline: 3. Converting the outline into a “standardized” shape; 4. Contour normalization to remove rotation, position and alignment to scale; 5. Projecting the contour into a shape model (e.g. a PCA model) of reduced dimensionality; 6. Contour stabilization in model space; 7. Deprojection of the contour from the model space; 8. Reapplying rotation, translation and scaling to the contour; and 9. Generating a new mask from the contours.
[0037] Each step 1 to 9 is described in more detail.
[0038] Contour Extraction: In one embodiment, contours are extracted from the class map, e.g., augmented in accordance with the description above. However, the threshold mask may be used in one embodiment. Contour extraction may be performed using known techniques. An open-source example of a contour extraction function is "cv::findContours" available from OpenCV (Open source Computer Vision) at opencv.org. This function implements aspects of S. Suzuki and K. Abe, "Topological structural analysis of digitized binary images by border following" Computer Vision, Graphics, and Image Processing 30 (1): 32-46 (1985). The function is capable of automatically determining the number of contours and the location of each contour.
[0039] [Fig.l] depicts a graph 100 representing an extracted contour 102 comprising a plurality of points numbered 0 to 30 for an instance of an object represented in a class map (not shown). [Fig.l] is a constructed example of what a detected contour may look like. The x and y axes represent the position (x,y) in the coordinate system of the image graph 100 having a point 104 that represents the base and a point 106 that represents the tip of the object (a fingernail). These points are provided for illustration purposes and are not part of the output of the contour extraction function, for example.
[0040] Conversion of the contour into a standardized form: [Fig.2] is a graph 200 representing a standardized contour 202 converted from the contour 102. The contours extracted using an extraction function (e.g., contour 102) typically have a different number of points for different object instances. Also, the location of the points is typically different. To simplify working with extracted contour points, the extracted points are converted to a curve using a curve fitting method and resampled (e.g., to predefined and relatively evenly distributed locations) so that the same number (e.g., 25) of points per nail (object) are defined for use (e.g., points 0 to 24) and each respective point is located in approximately the same area of the nail (e.g., point 0 (204) is always at the tip of the nail object). In one embodiment, curve fitting using cardinal splines is used (see en.wikipedia.org / wiki / Cubic_Hermite_spline#Cardinal_spline).
[0041] Contour Normalization: Standardized contour information includes data dimensions that may complicate further analysis and do not relate solely to the shape itself. To facilitate further analysis of object shape, for example using principal component analysis (PCA) techniques, it is desirable that the shape modeling only models the shape and nothing else. The normalization operations eliminate rotation, position, and scaling of the standardized contours. [Fig. 3] shows a graph 300 representing a normalized contour 302 constructed from the normalized contour 202.In one embodiment, a normalization process includes: 1) averaging all contour points to obtain an estimate of the center of the object (alternatively, the centroid could be calculated); 2) subtracting the center point from all points to normalize the position; 3) rotating all points by the negative nail angle to normalize the rotation, e.g., where the nail angle is obtained from the direction map output by the nail detection model; and 4) calculating the minimum / maximum X-values and uniformly scaling the points so that the X-values are between -1 and 1 to normalize the scale. In one embodiment, uniform scaling is used to maintain the side ratio.
[0042] Projecting the contour into the model space: In one embodiment, the object shape is constrained using PCA techniques, including a PCA model. The PCA model is trained with the applicable object shape data. The PCA model has input data requirements. It is noted that other approaches may be used for a shape model. For example, a different dimensionality reduction method could be used, such as training an autoencoder using a neural network.
[0043] In general, PCA is used to learn a different representation of the contour data, where the newly learned space contains dimensions that are ordered according to the variance they account for in the data. For a PCA model that models, for example, fingernail shapes, the input representation in one embodiment comprises a set of 2D points and the output representation comprises a set of numbers that more abstractly represents the shape. For example, the first number might affect the high-level overall shape, while subsequent numbers might progressively affect finer and finer details of the shape.
[0044] Dimensionality reduction (on the PCA model) can be performed to constrain a nail shape by eliminating dimensions that model small details, which can result in a smoother, more natural-looking nail shape. The resulting PCA model data is then deprojected back into the input space (e.g., contours) to obtain a new set of 2D points.
[0045] For this task, in one embodiment, the PCA model serves three purposes. The PCA model contains knowledge about a valid nail shape and can be used to constrain the detected nail contour to correct for any atypical or noisy contours. Stabilization can be performed in PCA space, allowing the shape to be stabilized over time (e.g., between frames of a video). The model can also measure the amount of error between the current shape and the closest nail shape in the model. If the error is large, this may indicate a false detection or a very inaccurate contour and it may reject that particular nail from rendering. It may be better not to render a nail than to render it with a truly inaccurate shape. An example where the shape is inaccurate includes occlusion by a non-nail object or when two nail objects overlap and cannot be distinguished.Additionally, inaccuracy can be caused by a number of other factors such as: poor lighting, motion blur, unclear boundary between the nail and the background, and anything that would make it difficult to segment the nail correctly.
[0046] To train the PCA model, the contours of the nail training data are preprocessed using the same standardization and normalization steps described above. For each nail object, the input to the PCA model is 25 points (50 dimensions in total when flattened, due to the x and y point values). Approximately 6500 nail shapes were used for model training.
[0047] Dimensionality reduction is performed during projection onto the PCA space. PCA provides a means to perform a mapping from the input space to the PCA space. The transformation involves multiplying the input data by a projection matrix. This matrix is obtained during training. For more details, see en.wikipedia.org / wiki / Principal_component_analysis#Dimensionality_ reduction. There is a trade-off between the model's ability to constrain shape and the model's expressiveness. If the number of dimensions is low, the model will constrain nail shape very well, but may have difficulty modeling all possible nail shapes and poses. If a large number of dimensions are used, it can better model the various nail shapes and poses, but will perform worse at constraining shape. Experimentation has shown that reducing it from 50 to 15 dimensions provides an acceptable balance in the above trade-off. This reduction can be performed after training the shape model and includes "trimming" the projection matrix based on the number of dimensions.
[0048] [Fig.4A] is a graph 400 representing two contours, namely an input shape comprising the points of contour 302 and an output shape (contour 402) determined from the output numbers of the PCA model after applying dimensionality reduction using PCA. In this example, since the input is already relatively smooth, the output is not significantly different from the input. [Fig.4B] is an enlargement representing a portion of [Fig.4A] at area 404 showing the numbered points of contour 302 and contour 402. The points of contour 302 are represented as a hollow circle and are immediately adjacent to a lower left corner of the point number.
[0049] [Fig. 5] is a graph 500 representing two contours, namely an input shape comprising the points of a contour 502 and an output shape (points of the contour 504) determined from the output numbers of the PCA model after applying dimensionality reduction using PCA. The input contour 502 of [Fig. 5] shows a different example where noise is intentionally added to the input curve to simulate noise detection. Using PCA dimensionality reduction eliminates almost all of the noise.
[0050] Stabilization in PCA Space: Optionally, in one embodiment, the shape data is stabilized over time. In a "live" mode (e.g., in a component of a VTO experience where frames of a video are rendered with effects applied to one or more objects within the frames), the PCA data (e.g., in PCA space) is stabilized to stabilize the shape over time (e.g., from frame to frame). In one embodiment, an exponential moving average is applied to the PCA data, where the weight is determined by the speed of the nails to minimize the perceived delay. In one embodiment, the speed is determined by calculating the change in center position of the contour from frame to frame. A history of N current and past speeds is maintained and the average is calculated to produce a final speed per nail. In one embodiment, N = 2.When the nails move, little stabilization is done and when . the nails are stationary, a large amount of stabilization is used. In one embodiment, the same weight is applied to all components. In another embodiment, the components are weighted differently. For example, the most important components (which model the overall shape) may be less stabilized than the least important components (which model the small details of the shape). Stabilization is not applicable to single-frame processing and rendering because no time domain is present in the data.
[0051] It is preferable to stabilize in the PCA data space, although it would be possible to stabilize after deprojection from the PCA space, for example after reapplying the transformations. An advantage of stabilizing within the PCA space is that it focuses only on stabilizing the essential aspects of the shape, rather than affecting rotation, position, etc. Position stabilization can produce a very noticeable delay, which is not a preferred result.
[0052] Deprojection from PCA space: In this step, the PCA data is deprojected back into the original space (contour space) to obtain the new contour.
[0053] Reapply rotation, translation and scaling: Rotation, translation and scaling are reapplied to the new contour.
[0054] Generating a new mask: A new nail mask image (e.g., a render image) is generated by tracing the new contours for each nail into a blank image. A new class map is also generated because the nail shapes have now changed.
[0055] Rendering may be performed in accordance with the new mask, for example, to render an effect within the region of the object defined by the updated outline, namely on the pixels therein, as identified by the new mask. In one example, the effect is a nail effect, for example, applying a nail color to simulate nail polish or a nail texture. Other nail effects may be applied in response to the shaped object. For example, the nail shape may be applied, for example, to extend the length to simulate a French manicure, an acrylic type or other type of nail shape material, etc. Results of the example object shaping
[0056] Figures 6A and 6B show, in connection with a first example input image (not shown), a rendered image without object shaping 600A, and a rendered image with object shaping 600B in accordance with one embodiment herein. Figures 7A and 7B show, in connection with a second example input image (not shown), a rendered image without object shaping 700A and a rendered image with object shaping 700B in accordance with one embodiment herein. Figures 8A and 8B show, in connection with a third example input image (not shown), a rendered image without object shaping 800A and a rendered image with object shaping 800B in accordance with one embodiment herein. Figures 9A and 9B show, in connection with a fourth example input image (not shown), a rendered image without object shaping 900A and a rendered image with object shaping 900B in accordance with one embodiment herein. Figures 6A and 6B, 7A and 7B, 8A and 8B and 9A and 9B are enlarged portions of hand images to more clearly show the differences provided by the object shaping.
[0057] [Fig. 10] is an illustration of a computer network providing an environment for various aspects according to embodiments herein. [Fig. 11] is a block diagram of a computing device of the computer network of [Fig. 10].
[0058] Referring to [Fig. 10], an exemplary computer network 1000 is shown in which a personal computing device 1002 is operated by a user 1004. The device 1002 is in communication via a communications network 1006 with remotely located server computing devices, namely server 1008 and server 1010. The user 1004 may be a consumer of nail-related products or services. Also shown are a second user 1012 and a second computing device 1014 configured for communication via the communications network 1006, such as to communicate with the server devices 1008 or 1010. The second user 1012 may also be a consumer of nail-related products or services.
[0059] Briefly, each of the computing devices 1002 and 1014 is configured to perform a virtual try-on of nail products or services. In one embodiment, a neural network for locating nail objects is stored and used onboard the computing device 1002 or 1014. In an alternative embodiment or it may be provided from the server 1008 such as via a cloud service, a web service, etc. from an image or images received from the computing device 1002 or 1014. In one embodiment, the neural network is as described herein and outputs masks and class maps when processing images. Thus, devices 1002, 1014, or 1008 may each include a nail object localization engine having computational circuitry configured to localize one or more nail objects in an input image of a hand or foot via one or more deep neural networks.
[0060] The one or more located objects may be reshaped using a trained shape model as described herein. Thus, devices 1002, 1014, or 1008 may each include a nail shaping engine having computational circuitry for reshaping, via a trained shape model, the one or more nail objects as located by the nail localization engine. In one embodiment, the trained shape model reduces the dimensionality of data to identify important features. In one embodiment, the trained model is defined and trained in accordance with principal component analysis techniques.
[0061] The reshaped nail objects are useful for rendering effects on the input images, such as for simulating a product or service providing a virtual try-on (VTO) experience to a user. Thus, devices 1002, 1014, or 1008 may each include a rendering component including computational circuitry for rendering an output image simulating a nail product or service applied to one or more nail objects in accordance with the reshaping by the nail shaping engine to provide a VTO experience. In one embodiment, devices 1002 and 1014 include respective displays for visualizing the VTO experience.
[0062] Device 1008 or 1010 may provide one or more cloud-based and / or web-based services such as one or both of a service for recommending nail-related products or services (e.g., nail salon services) and a service for purchasing nail-related products or services via an e-commerce transaction. Thus, for example, each of devices 1002, 1014, 1008, or 1010 may include a nail VTO component including computational circuitry for one or both of a recommendation of a nail product or nail service for a virtual try-on, and a purchase of the nail product or nail service via an e-commerce transaction. In one embodiment, the nail VTO component of devices 1002 and 1014 may communicate with one or both of servers 1008 and 1010 to obtain a recommendation or complete a purchase transaction.It is understood that the devices linked to payment services for carrying out an e-commerce transaction are not shown for the sake of simplicity, nor are the storage components storing product or service information for such recommendations and / or purchases.
[0063] Devices 1002 and 1014, in one embodiment, include respective cameras such as for capturing input images, which may include self-portrait images, including videos having a series of frames. The input images for processing may be frames of the video, including, for example, a current frame of a video. Each of elements 1002, 1014, or 1008 may include a stabilization component having computational circuitry for stabilizing the reshaping of the one or more nail objects between the series of frames of the video.
[0064] The computing device 1002 is shown as a portable mobile device (e.g., a smartphone or tablet). However, it may be another computing device such as a laptop, desktop computer, workstation, or other device. work, etc. Similarly, the device 1014 may take another form factor. The computing devices 1002 and 1014 may be configured using one or more native applications or browser applications, for example, to provide the VTO experience.
[0065] [Fig. 11] is a block diagram of the computing device 1002, in accordance with one or more embodiments of the present disclosure. The computing device 1002 includes one or more processors 1102, one or more input devices 1104, a gesture I / O device 1106, one or more communication units 1108, and one or more output devices 1110. The computing device 1002 also includes one or more storage devices 1112 storing one or more modules and / or data.Modules may refer to software modules, a component of such a device 1002 and may include a deep neural network model 1114 such as for a nail localization engine, a (trained) shape model 1116 such as for a nail shaping engine, a module for a rendering component 1118, a module for a stabilization component 1120, user interface modules such as components for a graphical user interface (GUI 11120) and a module for image acquisition 1124. The data may include one or more images for processing or as output from processing (e.g., images 1130), etc.
[0066] The modules such as when executed by the one or more processors 1102 provide the functionality to acquire one or more images such as a video and process the images to provide the VTO experience. In another example (not shown), one or more of the neural network model, trained shape model, rendering component, and stabilization component are located remotely (e.g., on the server 1008, 1010, or another computing device). The computing device 1002 may communicate one or more input images (e.g., from the images 1130) for processing and return.
[0067] The storage device(s) 1112 may store additional modules such as an operating system 1132 and other modules (not shown), including communication modules; graphics processing modules (e.g., for a GPU of processors 1102); a map module; a contacts module; a calendar module; a photo / gallery module; a photo (image / media) editor; a media player and / or streaming module; social networking applications; a browser module; etc. The storage devices may be referred to as storage units herein.
[0068] The one or more processors 1102 may implement functionality and / or execute instructions within the computing device 1002. For example, the processors 1102 may be configured to receive instructions and / or data from the storage devices 1112 in order to perform the functionality of the modules shown in [Fig. 11], among others (e.g., operating system, applications, etc.). The computing device 1002 may store data / information on the storage devices 1112. Some of the functionality is further described below. It is understood that the operations may not correspond exactly to the modules shown in the storage device 1112 of [Fig. 11] so that one module may assist in the functionality of another.
[0069] The computer program code for performing operations may be written in any combination of one or more programming languages, for example, an object-oriented programming language such as Java, Smalltalk, C++ or the like, or a conventional procedural programming language, such as the "C" programming language or similar programming languages.
[0070] The computing device 1002 may generate an output for display on a screen of the gesture I / O device 1106 or, in some examples, for display by a projector, monitor, or other display device. It will be understood that the gesture I / O device 1106 may be configured using various technologies (e.g., in connection with input capabilities: a resistive touchscreen, a surface acoustic wave touchscreen, a capacitive touchscreen, a projective capacitive touchscreen, a pressure-sensitive screen, an acoustic pulse recognition touchscreen, or other presence-sensitive screen technology; and in connection with output capabilities: a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a dot matrix display, an e-ink, monochrome, or similar color display.
[0071] In the examples described herein, the gesture I / O device 1106 includes a touchscreen device capable of receiving as input a touch interaction or gestures from a user interacting with the touchscreen. These gestures may include tapping, sliding, or swiping gestures, flicking gestures, pause gestures (e.g., when a user touches a same location on the screen for at least a threshold period) where the user touches or points to one or more locations of the gesture I / O device 1106. The gesture I / O device 1106 may also include non-tapping gestures. The gesture I / O device 1106 may output or display information, such as the graphical user interface, to a user.The gesture I / O device 1106 may present various applications, functions and capabilities of the computing device 1002, including, for example, acquiring images, viewing images, processing the images and displaying new images, messaging applications, telephone communications, contact and calendar applications, applications . web browsing applications, gaming applications, e-book applications, and financial, payment and other applications or functions, among others.
[0072] Although the present disclosure illustrates and discusses a gesture I / O device 1106 primarily as a display screen device with I / O capabilities (e.g., touch screen), other example gesture I / O devices may be used to detect movements and may not include a screen per se. In this case, the computing device 1002 includes a display screen or is coupled to a display device to present new images and GUIs. The computing device 1002 may receive gesture input from a touchpad / touch pad, one or more cameras, or another presence- or gesture-sensitive input device, where presence means aspects of a user's presence, including, for example, movement of all or part of the user.
[0073] One or more communication units 1108 may communicate with external devices (e.g., server 1008, server 1010, second computing device 1014) as for the described purposes and / or for other purposes (e.g., printing), such as via the communication network 1006, by transmitting and / or receiving network signals over one or more networks. The communication units may include various antennas and / or network interface cards, chips (e.g., global positioning system (GPS)), etc. for wireless and / or wired communications.
[0074] The input devices 1104 and the output devices 1110 may include one or more buttons, switches, pointing devices, cameras, a keyboard, a microphone, one or more sensors (e.g., biometric, etc.), a speaker, a buzzer, one or more LEDs, a haptic (vibrating) device, etc. One or more of these may be coupled via a universal serial bus (USB) or other communication channel (1138). A camera (an input device 1104) may be facing forward (i.e., on the same side) to allow a user to capture one or more images using the camera while looking at the gesture I / O device 1106 to take a "self-portrait."
[0075] The one or more storage devices 1112 may take various forms and / or configurations, for example, such as short-term memory or long-term memory. The storage devices 1112 may be configured for short-term storage of information as volatile memory, which does not retain stored content when power is removed. Examples of volatile memory include random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), etc. The storage devices storage 1112, in some examples, also includes one or more computer-readable storage media, e.g., for storing larger amounts of information than volatile memory and / or for storing such information long-term, retaining the information when power is removed. Examples of non-volatile memory include magnetic hard drives, optical discs, floppy disks, flash memory, or forms of electrically programmable memory (EPROM) or electrically erasable programmable memory (EEPROM).
[0076] Although not shown, a computing device may be configured as a training environment to train the neural network model 1114, for example with appropriate training and / or testing data.
[0077] The deep neural network may be adapted to a lightweight architecture for a computing device that is a mobile device (e.g., a smartphone or tablet) having fewer processing resources than a "larger" device such as a laptop, desktop, workstation, server, or other computing device of comparable generation.
[0078] It is understood that the second computing device 1014 may be configured similarly to the computing device 1002. The second computing device 1014 may have GUIs such as to request and display one or more images and diagnoses of skin signs from data stored on the server 1008 for different users, etc.
[0079] [Fig. 12] is a flowchart of operations 1200 such as for performing a method. The method, in one embodiment, is a computer-implemented method executed by a processor to perform one or more steps. At 1202, the operations render an effect on an object having an original shape as detected in an input image (e.g., by a deep neural network for object localization), wherein the effect is applied to the input image in accordance with an updated shape obtained from a trained shape model that constrains the original shape in response to important shape features identified by the trained shape model.
[0080] Embodiment 1: In this embodiment, the method comprises, for example, step 1202. Additional embodiments are apparent, including the following numbered embodiments.
[0081] Embodiment 2: According to embodiment 1, the rendering is a component of a method for providing an augmented reality experience, such as a VTO experience, in which the effect simulates the application of a product to the object.
[0082] Embodiment 3: In embodiment 1 or 2, the object is a fingernail or toenail and the effect is a nail effect simulating a nail product.
[0083] Embodiment 4: In embodiment 3, the object is a first object of a plurality of objects including a second object including another nail of a hand or a toe and wherein the method includes locating each of the first object and the second object and tracking the respective locations of the first object and the second object between frames of a video. The rendering further includes rendering a first effect on the first object and rendering a second effect, different from the first effect, on the second object, the first effect being applied in accordance with an updated first shape obtained from the trained shape model for the first object and the second effect being applied in accordance with an updated second shape obtained from the trained shape model for the second object. For example, embodiment 4 provides nails having different colors, nail decoration effects, or other effects.The location and tracking of the different nails are described at least with reference to class maps and augmented maps and the five nails of a single hand.
[0084] Embodiment 5: In any one of embodiments 1 to 4, the one or more steps further comprise: obtaining an outline of the object; using the trained shape model to determine important features of the original shape in response to the outline; constraining the outline in response to the important features; and generating a rendering mask using the outline as constrained, the rendering mask to be used when rendering the effect on the object. Embodiment 6: In embodiment 5, the one or more steps may further comprise: standardizing and normalizing the outline before applying the outline to the trained shape model, and, in connection with the outline as constrained by the trained shape model, canceling the normalization before generating the rendering mask.
[0085] Embodiment 6: In any one of embodiments 1 to 5, one or both of the following are true: the trained shape model reduces the dimensionality of data to identify important features in response to the outline of the original shape and the trained shape model is defined and trained in accordance with principal component analysis techniques.
[0086] Embodiment 7: In any one of embodiments 1 to 6, the input image comprises a plurality of objects of a same class (e.g., five nail objects) and the rendering step is performed for each of the plurality of objects as detected in the input image (e.g., to render nail effects on each nail object).
[0087] Embodiment 8: In any one of embodiments 1 to 7, the input image comprises a current frame of a series of frames of a video and wherein the one or more steps comprise, prior to rendering, stabilizing the shape updated in response to an object's velocity in the series of video frames, including the current frame.
[0088] Embodiment 9: In embodiment 8, the input image is a current frame of a video stream and the method comprises: detecting a location of the object in a previous frame of the video step; determining the velocity of the object according to the location in the previous frame and a location in the current frame; applying an exponential moving average to the contour as constrained when a weight is determined by the velocity.
[0089] Embodiment 10: In any one of embodiments 1 to 9, the trained shape model determines whether or not a shape detection error exists for the original shape according to the shape model. In such an embodiment, the rendering step may be performed according to an absence of a shape detection error for the object in the input image and the effect is not applied to the object in the presence of the shape detection error.
[0090] Computer device embodiments and computer program product embodiments corresponding to one or more of Embodiments 1-10 will be apparent.Embodiment 11 is further provided: A system comprising: a nail localization engine having computational circuitry for locating one or more nail objects in an input image of a hand or foot via one or more deep neural networks; a nail shaping engine having computational circuitry for reshaping, via a trained shape model, the one or more nail objects as located by the nail localization engine; and a rendering component having computational circuitry for rendering an output image simulating a nail product or nail service applied to the one or more nail objects in accordance with the reshaping by the nail shaping engine to provide a virtual try-on (VTO) experience. Additional embodiments included the following numbered embodiments.
[0091] Embodiment 12: In embodiment 11, the system includes a nail VTO component having computational circuitry for one or both of recommending the nail product or nail service for virtual try-on, and purchasing the nail product or nail service via an e-commerce transaction.
[0092] Embodiment 13: In embodiment 11 or 12, the input image comprises a current frame of a series of frames of a video and the system comprises a stabilization component having computational circuitry to stabilize the reshaping of the one or more nail objects between the series of frames of the video.
[0093] Embodiment 14: In embodiment 13 with respect to a single hand, the one or more nail objects comprise five objects including a pinky finger nail object, a ring finger nail object, a middle finger nail object, an index finger nail object and a thumb nail object and the nail localization engine comprises computational circuitry for locating each of the five objects and tracking the respective locations of the five objects between frames of a video and wherein the rendering component comprises computational circuitry for rendering a first effect simulating a product or service on at least one of the five objects and rendering a second effect, different from the first effect, on at least one different object among the five objects.
[0094] Embodiment 15: In any one of embodiments 11 to 14, one or both of the following are true: the trained shape model reduces the dimensionality of data to identify important features; and the trained model is defined and trained in accordance with principal component analysis techniques.
[0095] Embodiment 16: In any one of embodiments 11-15, the one or more deep neural networks are configured for anti-aliasing in each of the encoder and decoder components, the encoder components performing anti-aliasing when downsampling and the decoder components performing anti-aliasing when upsampling. Computer-implemented method embodiments and computer program product embodiments corresponding to one or more of embodiments 11-16 will be apparent. Object Location - Antialiasing
[0096] It is desirable for convolutional neural networks (CNNs) to be shift-equivariant. This means that shifting an input by a few pixels in a given direction should produce an output equal to the original output shifted in the same direction. However, due to aliasing, CNNs produce different outputs even for shifted versions of the same input. Antialiasing is discussed in Zhang, R. Making Convolutional Networks Shift-Invariant Again, International Conference on Machine Learning (ICML) 2019, (available at the time of filing at arxiv.org / abs / 1904,11486).
[0097] [Fig. 13] is a graphical illustration of a CNN 1300 in accordance with an embodiment of processing an image in accordance with an example of enhanced aliasing. The CNN 1300 is adapted for anti-aliasing using the CNN of US2020 / 0349711A1. The CNN 1300 is an example of a deep neural network for object localization and segmentation, e.g., for fingernail objects with enhanced anti-aliasing.
[0098] When performing object localization, in particular segmentation of objects, aliasing is introduced by resampling operators in a segmentation CNN, such as the deep neural network described in US2020 / 0349711A1. Resampling occurs at the downsampling and upsampling layers when the input is resized to a lower or higher resolution (under- and upsampling, respectively). In the nail CNN of US2020 / 0349711A1, downsampling occurs in the encoder at the convolutional layers with stride 2 (last convolutional layers of each “stage” in the unadapted CNN). Upsampling occurs in the decoder of the unadapted CNN when lower resolution feature maps are upsampled and combined with higher resolution feature maps (“upper” layers in the unadapted CNN).To counteract aliasing, antialiasing is introduced both in the 2-step convolution layers of each stage and in the upsampling layers of the upper layers, as described in more detail.
[0099] Figure 13 shows the CNN 1300, as adapted for antialiasing in accordance with one embodiment, and processing an input (image) 1302 using two branches. Figure 14 is a graphical illustration of a portion of CNN 1300. The first branch 1300A (top branch in Figure 13) includes blocks 1304-1324. The second (bottom) branch 1300B of Figure 13 includes blocks 1326-1338. It will be appreciated that these bright line distinctions may be modified. For example, block 1326 may be a block of the first branch 1300A. Block 1304 is a 2x block of area subsampling. Although not shown, BlurPool may be used herein instead of area subsampling. Blocks 1306 to 1320 (also called stage_lowl, stage_low2... stage_low8) are encoder-decoder backbone blocks (having one encoder phase and one decoder phase) as described in more detail.Blocks 1306 to 1320 are suitable for antialiasing as described in more detail. Block 1322 is an x2 upsampling block and block 1324 is a first branch merge block as described in more detail. Block 1322 is suitable for antialiasing as described in more detail. Block 1326 is also an x2 upsampling block. Block 1326 is suitable for antialiasing as described in more detail. Blocks 1328 to 1334 (also called stage_highl, stage_high2... stage_high4) are blocks of an encoder phase as described in more detail. Blocks 1328 to 1334, similar to blocks 1306 to 1320, are suitable for antialiasing as described in more detail. Block 1338 oversamples x 8, then provides the resulting feature map to decoder model 1340 described in more detail with reference to [Fig. 14]. Block 1338 is suitable for antialiasing, as are blocks 1322 and 1326 described in more detail below.
[0100] In one embodiment, the encoder-decoder backbone network is modeled on MobileNetV2 (See, Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted Residuals and Linear Bottlenecks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018, as further adapted for antialiasing, as described herein. In one embodiment, the encoder phase (blocks 1328 to 1332) of the second branch is also modeled on the encoder of MobileNetV2) as further adapted for antialiasing as described herein.
[0101] In one embodiment, encoder blocks 1306-1320 and 1328-1334 are adapted from similar blocks of the original nail CNN. At blocks 1306-1320 and 1328-1334, all original nail CNN operators of type: [Conv(stride=2) -> BatchNorm -> ReLU] become [Conv(stride=1) -> BatchNorm -> ReLU -> BlurPool(stride=2)] in the CNN 1300.
[0102] BlurPool is implemented as described in Zhang, R. Making Convolutional Networks Shift-Invariant Again, ICML 2019. An input feature map of shape HxLxC is first convolved with a low-pass filter. The size and type of the filter are a hyperparameter of the process it contains. Then, the BlurPool-processed feature map is downsampled by the stride. In one embodiment for nails, the stride is 2 for the CNN 1300.
[0103] In one embodiment, all nearest neighbor or bilinear upsampling operators in the original nail CNN become a custom BlurZoom operator (step=2) in the CNN 1300. [Fig. 15A] is a graphical illustration of a BlurZoom operation 1500 performed during upsampling, in accordance with one embodiment. In one embodiment, BlurZoom comprises two steps: a zero insertion operation 1500 as shown in [Fig. 15A] followed by a convolution with a low-pass filter. Figures 15B-15G are examples of tensor data related to BlurZoom operations in accordance with the embodiments described later. Zero insertion intersperses a group (e.g. 4 x 4) of existing values 1502 with a plurality of zero values, to produce an oversampled group (e.g. 8 x 8) where ¾ of the resulting values are zero values.The upsampled pool is then convolved with a filter (e.g., Gaussian or otherwise). In one embodiment, the BlurZoom low-pass filters are the same filters used in BlurPool. If the filter sums to 1, the operations are multiplied by 4. It can be understood that if a filter sums to 1, this implies that the filter preserves the average intensity of the output. Normally, most filters, including low-pass filters, sum to one. In BlurZoom, since zeros are inserted (or theoretically inserted as described below), the resulting average intensity is decreased. For example, when 3 zeros are inserted around . At each value, the average intensity is reduced to 1 / 4. A multiplication by 4 counteracts this reduction to maintain the average intensity preservation property of the filter.
[0104] In the BlurPool and BlurZoom operators, low-pass filters provide a form of anti-aliasing to achieve shift equivariance. Normally, stepwise convolution is not shift equivariant, because the output values can change significantly when the input is shifted slightly due to the discrete downsampling that occurs. BlurPool seeks to address this problem by applying a low-pass filter before downsampling. Intuitively, the low-pass filter can be thought of as "diffusing" each value into neighboring regions (of values), so that downsampling becomes less sensitive to shifts as the values are "diffusing."
[0105] Consistent with the teachings and techniques herein, in the case of BlurZoom, the low-pass filter "broadcasts" values into the newly inserted rows and columns (actually or theoretically), so that any subsequent operation receives an offset-equivariant input.
[0106] As described with reference to [Fig.15A], in one embodiment without any optimization, the BlurZoom operation is implemented by first inserting zeros around each value, and then applying a convolution to each value for low-pass filtering. For example, in the case of a 3x3 convolution, the computation involves 9 multiplications and 8 additions per value. To optimize this, in one embodiment, the BlurZoom operations may ignore zeros in the computation.
[0107] Figures 15B, 15C, 15D, 15E, and 15F are examples of tensor data (values) 1506, 1508, 1510, 1512 in connection with BlurZoom operations in accordance with embodiments for illustrating the optimization. Letters A through D in Figures B through E represent existing values before zero insertion. In the case of a 3x3 convolution with step-2 zero insertion, a central value (e.g., A in the center of the 3x3 window) will either be surrounded by zeros (1506), the central value will be zero and will be surrounded by existing values (e.g., non-zero values) in the corners (1506), or the existing values are scattered as shown in 1510 and 1512.
[0108] In the first case (1506), if the BlurZoom operations ignore zeros, the BlurZoom operations need only do one multiplication and no additions. In the second case, the BlurZoom operations need only do 4 multiplications and 3 additions. The latter two cases require 2 multiplications and one addition. This gives an average of about 2.25 multiplications and 1.25 additions per output value. In one embodiment, this can be further optimized by omitting the zero insertion step. In the above example, the operations BlurZoom could create a new tensor that is double in width and height, and calculate each value in the tensor using the above approach, determining which scenario from Figures 15B to 15E applies depending on whether the row and column numbers are odd or even.
[0109] It should be noted that tensor boundaries would be a special case. When the convolution requires a value outside the tensor, it would "mirror" the values near the edge. For example, as shown at 1514 in [Fig. 15F], if there are values in all positions, with an "A" in the top left corner of the tensor, then a 3x3 convolution at "A" would see the result at 1516 of [Fig.l5G]. Note that in a real-world scenario, most of these values are zero.
[0110] Thus, in one embodiment, the 3X3 convolution is optimized to reduce multiplication or addition operations in response to zero values. It is noted that zero values do not necessarily have to be present (i.e., inserted). Zeros are theoretically present when the operations consider zero expansion without inserting the values.
[0111] Table 1 shows a detailed layer-by-layer description. Table 1 presents a detailed summary of the nail segmentation model architecture. Each layer name corresponds to the blocks of Figures 13 and 14 as described herein. The height H and width VF refer to the input size of H x integral resolution. For the projection 1408 and dilated layers 1410, e.g. { 16, 8} . For stages stage3_low to stage7_low, the number of channels in parentheses corresponds to the first layer in the stage (not shown), which increases to the number without parentheses for subsequent layers in the same stage.
[0112] [Tables 1] layer name size of iOS output TF.js stage l_low 4 X JL 3x3, 32, not 2 stage2_low H 4 1x1.24. x 2 stage4_low H 16 XW 16' 1x1192(144) 3x3192(144) . 1x1, 32x3 stage5_low H 32 XW 32' W x 1384(192)' 3x3384(192) . 1x1.64 . x 4 stage6_low H 32 XW 32 ■1x1576(384)' 3x3576(384) . 1 x 1.96. x 3 stage7_low H 32 stage8_low 1 X 1960 N / A 3x 3960, dilated x 2 x 1 1x1320 stage l-4_high same as stage l-4_low with I / O size x2 projection / ? x ~ lx 1320 dilated_p 7 X y [3x3320, dilated] % 2
[0113] The CNN decoder 1300 is shown in the middle and bottom right of Figure 13 (e.g., blocks 1324 and 1336 (comprising the fusion blocks) and upsampling blocks 1322 and 1326 adapted for antialiasing) and a detailed view of the decoder fusion model for each of blocks 1324 and 1336 is shown in Figure 14. For an original input of size HX IV, , the decoder merges the JL X ~~ features of stage_low4 (from block 1312) with the upsampled features of block 1322 derived from stage_low8, then upsamples (block 1326) and merges the resulting features via merge block 1336 with the S x — Par-8 8 particularities of stage_low4 (block 1334).
[0114] Figure 14 shows the fusion module 1400 used to fuse the upsampled high and low resolution semantic information features represented by the feature map F { (1402) with the low and high resolution semantic information features represented by the feature map F2 (1404) to produce high resolution fused features represented by the feature map F2 (1406) in the decoder using blocks 1408, 1410, 1412 and adder 1414. In connection with block 1324, the feature map F (1402) is output from block 1322 and the feature map F2 (1404) is output from block 1312. The feature map F2 (1406) of block 1324 is upsampled in 1326 to be provided to block 1336 as feature map F^ (1402) in the instance of model 1400 of that block.As noted, the upsampling performed at blocks 1322 and 1326 is suitable for antialiasing. At block 1336, the F2 feature map (1404) is received as output from block 1334 and the F2 feature map (1406) is output at block 338. Block 1338 upsamples to the input resolution / 4 and then provides the resulting feature map to decoder model 1340. Block 1338 is suitable for antialiasing as noted. Decoder model 1340 produces three types of information for the image (e.g., a 3-channel output 1342) as described below.
[0115] As shown in Figure 14, a 1 x 1 convolutional classifier 1412 is applied to the upsampled Fl features, which are used to predict downsampled labels. This "Laplacian pyramid" of outputs (See, Golnaz Ghiasi and Charless C. Fowlkes. Laplacian reconstruction and refinement for semantic segmentation. In ECCV, 2016) optimizes smaller, higher-resolution receptive field feature maps to focus on refining predictions from larger, lower-resolution receptive field feature maps. Thus, on the 1400 model, the feature map (not shown) from block 1412 is not used as output per se. Rather, in training, the loss function is applied as a pyramid output regularization.
[0116] Block 1342 represents an overall output of decoder 1340 that includes three channels corresponding to the outputs of the blocks of the three decoder branches (not shown). A first channel includes a per-pixel classification (e.g., a foreground / background mask or object segmentation masks), a second channel includes a classification of the segmented masks into individual fingertip classes, and a third channel includes a 2D directionality vector field per pixel of the segmented mask (e.g., (%,y) per pixel).
[0117] In one embodiment, decoder 1340 uses multiple output decoder branches (e.g., three) to provide directionality information (e.g., base-to-tip vectors in the third channel) needed for rendering on fingernail tips, and fingernail class predictions (in the second channel) needed to find fingernail instances using connected components. These additional decoders are trained to produce penalized dense predictions only in the annotated fingernail region of the image. Each branch employs a respective loss function according to the example. While, in one embodiment, a normalized exponential (Softmax) function is used in two of the branches, another activation function for segmentation / classification may be used.It will be understood that the dimensions here are representative and can be adapted to different tasks. For example, in the 1340 decoder, two branches concern 10 classes (one per nail on two hands for example), and are sized accordingly.
[0118] In one embodiment, the CNN 1300 provides a deep neural network for a nail localization engine to locate one or more nail objects in an input image of a hand or foot via one or more deep neural networks.
[0119] In one embodiment, an output of the CNN 1300 is provided to a rendering component, e.g., to use one or more of the mask, the class map, and the directional information to render effects on the detected objects. For example, nail effects are applied to the detected nails from the CNN 1300, e.g., to provide a virtual try-on (VTO) experience.
[0120] In one embodiment, an output of the CNN 1300 is provided to a nail shaping engine for reshaping, via a trained shape model, the one or more nail objects as located by the nail localization engine. The reshaped nail objects are provided to a rendering component to render an output image simulating a nail product or nail service applied to the one or more nail objects in accordance with the reshaping by the nail shaping engine to provide a virtual try-on (VTO) experience.
[0121] [Fig. 16] is a flowchart of operations 1600 such as for a method according to one embodiment. The method may include executing by one or more processors the steps shown in [Fig. 16]. In one embodiment lisation 17, a method comprises: (at 1602) localizing one or more nail objects in an input image of a hand or foot via one or more deep neural networks (e.g., of a tracker engine), wherein the one or more deep neural networks are configured for antialiasing in each of the encoder and decoder components; and (at 1604) a rendering component having computational circuitry to render an output image simulating a nail product or service applied to the one or more nail objects, in response to the localization by the nail localization engine, to provide a virtual try-on experience (VTO). These and other embodiments will be apparent, including the following numbered embodiments.
[0122] Embodiment 18: In embodiment 17, the method comprises one or both of recommending the nail product or nail service for virtual try-on, and facilitating a purchase of the nail product or nail service via an e-commerce transaction.
[0123] Embodiment 19: In embodiment 17 or 18, the one or more deep neural networks are configured for antialiasing in each encoder stage having a downsampling operation and each decoder block having an upsampling operation.
[0124] Embodiment 20: In embodiment 19, each decoder block having an upsampling operation is configured with an operator of the type: BlurZoom(step=2), where the BlurZoom operator includes an actual or theoretical zero insertion operation to interleave rows and columns around existing values and a low-pass filtering operation to diffuse the existing values into the newly inserted rows and columns to provide any subsequent operation with an offset-equivariant input.
[0125] Embodiment 21: In embodiment 20, the low-pass filtering operation is optimized to reduce multiplications and additions in response to actually or theoretically inserted zero values.
[0126] Embodiment 22: In any one of embodiments 19 to 21, each encoder stage having a downsampling operation is configured with operators of the type: Conv(stride=l) -> BatchNorm -> ReLU -> BlurPool(stride=2), where the BlurPool operator comprises applying a low-pass filter to diffuse each value into neighboring regions of the values.
[0127] Embodiment 23: In any one of embodiments 17 to 22, the method comprises reshaping, via a trained shape model (e.g., a nail shaping engine), the one or more nail objects as located by the nail localization engine; and wherein the rendering component renders the output image in accordance with the reshaping of the one or more objects nails by the nail shaping motor.
[0128] Embodiment 24: In embodiment 23: the input image comprises a current frame of a series of frames of a video and the method comprises stabilizing the reshaping of the one or more nail objects between the series of frames of the video.
[0129] Embodiment 25: In embodiment 23 or embodiment 24: one or both of: the trained shape model reduces the dimensionality of data to identify important features; and the trained model is defined and trained in accordance with principal component analysis techniques. Computer-implemented method embodiments and computer program product embodiments corresponding to one or more of embodiments 18-26 will be apparent. Further, any of the anti-aliasing embodiments 18-26 may be combined with any of the object shaping (contouring) and / or stabilization embodiments 1-17 (e.g., by combining corresponding system, method, or computer program product embodiments.Thus, antialiasing techniques can be used in localization operations in combination with contouring and / or object stabilization techniques.
[0130] In addition to the computing device and method aspects, those skilled in the art will understand that computer program product aspects are disclosed, where instructions are stored in a non-transitory storage device (e.g., memory, CD-ROM, DVD-ROM, disk, etc.), and which, when executed, cause a computing device to perform any of the method aspects stored herein.
[0131] Although the computing devices are described with reference to processors and instructions that, when executed, cause the computing devices to perform operations, it is understood that other types of circuitry than programmable processors may be configured. Hardware components including specifically designed circuitry may be employed, such as, but not limited to, an application-specific integrated circuit (ASIC) or other hardware designed to perform specific functions, which may be more efficient than a general-purpose central processing unit (CPU) programmed using software.Thus, an apparatus aspect herein generally relates to a system or device having circuitry (sometimes referenced as computational circuitry) that is configured to perform certain operations described herein, such as but not limited to those of a method aspect herein, whether the circuitry is configured via programming or via its hardware design.
[0132] The practical implementation may include all or part of the features described herein. These and other aspects, features, and various combinations may be expressed as methods, apparatus, systems, means for performing functions, program products, and other ways, combining the features described herein. A number of embodiments have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the processes and techniques described herein. In addition, other steps may be provided, or steps may be deleted, from the described process, and other components may be added to or removed from the described systems. Accordingly, other embodiments fall within the scope of the following claims.
[0133] Throughout the description and claims of this specification, the terms "include" and "contain" and variations thereof mean "including but not limited to" and are not intended to (and do not) exclude other components, integers, or steps. Throughout this specification, the singular includes the plural, unless the context otherwise requires. In particular, when the indefinite article is used, the specification is to be understood as contemplating plurality as well as singularity, unless the context otherwise requires.
[0134] Any features, characteristics, integers, compounds, moieties, or chemical groups described in conjunction with a particular aspect, embodiment, or example of the invention are to be understood as being applicable to any other aspect, embodiment, or example, unless inconsistent therewith. Any features disclosed herein (including any accompanying claims, abstracts, and drawings), and / or any steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive. The invention is not limited to the details of the preceding examples or embodiments.The invention extends to any new feature, or any new combination, of the features disclosed in this specification (including any accompanying claims, abstracts and drawings) or to any new feature, or any new combination, of the steps of any disclosed method or process.
Claims
Claims
1. A computer-implemented method comprising executing on a processor one or more steps comprising: rendering an effect on an object having an original shape as detected in an input image, wherein the effect is applied to the input image in accordance with an updated shape obtained from a trained shape model that constrains the original shape in response to significant shape features identified by the trained shape model, the rendering being a component of a method for providing an augmented reality experience in which the effect simulates the application of a product to the object, the object being a fingernail or toenail and the effect being a nail effect simulating a nail product, the input image being a current frame of a video stream and the method comprising: - detecting a location of the object in a previous frame of the video stream;- determining the velocity of the object according to the location in the previous frame and a location in the current frame; and - applying an exponential moving average to a contour of the object as constrained where a weight is determined by the velocity.;
2. The method of claim 1, wherein the object is a first object of a plurality of objects including a second object including another fingernail or toenail and wherein the method comprises locating each of the first object and the second object and tracking respective locations of the first object and the second object between frames of a video and wherein the rendering further comprises rendering a first effect on the first object and rendering a second effect, different from the first effect, on the second effect, the first effect applied in accordance with an updated first shape obtained from the trained shape model for the first object and the second effect applied in accordance with an updated second shape obtained from the trained shape model for the second object.
3. The method of claim 1, further comprising: obtaining an outline of the object; using the trained shape model to determine important features of the original shape in response to the outline; constraining the outline in response to the important features bearings; and generating a render mask using the outline as constrained, the render mask being intended to be used to render the effect on the object.
4. The method of claim 1, further comprising standardizing and normalizing the contour before applying the contour to the trained shape model, and, in connection with the contour as constrained by the trained shape model, undoing the normalization before generating the rendering mask.
5. The method of claim 1, wherein one or both of: the trained shape model reduces the dimensionality of data to identify important features in response to the outline of the original shape; and the trained model is defined and trained in accordance with principal component analysis techniques.
6. The method of claim 1, wherein the input image comprises a plurality of objects of a same class and the rendering step is performed for each of the plurality of objects as detected in the input image.
7. The method of claim 1, wherein the input image comprises a current frame of a series of frames of a video and wherein the method comprises, prior to rendering, stabilizing the updated shape in response to a velocity of the object in the series of frames of the video, including the current frame.