Image Processing for Tracking Multiple Targets Using Convolutional Neural Networks
By developing a high-resolution neural network model on mobile devices, combined with the cascading semantic segmentation model architecture, the problem of real-time nail tracking and rendering in video streams is solved, and the effect of real-time tracking and rendering of nail polish on mobile devices is achieved, and the effect of real-time tracking and rendering of nail polish on mobile devices is achieved, with good generalization capabilities.
Patent Information
- Application Number
- CN202080039718.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-30
- Filing Date
- 2020-04-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2040-04-29
AI Technical Summary
The prior art is difficult to accurately locate and identify nails in real time in video streams and provide augmented reality rendering, especially when processing small objects on mobile devices, inadequate generalization capabilities.
A high-resolution neural network model was developed, combining the cascade semantic segmentation model architecture, using deep learning and shallow learning of low-resolution and high-resolution features, end-to-end nail tracking and rendering through convolutional neural networks (CNNs), providing foreground/background and object class segmentation and directional information, and training using specific loss functions.
It realizes real-time tracking of nails and rendering nail polish on mobile devices, can handle small objects, has good generalization capabilities, and is suitable for semantic segmentation and image updates of multiple objects.
Smart Images

Figure CN113924597B_ABST
Abstract
Description
Technical Field
[0001] The following relates to processing images, including video images, using a computing device adapted to a Convolutional Neural Network (CNN), where such a computing device can include a consumer-oriented smartphone or tablet, and more particularly to processing (e.g., semantic segmentation) multiple objects tracked by a CNN, such as nails in a video. Background Art
[0002] The nail tracking problem is to locate and identify nails in real time with pixel accuracy from a video stream. Additionally, rendering techniques need to be supported to adapt to images from the video stream, such as providing augmented reality. It may be necessary to locate and identify objects other than nails in the image, including in the video stream. Summary of the Invention
[0003] An end-to-end solution is proposed for simultaneously tracking nails and rendering nail polish in real time. A brand-new dataset with semantic segmentation and landmark labels is collected. A high-resolution neural network model is developed for mobile devices and trained using the new dataset. In addition to providing semantic segmentation, the model also provides directional information, such as indicating a direction. Post-processing and rendering operations are provided for nail polish trials, which use at least some of the outputs of the model.
[0004] Although described with respect to nails, other objects can be similarly processed for segmentation and for image updates. Such other objects may also be small objects with simple boundaries (e.g., nails, toenails, shoes, cars (passenger vehicles), license plates, or car parts on a car, etc.). The term "small" here is a relative term related to the scale and the size of the entire image. For example, a nail is relatively small compared to the size of a hand captured in an image that includes the nail. A car in a group of cars imaged from a distance is similarly small compared to a group of plums (or other fruits) imaged on a table. The model is well-suited for generalization to classify a set of objects with known counts and clusters (as here, classifying the fingertips of a hand).
[0005] A computing device is provided that includes a processor and a storage device coupled thereto, the storage device storing a CNN and instructions that, when executed by the processor, configure the computing device to: process an image including multiple objects using the CNN, the CNN being configured to semantically segment the multiple objects within the image, the CNN including a cascaded semantic segmentation model architecture having: a first branch of deep learning that provides low-resolution features; and a second branch of shallow learning that provides high-resolution features; wherein the CNN combines corresponding predictions from the first branch and the second branch to output information including foreground / background and object class segmentation.
[0006] The CNN can combine the corresponding predictions from the first branch and the second branch such that the information output from the CNN also includes directional information.
[0007] The first branch can include an encoder-decoder backbone that produces the corresponding predictions of the first branch. The corresponding predictions of the first branch include a combination of the initial predictions produced after the encoder stage of the first branch and the further predictions produced after further processing in the decoder stage of the first branch. A first branch fusion block can be used to combine the initial predictions and the further predictions to produce the corresponding predictions of the first branch for further combination with the corresponding predictions of the second branch.
[0008] The corresponding predictions of the second branch can be produced after the processing in the encoder stage of the second branch and are cascaded with the first branch. A second branch fusion block can be used to combine the corresponding predictions (F1) of the first branch with the corresponding predictions (F2) of the second branch. F1 can include upsampled low-resolution, high-semantic information features and F2 can include high-resolution, low-semantic information features. Thus, the second branch fusion block combines F1 and F2 together to produce high-resolution fusion features F2' in the decoder stage of the second branch. The CNN can use a convolutional classifier applied to the corresponding prediction F1 to generate downsampled class labels. To process F2, the CNN can use multiple output decoder branches to generate foreground / background and object class segmentations as well as directional information.
[0009] The multiple output decoder branches can include: a first output decoder branch having a 1x1 convolutional block and an activation function that produces a foreground / background segmentation; a second output decoder branch having a 1x1 convolutional block and an activation function that produces an object class segmentation; and a third output decoder branch having a 1x1 convolutional block that produces directional information.
[0010] The CNN can be trained using a loss max pooling (LMP) loss function for overcoming pixel-wise class imbalance in semantic segmentation to determine the foreground / background segmentation.
[0011] The CNN can be trained using a negative log likelihood (NLL) loss function to determine the foreground / background and object class segmentations.
[0012] The CNN can be trained using a Huber loss function to determine the directional information.
[0013] Each object can include a base and a tip, and the directional information can include a base-tip direction field.
[0014] The MobileNetV2 encoder-decoder structure can be used to define the first branch, and the encoder structure from the MobileNetV2 encoder-decoder structure can be used to define the second branch. The CNN can initially be trained using training data from ImageNet and then using an object tracking dataset to train on multiple objects labeled with ground truth.
[0015] These instructions can further configure the computing device to perform image processing to generate an updated image from the image using at least some of the output information. To perform the image processing, at least some of foreground / background and object class segmentation and directional information can be used to change the appearance, such as the color of multiple objects.
[0016] The computing device can include a camera and be configured to: present a user interface to receive an appearance selection applied to multiple objects, and receive a selfie video image from the camera to be used as the image; process the selfie video image to generate an updated image using the appearance selection; and display the updated image to simulate augmented reality.
[0017] The computing device can include a smartphone or a tablet.
[0018] The image can include at least a part of a hand with nails, and the multiple objects can include nails. The CNN can be defined to provide a Laplacian pyramid of output information.
[0019] A computing device is provided that includes a processor and a storage device coupled thereto, the storage device storing instructions that, when executed by the processor, configure the computing device to: receive CNN outputs of foreground / background and object class segmentation and directional information for each of multiple objects to be semantically segmented by the CNN, the CNN having processed an image including the multiple objects; and process the image to generate an updated image by drawing a gradient of a selected color on each of the multiple objects segmented according to the foreground / background segmentation (and object class segmentation), the selected color being drawn perpendicular to the corresponding direction of each object indicated by the directional information.
[0020] The computing device can be configured to apply a corresponding specular reflection component to each of the multiple objects on the gradient and blend the results.
[0021] The computing device may be configured to stretch, before drawing, the corresponding regions of each of the plurality of objects identified by foreground / background segmentation to ensure that edges such as its tips are included for drawing. The computing device may be configured to, before drawing, color at least some adjacent regions outside the corresponding regions of each of the stretched plurality of objects using an average color determined from the plurality of objects; and blur the corresponding regions and the adjacent regions of each of the stretched plurality of objects.
[0022] The computing device may be configured to receive a selected color to be used in drawing.
[0023] A computing device is provided that includes a processor and a storage device coupled thereto, the storage device storing a CNN and instructions that, when executed by the processor, configure the computing device to: process an image including a plurality of objects using the CNN, the CNN being configured to semantically segment the plurality of objects within the image, the CNN including a cascaded semantic segmentation model architecture having: a first branch of deep learning that provides low-resolution features; and a second branch of shallow learning that provides high-resolution features; wherein the CNN combines corresponding predictions from the first branch and the second branch to output information including foreground / background and object class segmentation, and wherein the CNN is trained using a loss average polling loss function.
[0024] The image includes a plurality of pixels, and the plurality of objects within the image are represented by a small number of the plurality of pixels. The CNN may combine corresponding predictions from the first branch and the second branch to further output information including object class segmentation, and wherein the CNN is further trained using an NLL loss function. The CNN may combine corresponding predictions from the first branch and the second branch to further output information including directional information, and the CNN may be further trained using a Huber loss function or an L2 loss function.
[0025] The CNN may be defined as a Laplacian pyramid that provides output information.
[0026] A computing device is provided that includes a processor and a storage device coupled thereto, the storage device storing instructions that, when executed by the processor, configure the computing device to: provide a graphical user interface (GUI) for an annotated image dataset to train the CNN, the GUI having an image display portion that displays the corresponding images to be annotated, the display portion being configured to receive an input for outlining (segmenting) the corresponding objects shown in the corresponding images, and receive an input indicating directional information for each of the corresponding images; receive an input of the annotated images; and save the images associated with the annotations to define a dataset.
[0027] A computing device may be configured to provide control to receive input for semantically classifying each of various objects.
[0028] A CNN may be configured to semantically segment multiple objects within an image. The CNN includes a cascaded semantic segmentation model architecture having: a first branch of deep learning that provides low-resolution features; and a second branch of shallow learning that provides high-resolution features. Wherein, the CNN combines corresponding predictions from the first branch and the second branch to output information including foreground / background and object class segmentation.
[0029] A computing device may be configured to have any one of the computing device aspects or features herein. Clearly, corresponding method aspects and features as well as corresponding computer program product aspects and features are provided for each computing device aspect and feature. These and others will be obvious to those of ordinary skill in the art. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a screenshot of a graphical user interface (GUI) according to an example, using which annotation data of a data set is defined.
[0031] Figure 2 is a part of a screenshot of a GUI according to an example, using which annotation data of a data set is defined.
[0032] Figure 3 is an illustration of a CNN for processing an image according to an example.
[0033] Figure 4 and Figure 5 are respectively Figure 3 illustrations of parts of the CNN.
[0034] Figure 6 is a 4×4 image array processed using a CNN according to an example herein, showing foreground and background masks and directional information.
[0035] Figures 7A - 7C is Figure 6 a magnified portion of
[0036] Figure 8 is a 4×4 image array processed using a CNN according to an example herein, showing an example of the application of an object class segmentation mask to each nail.
[0037] Figure 9 is Figure 8 a magnified portion of
[0038] Figure 10 is a flowchart of an operation
[0039] Figure 11 It is an illustration of pseudocode for operation.
[0040] The inventive concept is best described by certain embodiments of the present invention, which are described herein with reference to the accompanying drawings, wherein like reference numerals always refer to like features. It should be understood that the term "invention" as used herein is intended to imply the inventive concept underlying the embodiments described below, and not merely the embodiments themselves. It should also be understood that the general inventive concept of the present invention is not limited to the illustrative embodiments described below, and the following description should be read in such a perspective. More than one inventive concept may be shown and described, and unless otherwise stated, each inventive concept may be independent or combined with one or more other inventive concepts. Detailed implementation
[0041] An end-to-end solution is proposed for simultaneously and real-time tracking the rendering of nails and nail polish. A brand-new dataset with semantic segmentation and landmark labels is collected. A high-resolution neural network model is developed for mobile devices and trained using the new dataset. In addition to providing semantic segmentation, the model also provides directional information, such as indicating direction. Post-processing and rendering operations are provided for nail polish trials, which at least use some outputs of the model.
[0042] Although described with respect to nails, other objects can be similarly processed for segmentation and for image updating. Such other objects may also be small objects with simple boundaries (such as nails, toenails, shoes, cars (passenger cars), license plates, or car parts on a car, etc.). The term "small" here is a relative term related to the scale and the size of the entire image. For example, compared to the size of the hand captured in an image including a nail, the nail is relatively small. A car in a group of cars imaged from a distance is similarly small compared to a group of plums (or other fruits) imaged on a table. The model is well-suited for generalization to classify a set of objects with known counts and clusters (as here, classifying the fingertips of a hand).
[0043] The trained model is deployed on two hardware platforms: iOS TM Via Core ML TM (e.g., implemented by a native application on Apple products, such as iPhone TM supporting such an environment), and via the web browser of Tensorflow.js[1] (a more agnostic platform). The following are trademarks of Apple Inc.: iOS, Core ML, and iPhone. The model and post-processing operations are flexible enough to support the higher-computing native iOS platform as well as the more resource-constrained web platform, with only minor adjustments to the model architecture and no significant negative impact on performance.
[0044] The selected features are as follows:
[0045] · A dataset was created, including 1438 images from photos and videos, and annotated with foreground-background, each finger category, and base-tip orientation field labels.
[0046] · A new neural network structure for semantic segmentation was proposed, which is suitable for running on mobile devices and for precisely segmenting small objects.
[0047] · It has been demonstrated that the loss maximum pooling robustly produces precise segmentation masks for small objects, which results in spatial (or per-pixel) class imbalance.
[0048] · Post-processing operations were developed, which use multiple outputs from the nail tracking model to segment the nails and locate individual nails, as well as to find their 2D orientations.
[0049] · The post-processing (including rendering) operations use these individual nail positions and orientations to render gradients and hide the light-colored distal edges of natural nails.
[0050] 5.1 Related Work
[0051] MobileNetV2 [2] forms the basis of the encoder of the encoder-decoder neural network architecture. This work is based on MobileNetV2 and uses it as the backbone in the cascaded semantic segmentation model architecture. In addition, the model is agnostic to the specific encoder model used, so any existing valid model from the literature [3, 4, 5, 6] can be used as a direct substitute for the encoder, and any future valid model, including hand-designed and automatically discovered models (e.g., through network pruning), can also be used. MobileNetV2 meets the efficiency requirements to enable the storage and execution of the model on smaller or fewer resources available in, for example, smartphones (e.g., with less graphics processing resources than larger computers such as laptops, desktops, gaming computers, etc.).
[0052] The loss maximum pooling (LMP) loss function is based on [7], where the p-norm parameter is fixed at p = 1 because this simplifies the function while maintaining the performance within the standard error range of the optimal p-norm parameter performance according to [7]. Applying LMP to the inherently class-imbalanced task of nail segmentation, the experiments further support the effectiveness of LMP in overcoming pixel-level class imbalance in semantic segmentation.
[0053] The cascaded structure is related to ICNet[8] as the neural network model here combines shallow / high-resolution and deep / low-resolution branches. Different from ICNet, the model is designed to run on mobile devices, so the encoder and decoder are completely redesigned according to this requirement.
[0054] 5.2 Dataset
[0055] Due to the lack of previous work specifically for nail tracking, a brand-new dataset was created for this task. Egocentric data was collected from participants who were asked to take photos or videos of their hands as if they were showing off their nails on social media.
[0056] Dense semantic segmentation labels were created using polygons, which are an easy-to-annotate and precise label type for rigid objects such as nails. Since the model is trained on dense labels, the polygon annotation method can also be replaced by per-pixel annotation. Figure 1 and Figure 2 Shown in is an example of an interface 100 for creating nail annotations through a combination of three label types. Figure 1 Shown is an interface 100 with a part 102 that displays and receives the input of an image to be annotated for the dataset. The interface 100 also includes a part 104 that has multiple controls, such as radio button controls for setting data (e.g., flags). Other controls in part 104 can be used to define polygons and mark landmarks (such as tip landmark 106A and base landmark 106B), etc.
[0057] The interface 100 thus enables:
[0058] 1. A polygon that encloses the nail pixels (i.e., separates the foreground nail from the background).
[0059] 2. A per-polygon class label to identify an individual nail. Each polygon in the dataset represents a nail and is classified into one of ten nail categories, namely "left little finger", "right thumb", etc., see 102 in. Figure 2 in.
[0060] 3. Base and tip landmarks that define the orientation of each polygon. The nail base / tip landmarks are used to generate a dense orientation field that has the same spatial resolution as the input image, and each pixel has a pair of values representing the x and y directions from the base to the tip of the nail to which the pixel belongs.
[0061] The new annotated dataset includes a total of 1438 annotated images, and the participants contributing the images are split into training, validation, and test sets (i.e., the images of each participant belong only to training, validation, or test). The split dataset contains 941, 254, and 243 images for training, validation, and test respectively. In the experiment, the model is trained on the training set and evaluated on the validation set.
[0062] 5.3 Model
[0063] At the core of the nail tracking system (e.g., a computing device configured as described herein) is an encoder-decoder convolutional neural network (CNN) architecture trained to output foreground / background and nail class segmentation as well as directional information (e.g., base-tip direction field). The model architecture is related to ICNet [8], but changes were made to enable the model to run fast enough on mobile devices and produce multi-task outputs. A top-level view of the model architecture is as Figure 3 shown.
[0064] Figure 3 Model 300 is shown processing an input (image) 302 using two branches. The first branch 300A ( Figure 3 the upper branch in ) includes blocks 304 - 324. Figure 3 The second branch 300B (lower part in ) includes blocks 326 - 338. It should be understood that these bright line differentiations can be modified. For example, block 326 can be a block of the first branch 300A. Block 304 is a downsampling ×2 block. Blocks 306 - 320 (also known as stage_low1, stage_low2,... stage-low8) are blocks of an encoder-decoder backbone (with an encoder phase and a decoder phase) described further below. Block 322 is an upsampling ×2 block, and block 324 is a first branch fusion block described further below. Block 326 is also an upsampling X2 block. Blocks 326 - 332 (also known as stage_high1, stage_high2,... stage-high4) are blocks of the encoder phase described further below. The encoder-decoder backbone is based on MobileNetV2 [2]. More details are shown in Table 1. The encoder phase of the second branch (boxes 328 - 332) is also modeled based on the encoder of MobileNetV2 [2].
[0065] The encoder of the model is initialized with the model weights pre-trained by MobileNetV2 [2] on ImageNet [9]. A cascade of two MobileNetV2 encoder backbones with α = 1.0 (i.e., the encoder phase) is used, both pre-trained on 224×224 ImageNet images. The encoder cascade (from each branch) consists of a shallow network (stage_high1...4) with a high-resolution input and a deep network (stage_low1...8) with a low-resolution input, both of which are prefixes of the complete MobileNetV2. For the low-resolution encoder of the first branch at stage 6, the stride is changed from 2 to 1, and to compensate for this change, dilated 2× convolutions are used in stages 7 and 8. Thus, the output stride of the low-resolution encoder is 16× with respect to its input, rather than 32× in the original MobileNetV2. A detailed layer-by-layer description is shown in Table 1. Table 1 shows a detailed overview of the nail segmentation model architecture. Each layer name corresponds to the blocks in Figure 3 and Figure 4 as described herein. The height H and width W refer to the full-resolution H×W input size. For the projection 408 and dilation layer 410, p ∈ {16, 8}. For stages stage3_low to stage7_low, the number of channels in parentheses is used for the first layer of the stage (not shown), which increases to the non-parenthesized number for the subsequent layers in the same stage.
[0066]
[0067] Table 1
[0068] The decoder of model 300 is shown in the middle and bottom right of Figure 3 (e.g., blocks 324 and 336 (including the fusion block) and upsampling blocks 322 and 326), and a detailed view of the decoder fusion model for each of blocks 324 and 336 is shown in Figure 4 . For the original input of size H×W, the decoder fuses the features from stage_low4 (from block 312) with the upsampled features from block 322 derived from stage_low8, then upsamples (block 326), and fuses the resulting features with the features of stage_high4 via the fusion block 336 (block 334).
[0069] Figure 4Shown is a fusion module 400 that uses blocks 408, 410, 412, and adder 414 in a decoder to fuse upsampled low-resolution, high-semantic information features represented by feature map F1(402) with high-resolution, low-semantic information features represented by feature map F2(404) to produce high-resolution fused features represented by feature map F2′(406). Regarding block 324, feature map F1(402) is output from block 322, and feature map F2(404) is output from block 312. At 326, the feature map F2′(406) from block 324 is upsampled to be provided as feature map F1(402) in the block instance of model 400 to block 336. In block 336, the feature map F2(404) received from block 334 is output, and the feature map F2′(406) is provided as an output to block 338. Block 338 upsamples the input resolution / 4 and then provides the resulting feature map to decoder model 340. Decoder model 340 is shown in Figure 5 . Decoder model 340 produces three types of information of the image (e.g., 3-channel output 342), as further described with respect to Figure 5 .
[0070] As Figure 4 shown, a 1×1 convolutional classifier 412 is applied to the upsampled F1 features, which is used to predict the downsampled labels. Similar to
[10] , this output's "Laplacian pyramid" optimizes higher-resolution, smaller receptive field feature maps to focus on improving predictions from low-resolution, larger receptive field feature maps. Thus, in model 400, the feature map (not shown) from block 412 is not itself used as an output. Instead, during training, the loss function is applied in the form of pyramid output regularization (i.e., the loss applied in Figure 5 ).
[0071] Box 342 represents a global output from the decoder, which includes three channels corresponding to the outputs of the blocks of the three branches 502, 504, and 506 from Figure 5 . The first channel includes per-pixel classification (e.g., foreground / background mask or object segmentation mask), the second channel includes classifying the segmentation mask into individual fingertip classes, and the third channel includes a field of 2D directional vectors per segmentation mask pixel (e.g., per pixel (x, y)).
[0072] As Figure 5As shown, the decoder uses multiple output decoder branches 502, 504, and 506 to provide the directional information (e.g., the vector from the base to the tip in the third channel) required for rendering on the nail tip, as well as the nail class prediction (in the second channel) required to find the nail instance using connected components. These additional decoders are trained to produce dense predictions that are only discharged in the annotated nail regions of the image at a disadvantage. Each branch employs a corresponding loss function according to this example. While the normalized exponential function (softmax) is shown in branches 502 and 504, another activation function for segmentation / classification can be used. It should be understood that the dimensions here are representative and can be adapted to different tasks. For example, in Figure 5 branches 502 and 504 involve 10 classes and determine the dimensions accordingly.
[0073] The binary (i.e., nail vs. background) prediction, together with the direction field prediction, is visualized in Figure 6 That is, Figure 6 shows a 4×4 array 600 of updated images generated from the processed input image. The foreground / background mask is used to identify the corresponding nails for coloring. The nail regions are colored pixel by pixel (although shown in grayscale here) to show consistency with the ground truth as well as false positive and false negative identifications in the foreground / background mask. The updated images of the array 600 also show the directional information. Figure 6 Figures 6A, 6B, and 6C show magnified images 602, 604, and 606 from the array 600, with annotations where the white arrows point to false positive regions and the black arrows point to false negative regions. In image 604, a common failure mode is shown where an invisible hand pose results in over-segmentation. In image 606, an example of under-segmentation due to an invisible lighting / nail color combination is shown. It is expected that both of these failure cases can be improved by adding relevant training data.
[0074] The individual class prediction for each hand / finger combination (e.g., left little finger) is only visualized in the 4×4 array 800 in Figure 8 That is, Figure 9 shows a magnified image 802 with an annotation (white arrow 900) indicating that one class (ring finger) leaks into another class (middle finger). The reason for the class leak is that the nails overlap due to the camera's perspective. This can be improved by dense CRF or guided filter post-processing.
[0075] 5.4 Inference (Training Details)
[0076] The neural network model is trained using PyTorch
[11] . The trained model is deployed to iOS using Core ML and to web browsers using TensorFlow.js [1].
[0077] Data augmentation includes contrast normalization, frequency noise alpha blending augmentation, and random scale, aspect ratio, rotation, and cropping augmentations. Contrast normalization adjusts the contrast by scaling each pixel value I ij to 127 + α(I ij - 127), where α ∈ [0.5, 2.0]. Frequency noise alpha blending mixes two image sources using a frequency noise mask. There is uniform random sampling ratio magnification starting from [1 / 2, 2], aspect ratio stretching magnification starting from [2 / 3, 3 / 2], rotation magnification starting from ±180°, and a square image with side length 14 / 15 is randomly cropped from the shorter side length of the given downsampled training image.
[0078] Considering the current software implementations, namely Core ML and TensorFlow.js, and the current mobile device hardware, the system can run in real-time (i.e., at 10 FPS) at all resolutions of 640×480 (local mobile) and 480×360 (web mobile). The model is trained on input resolutions of 448×448 and 336×336 respectively. All input images are normalized by the mean and standard deviation of the ImageNet dataset. The MobileNetV2 encoder backbone is pre-trained on ImageNet for 400 epochs using SGD with Nestrov momentum of 0.9, and the initial learning rate of 10^(-2) is reduced by 10 times at 200 and 300 epochs.
[0079] The encoder-decoder model is trained 400 times on the nail tracking dataset. To preserve the pre-trained weight values, for all pre-trained layers, i.e., stage_high1..4 and stage_low1..8, a lower initial learning rate of 5×10 -3 is used, while for all other layers, an initial learning rate of 5x10 -2 is used. After the previous work
[12] , according to a polynomial decay learning rate schedule is used, where l_t is the learning rate at iteration t and T is the total number of steps. The batch size used is 32. The optimizer is SGD with Nestrov momentum of 0.99 and model weight decay of 10 -4 . There is clipped gradient at 1.0. The LMP loss function calculates the loss as the average loss of the 10% pixels with the highest loss values.
[0080] 5.5 Discussion of the objective function
[0081] To handle the class imbalance between the background (high-representative class) and the nails (low-representative class), in the objective function, by sorting the loss magnitudes per pixel, max-loss pooling [7] is used for all pixels in the mini-batch, and the average of the top 10% of the pixels is taken as the mini-batch loss. It was found that compared to the baseline where the nail class was weighted only 20× higher than the background, the gain using max-loss pooling was ≈2% mIoU, where the improvement in mIoU was reflected in the sharper appearance of the nail edges along the class boundaries (the original baseline always over-segmented).
[0082] Three loss functions corresponding to the three outputs of the model as shown in Figure 5 were used. Both the nail class and the foreground / background predictions minimize the negative log-likelihood of the multinomial distribution given in Equation 1, where c is the ground truth class, is the pre-softmax prediction of the model for class c, and is the loss of the pixel at (x, y) = (i, j).
[0083]
[0084] For class prediction, c ∈ {1, 2,.., 10}, while for foreground / background prediction, c ∈ {1, 2}. LMP was only used for foreground / background prediction; since the nail class predictions are only valid within the nail regions, these classes are balanced and do not require LMP.
[0085]
[0086] In Equation 2, the threshold τ is the loss value of the highest loss pixel. The [·] operator is the indicator function.
[0087] For the orientation field output, for each pixel within the ground truth nails, the Huber loss was applied to the tip orientation of the nail on a normalized basis. This was to de-emphasize the field loss when it was approximately correct, since all that was required for rendering was the approximate correctness of the base-tip orientation, which prevented the orientation field loss from detracting from the binary and class nail segmentation losses. Other loss functions (such as L2 and L1 errors) could also be used in the system instead of the Huber loss.
[0088]
[0089] In Equation 3, the indices (i, j) cover all spatial pixel positions, while k ∈ {0, 1} indexes the (x, y) directions of the base-tip orientation vector. Additionally, each scalar field prediction is normalized such that the vector is a unit vector, i.e., The field direction labels are also normalized such that For the orientation field and the fingernail - like loss, there is no class - imbalance problem, so they are just the mean of their respective losses, i.e., and where N class = H×W and N field = 2×H×W. The overall loss is l = l fgbg + l class + l field .
[0090] 5.6 Post - processing and rendering
[0091] The output from the model can be used to process the input image and generate and update images. In Method 1 (see Figure 10 ), a post - processing and rendering method is described that uses the output of the tracking prediction of a CNN model to paint realistic nail polish on a user's fingernails. This method uses the single - fingernail position and directional information predicted by the fingernail tracking module (using a CNN model) to render the gradient and hide the light - colored distal edge of the natural fingernail.
[0092] Figure 10 Operation 1000 of a computing device is shown. The computing device includes a CNN model as shown and described here and instructions that configure the computing device. Operation 1000 shows the computing device presenting a user interface (e.g., a GUI) at step 1002 to receive an appearance selection to be applied to multiple objects (e.g., fingernails). At 1004, the operation receives, for example, a source image from the camera of the computing device. The source image can be a self - portrait still image or a self - portrait video image as the image to be processed. At 1006, the instructions configure the computing device to process the image to identify multiple objects, at 1008 process the image to apply the appearance selection, and at 1010 generate an updated image showing the applied appearance selection. There may be an updated image (at 1012) to simulate augmented reality.
[0093] Figure 11 "Method 1" including pseudocode 1100 is shown, for operations that can be used after processing by a CNN using the output from the CNN. Method 1 shows post - processing and nail - polish rendering operations. These operations first paint a gradient of the user - selected color on each fingernail perpendicular to the fingernail direction and covered by the fingernail cap. Then, it copies the specular reflection components from the original fingernail and blends them on top of the gradient.
[0094] 6 Miscellaneous
[0095] It can be understood that pre - processing, such as generating an input of the required size, centering the required part of the image, correcting the illumination, etc., can be used before the model processes.
[0096] Although described with respect to fingernails, one of ordinary skill in the art may follow other objects as described and make modifications to the teachings herein. Although color appearance effects are described as being applied to produce updated images, other appearance effects may be used.
[0097] Appearance effects may be applied at or near the location of the object being tracked. In addition to aspects of computing devices, one of ordinary skill in the art will understand that aspects of a computer program product are disclosed, where instructions are stored in a non-transitory storage device (e.g., memory, CD-ROM, DVD-ROM, RAM, tape, disk, etc.) and executed by a processor to configure the computing device to perform any method aspects stored herein. The processor may be a CPU, GPU, or other programmable device or a combination of one or more of any such devices. As described herein, Core ML from Apple's iOS-based iPhone products is used to prepare an implementation.
[0098] Actual implementations may include any or all of the features described herein. These and other aspects, features, and various combinations may be expressed as methods, apparatus, systems, means for performing functions, program products, and otherwise combining the features described herein. Multiple embodiments have been described. However, it should be understood that various modifications may be made without departing from the spirit and scope of the methods and techniques described herein. Additionally, other steps may be provided from the described processes, or steps may be eliminated, and other components may be added to the described systems or other components may be removed from the described systems. Accordingly, other embodiments are within the scope of the appended claims.
[0099] In the description and claims of this specification, the words "comprise" and "contain" mean "including but not limited to", and they are not intended to (nor do they) exclude other components, integers, or steps. In this specification, unless the context otherwise requires, the singular includes the plural. In particular, in the case of using an indefinite article, unless the context otherwise requires, the specification should be understood to contemplate both the plural and the singular. For example, the term "and / or" with respect to "A and / or B" herein means one of A and B and both A and B.
[0100] Features, integer features, compounds, chemical moieties or groups described in conjunction with a particular aspect, embodiment or example of the invention are to be understood to be applicable to any other aspect, embodiment or example unless incompatible therewith. All features disclosed herein (including any accompanying claims, abstract and drawings) and / or all steps of any method or process so disclosed may be combined in any combination, except combinations in which at least some of such features and / or steps are mutually exclusive. The invention is not limited to the details of any foregoing examples or embodiments. The invention extends to any new one or any new combination of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any new one or any new combination of the steps of any method or process disclosed.
[0101] 7 Conclusion
[0102] A model for nail tracking and nail polish rendering operations is presented. Using current software and hardware, a user computing device (e.g., a smartphone or tablet) can be configured to run in real-time on iOS and web platforms. The use of LMPs combined with a cascaded model architecture design simultaneously enables pixel-accurate nail predictions at resolutions up to 640×480. The proposed post-processing operations leverage multiple output predictions of the model to render gradients on a single nail and hide light-colored distal edges when rendering on top of the natural nail by stretching the nail cap in the direction of the nail tip.
[0103] References
[0104] Each of the references [1] to
[13] listed below is incorporated herein by reference:
[0105] [1] Daniel Smilkov, Nikhil Thorat, Yannick Assogba, Ann Yuan, Nick Kreeger, Ping Yu, Kangyi Zhang, Shanqing Cai, Eric Nielsen, David Soergel, Stan Bileschi, Michael Terry, Charles Nicholson, Sandeep N. Gupta, Sarah Sirajuddin, D. Sculley, Rajat Monga, Greg Corrado, Fernanda B. Viégas, and MartinWattenberg.Tensorflow.js:Machine learning for the web and beyond.arXivpreprint arXiv:1901.05350,2019.
[0106] [2]Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018.
[0107] [3]Xiangyu Zhang, Xinyu Zhou, Mengxiao Lin, and Jian Sun. Shufflenet: An extremely efficient convolutional neural network for mobile devices. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
[0108] [4]Robert J Wang, Xiang Li, and Charles X Ling. Pelee: A real-time object detection system on mobile devices. In Advances in Neural Information Processing Systems 31, 2018.
[0109] [5]Forrest N. Iandola, Song Han, Matthew W. Moskewicz, Khalid Ashraf, William J. Dally, and Kurt Keutzer. Squeezenet: Alexnet-level accuracy with 50x fewer parameters and <0.5mb model size. arXiv:1602.07360, 2016.
[0110] [6]Barret Zoph,Vijay Vasudevan,Jonathon Shlens,and Quoc V.Le.Learning transferable architectures for scalable image recognition.In The IEEE Conference on Computer Vision and Pattern Recognition(CVPR),2018.
[0111] [7]Samuel Rota Bulò,Gerhard Neuhold,and Peter Kontschieder.Loss max-pooling for semantic image segmentation.In The IEEE Conference on Computer Vision and Pattern Recognition(CVPR),2017.
[0112] [8]Hengshuang Zhao,Xiaojuan Qi,Xiaoyong Shen,Jianping Shi,and Jiaya Jia.Icnet for realtime semantic segmentation on high-resolution images.In ECCV,2018.
[0113] [9]J.Deng,W.Dong,R.Socher,L.-J.Li,K.Li,and L.Fei-Fei.ImageNet:A Large-Scale Hierarchical Image Database.In The IEEE Conference on Computer Vision and Pattern Recognition(CVPR),2009.
[0114]
[10] Golnaz Ghiasi and Charless C.Fowlkes.Laplacian reconstruction and refinement for semantic segmentation.In ECCV,2016.
[0115]
[11] Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. Automatic differentiation in pytorch. In NIPS-W, 2017.
[0116]
[12] Liang-Chieh Chen, George Papandreou, Iasonas Kokkinos, Kevin Murphy, and Alan L. Yuille. Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. arXiv preprint arXiv:1606.00915, 2016.
[0117]
[13] C. Grana, D. Borghesani, and R. Cucchiara. Optimized block-based connected components labeling with decision trees. IEEE Transactions on Image Processing, 2010.
Claims
1. A computing device, comprising a processor and a storage device coupled thereto, the storage device storing a convolutional neural network (CNN) and instructions that, when executed by the processor, configure the computing device to: Process an image including a plurality of objects using the CNN, the CNN being configured to semantically segment the plurality of objects within the image, the CNN including a cascaded semantic segmentation model architecture having: A first branch of deep learning that provides low-resolution features; and A second branch of shallow learning that provides high-resolution features; Among them, The CNN combines corresponding predictions from the first branch and the second branch to output each of the following: a first channel including output information of foreground / background, a second channel including output information of object class segmentation for each of the plurality of objects, and a third channel including output information of directional information for each object, the directional information including a two-dimensional directional vector field of each object from a first end to a second end.
2. The computing device according to claim 1, wherein, The plurality of objects includes nails.
3. The computing device according to claim 1 or claim 2, wherein, The first branch includes an encoder-decoder backbone that generates the corresponding prediction of the first branch.
4. The computing device according to any one of claims 1 to 3, wherein, The corresponding prediction of the second branch is generated after processing in the encoder stage of the second branch and is cascaded with the first branch.
5. The computing device according to claim 4, wherein, F1 includes upsampled low-resolution, high-semantic information features, F2 includes high-resolution, low-semantic information features, and wherein the second branch fusion block combines F1 and F2 to generate high-resolution fusion features F2' in the decoder stage of the second branch.
6. The computing device according to claim 5, wherein, To process F2, the CNN uses a plurality of output decoder branches to generate the foreground / background and object class segmentation as well as the directional information.
7. The computing device according to any one of claims 1 to 3, wherein, The CNN is trained using a loss max pooling (LMP) loss function for overcoming per-pixel class imbalance in semantic segmentation to determine the foreground / background segmentation.
8. The computing device according to any one of claims 1 to 3, wherein, The CNN is trained using a Huber loss function to determine the directional information.
9. The computing device according to any one of claims 1 to 3, wherein, The first end includes a base, and the second end includes a tip, and the directional information includes a base-tip direction field.
10. The computing device according to any one of claims 1 to 3, wherein, The first branch is defined using a MobileNetV2 encoder-decoder structure, and the second branch is defined using the encoder structure from the MobileNetV2 encoder-decoder structure; And wherein the CNN is initially trained using training data from ImageNet and then trained using an object tracking dataset for the plurality of objects labeled with ground truth.
11. The computing device according to claim 10, wherein, To perform image processing, at least some of the foreground / background and object class segmentation as well as the directional information are used to change the appearance.
12. An image processing method, comprising: Processing an image including a plurality of objects using a convolutional neural network (CNN), the CNN being configured to semantically segment the plurality of objects within the image, the CNN including a cascaded semantic segmentation model architecture having: A first branch of deep learning that provides low-resolution features; and A second branch of shallow learning that provides high-resolution features; wherein the CNN combines corresponding predictions from the first branch and the second branch to output each of the following: a first channel including output information of foreground / background, a second channel including output information of object class segmentation for each of the plurality of objects, and a third channel including output information of directional information for each object, the directional information including a two-dimensional directional vector field of each object from a first end to a second end.
13. The method according to claim 12, wherein, The plurality of objects includes nails.
14. The method according to claim 13, comprising performing image processing to generate an updated image from the image using at least some of the information output from the CNN; and wherein, Performing image processing uses at least some of the foreground / background, object class segmentation, and directional information to change the appearance.
15. The method according to claim 14, wherein, To change the appearance, realistic nail polish is painted on the plurality of objects to render a gradient and hide the light-colored distal edges of the natural nails.
16. The method according to claim 14 or 15, comprising: Presenting a user interface to receive an appearance selection applied to a plurality of objects; Receiving a selfie video image from a camera to be used as the image; Processing the selfie video image to generate the updated image using the appearance selection; And displaying the updated image to simulate augmented reality.