SYSTEMS AND METHODS IN THE FIELD OF SELF-SUPERVISED DETECTION OF FACIAL FLAGSHIPS

The self-supervised detection framework with CARB and MIM network efficiently refines selective local correspondences, addressing inefficiencies in existing methods, achieving significant performance gains in facial landmark detection.

FR3163471A3Active Publication Date: 2025-12-19LOREAL SA
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
FR2024006398
Authority / Receiving Office
FR · FR
Patent Type
Utility models
Current Assignee / Owner
Filing Date
2024-06-17
Publication Date
2025-12-19
Estimated Expiration
2034-06-17

AI Technical Summary

Technical Problem

Existing self-supervised learning methods for facial landmark detection are inefficient due to their reliance on memory-intensive hypercolumns and the need to establish correspondences between all pairs of spatial features, neglecting the importance of selective matching and refining only critical local correspondences.

Method used

A self-supervised detection framework that employs a Correspondence Approximation and Refinement Block (CARB) using a Masked Image Modeling (MIM) network with a Correspondence Approximation and Refinement Block (CARB) to selectively refine only selected local matches, leveraging a locality-constrained loss to optimize network parameters.

Benefits of technology

The framework achieves superior performance in landmark matching and detection tasks, outperforming existing methods by 20% to 44% on landmark matching and 9% to 15% on landmark detection, while reducing computational and memory costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000037_0000
    Figure 00000037_0000
  • Figure 00000037_0001
    Figure 00000037_0001
  • Figure 00000037_0002
    Figure 00000037_0002
Patent Text Reader

Abstract

SYSTEMS AND METHODS IN THE FIELD OF SELF-SUPERVISED DETECTION OF FACIAL BOUNDARIES. Systems and methods for self-supervised learning (SSL) in facial detection networks are proposed. In one embodiment, a facial detection network comprises encoder components configured to encode facial features, the encoder components including trained components of a masked image modeling (MIM) network configured to process non-overlapping patches determined from the input image, the MIM network trained with an SSL objective; and decoder components configured by training to determine local matches between features to determine estimates for facial landmarks. In one embodiment, the MIM network is an MAE network.In one embodiment, the decoder components are derived from those of a second trained network comprising the encoder components as trained but fixed, wherein the decoder components of the second network are trained using locality constraint repulsion loss (LCR). Methods are proposed for SSL training of the encoder and decoder components. Figure for abstract: none.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: SYSTEMS AND METHODS IN THE FIELD OF SELF-SUPERVISED DETECTION OF FACIAL BOUNDARIES

[0001] This disclosure relates to computer vision and image processing using deep neural networks and more particularly to systems and methods for self-supervised detection of facial landmarks. CONTEXT

[0002] Self-supervised estimation of landmarks is a difficult task that requires the training of locally distinct feature representations to identify sparse facial landmarks in the absence of annotated data. To solve this task, existing state-of-the-art methods (EDA) (1) extract coarse features from ridges that are trained with instance-level Self-Supervised Learning (SSL) paradigms, which neglect the dense prediction nature of the task, (2) aggregate them into memory-intensive hypercolumn formations, and (3) supervise lightweight projector arrays to naively establish complete local correspondences between all pairs of spatial features. SUMMARY

[0003] It is proposed (for example, in embodiments) to provide systems and methods for self-supervised detection of facial landmarks that leverage a regional SSL method, operate on a vanilla feature map instead of expensive hypercolumns, and employ a Correspondence Approximation and Refinement Block (CARB) that exploits a simple density peak aggregation algorithm and the proposed locality-constrained loss of repulsion to directly refine only selected local matches. Extensive experiments have demonstrated that such a framework is highly efficient and robust, outperforming existing EDA methods by large margins of ~20% to 44% on landmark matching and ~9% to 15% on landmark detection tasks.Given that multiple new features are provided, it will emerge that not all embodiments can incorporate each of the new features (for example, an embodiment may incorporate only one of the new features).

[0004] Self-supervised learning (SSL) systems and methods are proposed for facial recognition networks. In one embodiment, a facial recognition network includes encoder components configured to encode Facial features are analyzed using encoder components, including trained components from a Masked Image Modeling (MIM) network configured to process non-overlapping patches derived from the input image. The MIM network is trained with an SSL objective. Decoder components are configured by training to determine local matches between features to generate estimates for facial landmarks. In one embodiment, the MIM network is an MAE network. In another embodiment, the decoder components are derived from those of a second trained network comprising the encoder components as trained but fixed, where the decoder components of the second network are trained using locality-constrained repellence (LCR). Methods are proposed for SSL training of the encoder and decoder components. BRIEF DESCRIPTION

[0005] [Fig.1A] Fig.1A shows a self-supervised detection framework for facial landmarks (for example, as a plurality of stages with representative input and output) in accordance with the prior art.

[0006] [Fig.1B] Fig.1B shows a self-supervised detection framework for facial landmarks (e.g., in the form of a plurality of stages (e.g., in the form of system components (e.g., in the form of software structures)) with representative input and output) according to one embodiment of the present.

[0007] [Fig.2] Fig.2 shows an embodiment of an automatic detection frame supervised by facial features similar to [Fig.1B].

[0008] [Fig.3] Fig.3 shows t-SNE tracings of landmark representations for each of the two prior art frames and an embodiment of one frame disclosed herein.

[0009] [Fig.4] [Fig.4] is a schematic diagram of the operating principle of a single effects pipeline threaded in accordance with a prior art embodiment.

[0010] [Fig.5A] The [Fig.5A] is an illustration showing a representative frame t (an example of a face image) including a face and a background.

[0011] [Fig.5B] The [Fig.5B] is an output illustration of a face tracking component of an effects pipeline showing a cropped face image and face point groups according to one embodiment.

[0012] [Fig.6] Fig.6 is an illustration of a computer environment, in accordance with an embodiment, such as for performing a virtual fitting. DETAILED DESCRIPTION

[0013] Facial landmark detection is a computer vision task involving the identification and localization of specific key points corresponding to particular positions on a human face. Facial landmarks form the core many classic downstream tasks such as 3D facial reconstruction, facial recognition, facial emotion / expression recognition and more contemporary applications such as facial beauty prediction and face makeup try-on or virtual try-on (“VTO” for Virtual Try On).

[0014] Although extremely useful, training facial landmark detectors requires numerous accurate annotations per sample, making it a laborious and expensive endeavor. Furthermore, landmarks are not always well-defined semantically, which makes their annotations prone to inconsistencies and can severely limit the development of accurate landmark models. Motivated to avoid these drawbacks, recent work has incorporated unsupervised and self-supervised learning (SSL) paradigms into its methods. Pre-trained SSL models have been shown to produce highly effective feature representations without the use of labeled data and often outperform their supervised counterparts on target tasks.

[0015] Facial landmark detection and matching tasks rely on the formation of locally distinct features to differentiate (1) facial regions (e.g., eye vs. lip), (2) components of facial parts (e.g., left vs. right corners of the lip), and (3) specific pixels of each landmark. In configurations where annotations are severely limited, some recent methods follow a two-stage training protocol. In the first stage, the backbone is trained with a typical SSL objective. In the second stage, the backbone is fixed, and a separate array of light projectors is trained to encode local matches, i.e., the relationships between different regions within the same image.

[0016] Previous work has adopted multi-view SSL protocols, which may be less efficient on bitterness estimation tasks due to several factors. First, these augmentation and comparison pretext tasks incentivize the network to learn category-specific signals, but the task framework here operates only on a single category, namely the human face. Second, contrastive learning requires a large set of diverse negative samples to avoid crashing. Finally, the training objectives may not directly encourage the model to learn the complex facial cues in positive face samples to differentiate facial regions, which are necessary for dense tasks such as bitterness detection and matching.

[0017] On the other hand, the Masked Image Modeling (MIM) protocol requires the network to reconstruct masked regions from a limited context. For example, for an input image, 75% of the image is masked, leaving patches representing 25% with which to reconstruct the input image.

[0018] In accordance with a teaching herein, based on an observation that non-flavored regions (e.g., the cheeks and forehead) are larger and more uniform than sparse and distinctive flavored regions (e.g., the eyes and corners of the lips), it is hypothesized, without the Applicant being bound by such a hypothesis, that the reconstruction of masked flavored regions leads to the formation of effective representations of facial flavors. In one embodiment, the masked autoencoder (MAE) as described in He, Kaiming, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollar, and Ross B. Girshick, “Masked Autoencoders are Scalable Vision Learners.” » 2022 1EEE / CVE Conference on Computer Vision and Pattern Recognition (CVPR) (2021): 15979-15988 (which is incorporated here in its entirety by reference), is adopted as the backbone network in the first stage of the framework.It would be appreciated if another network following a MIM protocol could be used as a first-tier backbone network.

[0019] In an MAE-based network as described in He et al., an encoder maps the observed signal to a latent representation, and a decoder reconstructs the original signal from the latent representation. An asymmetric design can be employed. The asymmetric design allows the encoder to operate only on the observed partial signal (e.g., without mask tokens), and a lightweight decoder is used for the complete reconstruction of the signal from the latent representation and the mask tokens. Following the vision transformer approach (e.g., a Vision Transformer (ViT) using self-attention mechanisms to process images), an input image is divided into non-overlapping patches. The MAE samples the patches to determine which patches to mask (or not mask, as appropriate). The MAE encoder of He et al.incorporates patches by linear projection with added positional embeddings. The result is processed by a series of transformer blocks. Only unmasked patches are processed by the encoder; masked patches are discarded. No mask tokens are used in the encoder. A complete set of tokens (i.e., the encoded visible patches supplemented by mask tokens) are the inputs to the decoder. A shared, learned vector indicating the presence of a missing patch to be predicted defines each mask token. Positional embeddings are added to all tokens in this complete set, providing location information to the masking tokens, for example. The decoder includes its respective set of transformer blocks. The MAE decoder of He et al.can only be used during initial training (e.g., pre-training) to perform the image reconstruction task in order to obtain a task-trained encoder. By. For example, only the encoder is used to produce image representations for recognition purposes. The decoder architecture can be designed independently. Smaller decoders (e.g., transformer blocks) that are narrower and shallower than the encoder can be used, ensuring asymmetry. Therefore, and in accordance with He et al., a reduced set of inputs is processed by the encoder, and the full set of tokens is processed by a lightweight decoder, reducing training time.

[0020] For the second stage, both CL (Cheng, Zezhou, Jong-Chyi Su and Subhransu Maji. “On equivariant and invariant learning of object landmark representations.” In Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 9897-9906. 2021, incorporated herein in full by reference) and LEAD (Karmali, Tejan, Abhinav Atrishi, Sai Sree Harsha, Susmit Agrawal, Varun Jampani and R. Venkatesh Babu. “Lead: Self-supervised landmark estimation by aligning distributions of feature similarity.” In Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision, pp. 623-632. 2022, incorporated herein in full by reference) use objectives to establish correspondences between each pair of feature descriptors within the same image.Based on the earlier observation that non-referential regions are larger and more uniform, the question arose: is it necessary to establish correspondences between all pairs of feature descriptors? It is further hypothesized, though not limited to such a hypothesis, that selectively refining important correspondences makes more efficient use of the network parameters. To this end, in one embodiment, a new Correspondence Approximation and Refinement Block (CARB) is employed in the second stage. First, the MAE output (first-stage MIM output) is differentiated into attentive tokens (markers and important facial regions) and inattentive tokens (insignificant facial regions or background) using first-stage matching signals.Next, in one embodiment, an aggregation algorithm operates on inattentive tokens and approximates member tokens using the aggregation center. Finally, in another embodiment, a lightweight projector network is supervised using a novel locality-constrained repulsion loss (LCR) that penalizes strong erroneous matches between different token types weighted by spatial proximity. Here, only selected matches are directly refined since the loss only operates on attentive tokens and inattentive aggregation center proxy servers.

[0021] Figure 1A shows, at a high level, an SSL 100A frame in accordance with the prior art. The frame shows components of a system comprising a computer device (or more than one such device) configured as if by software. The software includes instructions stored on a non-transient storage medium (e.g., computer device memory) which, when executed (e.g., by at least one processor in the system), cause the system (e.g., its computer device) to perform operations of a method. Figure 1B shows, at a high level, an SSL 100B framework, according to one embodiment of a proposed Selective Correspondence Enhancement (SCE) framework with MAE (SCE-MAE). In each case, an example query 102A and an unannotated dataset 102B of facial images (e.g., in different poses) are provided as input.

[0022] It will be appreciated if each framework shown or described herein can be implemented using respective components in a system comprising a computing device (or more than one such device) configured using respective software. The software comprises instructions stored on a non-transient storage medium (for example, memory of the computing device) which, when executed by at least one processor of the computing device, cause the computing device (for example, the system) to perform operations of a respective method.

[0023] Stage 1 104A of frame 100A, in accordance with the prior art, uses instance-level multiview SSL paradigms that produce less distinct initial local features as output. For stage 1 102A, a query image is transformed (e.g., randomly) and processed via respective networks (e.g., 106) using contrastive loss of learning and similarity 108.

[0024] Stage 1 104B of the SCE-MAE 100B frame exploits MAEs to naturally form better initial features that result in well-defined boundaries between facial landmarks. The query image 102A is randomly masked into patches and processed with a ViT-based transformer block (e.g., an encoder 110) and a transformer block decoder 112, which is discarded after a first-stage training, the decoder being trained to regenerate the output query image.

[0025] Representative t-SNE plots (114A and 114B) illustrate the differences in boundaries (these data plots are representative only and are not necessarily true for a left eye, a right eye, a nose, a corner of the left lip and a corner of the right lip. Plots 114A and 114B are shown in greyscale for ease of patent reproduction but could be in colour to better distinguish one type of bitter from another).

[0026] Stage 2 116A according to the prior art operates on memory-intensive hypercolumns and monitors each pair of features to obtain a complete match. Stage 2 116B of the SCE- frame embodiment MAE 100B employs a matching approximation and refinement block (CARB) that operates on the original MAE output and directly refines only the selected matching pairs. For the example query, SCE-MAE outputs a more focused and sharper similarity map, demonstrating the superiority of the final features. The representative outputs 118A and 118B show differences in the query results for facial landmarks (e.g., a nasal region) for each of the 100A and 100B frames.

[0027] Related work for self-supervised learning (SSL). By solving unique pretext tasks, SSL methods are able to learn representations of discriminating features from unlabeled data. Previous work has explored pretext tasks such as rotation angle prediction and retrieving the original image from randomly permuted patches. Recently, SSL methods based on invariant and contrastive learning have gained popularity due to their ability to capture high-level semantic concepts from data. Invariant learning aims to learn invariant transformation features by forcing the representations of two randomly augmented views of the same image to be similar. Contrastive learning defines different views of an anchor image as positive and views of different images as negative.The goal here is to combine the anchor and positive representations while separating the anchor and negative representations. These methods operate at the encoded image or instance level and can be categorized as SSL augmentation and comparison methods.

[0028] The masked image modeling (MIM) protocol has seen considerable growth. These methods operate at the regional level and learn to recover masked regions from contextual information contained in unmasked patches. It has been empirically shown that by using non-extreme masking ratios or patch sizes in masked autoencoders (MAEs), representation abstractions capture robust high-level information, while extreme masking ratios capture more low-level information. With higher-than-normal masking ratios, the MAE performs dense reconstruction, making them inherently suitable for dense prediction tasks.

[0029] For the first stage of self-supervised face landmark detectors, others have used pre-trained backbones that do not explicitly operate at the sub-image (region / pixel) level. On the other hand, the sparse nature of facial landmarks perfectly matches the MIM objective of reconstructing the entire view from patches. unmasked, which can lead to greater fidelity of crude local peculiarities.

[0030] Related work for unsupervised landmark prediction. Several approaches have been adopted to address landmark prediction without annotated data. Equivalence learning exploits transformation equivalence as a free-supervised signal to learn landmark embeddings. Since an undesirable constant vector output would satisfy the objective, it is proposed to add a loss of diversity or to allow the application of similarity through intermediate auxiliary images to solve the problem. Another approach is generative modeling where landmarks are discovered by training networks with a reconstruction objective such as reconstructing a human image with a different pose.

[0031] Other works such as ContrastLandmark (CL) [9] and LEAD

[19] have adopted SSL methods to extract coarse features that capture the broad semantic concept and further process them to establish regional / local correspondences. These other works construct hypercolumns and compact them using proximity-guided and correspondence-guided reduction objectives, respectively. Although both methods reduce the final size of the representation, hypercolumns are enormous structures in terms of memory, and their exploitation is a computationally intensive process. Moreover, each pair of spatial features is subjected to the optimization objective, neglecting the possibility that some local correspondences may not contribute as much to the downstream task.

[0032] On the contrary, by using an embodiment of the SCE-MAE frame, it is not necessary to exploit expensive hypercolumns, and the SCE-MAE frame directly identifies and processes only salient local correspondences.

[0033] We are interested in an embodiment of the SCE-MAE 200 frame illustrated in [Fig. 2]. In short, the embodiment schematically comprises a first stage 202 of the masked image modeling type, which is implemented as an MAE, followed by a second stage 204 in which the processing proceeds by selective matching through the process of reducing the effective number of final matching pairs. The second stage is defined in the embodiment according to an example of a matching approximation and refinement block, which is trained using a (new) locality-constrained repulsion loss and is intended to directly refine only the selected matches.

[0034] Re-examination of masked image modeling. Masked image modeling (MIM) is an SSL paradigm that involves reconstructing the original image from unmasked patches. Taking MAE by He et al. as an example, given an input image x, the encoder first splits the image into Non-overlapping patches -UxP are used, to which a positional embedding is added. A class token is attached to the patch tokens but will not be affected by the subsequent masking procedure. A binary mask M is randomly sampled to determine the masked regions. The unmasked patches are denoted by ^p~xP ° M, where ^p~xP represents the Hadamard product, and are processed by the encoder to produce the patch embeddings as output. Finally, MAE uses a special [MASK] embedding to fill the masked positions, fp = ^p + [MAE0YE] o Q_, and reconstructs x from fp, minimizing the squared error. The average is calculated at the pixel level using a lightweight decoder. The reconstruction task requires the network to capitalize on the limited semantic context provided by the unmasked patches and the positional information supplied. This encourages the network to forge optimal discriminatory features to differentiate and locate important landmark regions.

[0035] Selective matching configuration: Attentive-inattentive separation. The second stage 204 of the 200 framework aims to efficiently establish local correspondences to ensure that the representations reflect the extent of similarity and dissimilarity between different facial regions. To achieve this, the second stage aims to perform selective matching, that is, eliminating the direct refinement of unimportant non-bitter correspondences and focusing on optimizing those that are critical for disambiguating bitterness. In one embodiment, a first stage identifies potential bitter and non-bitter regions. Due to the observable opposite nature of facial bitters (sparse and distinct) and non-bitter regions (dense and uniform), it is assumed (without restriction) that bitters are roughly distinguishable using the first-stage dorsal features.

[0036] With further reference to the embodiment in Figure 2, it should be noted that the first stage 202 is a pre-trained ViT backbone that is fixed (i.e., not further trained once its initial training is complete), and whose pre-trained decoder is removed, leaving the trained encoder components. The x input 206 is supplied to the first stage 202 of MAE and to a patch block 208. Examples of x input and a "patch x" are shown figuratively in 206A and 208A respectively and are further represented as a grouping of tokens (e.g., 212A), watchful tokens (e.g., 213A), and a class token (CLS) (e.g., 21 IB) as described in more detail below. The 212A tokens are supplemented by positional embeddings, for example using an element addition function.

[0037] The first stage 202 comprises a plurality of transformer blocks (ViT) in the form of an encoder layer 214 producing as output a token grouping (e.g., 212B) of the attentive tokens and the CLS token. In 216, a CLS similarity block processes the token grouping 212B and produces as output a grouping (e.g., 212C) of attentive tokens, inattentive tokens, and a CLS token. Another output is an attentive mask 218. In one embodiment, an all-pair attentive mask M of size PxP, where P is the number of tokens, stores a 0 (inattentive) or 1 (attentive) signifying a token type after the attentive-inattentive split. An input (i,j) represents the matching type between the token pair at (i,j). Based on the type of token pair, a repulsion coefficient matrix is ​​constructed in an embodiment as described below.The matrix is ​​useful for loss-based training, but is not used after training, e.g., when second-stage 204 training is completed.

[0038] The aggregation block 220 processes the grouping 212C to produce as output a grouping (for example, 212D) comprising aggregation centers, attentive tokens, and the CLS token. A final encoder layer block 220 processes the grouping 212D for supply to the second stage 204 (for example, as aggregation centers, attentive tokens, and CLS tokens (not shown)). Thus, through processing, the MAE patch tokens are split into attentive tokens (shown as bordered circles) and inattentive tokens (shown as bordered circles with an X) based on their similarity to the CLS token (shown as a black circle). Inattentive tokens are aggregated into K clustering centers (e.g., shown as a black square or bordered square as an example of clusters in the 212D grouping, although more than 2 such centers may be determined). Further details are provided below.

[0039] The CLS token (for example, 213B) represents the image and is obtained by aggregating information from the other patch tokens over several layers. Since landmarks are sparse and have a more distinct texture, it is expected (without restriction) that the corresponding tokens will have a significant influence on the representation of the CLS tokens. The first-stage block (i.e., the MAE frame as pre-trained and frozen) used to train the second stage 204 is configured, in one embodiment, to compute a similarity vector between the CLS token and all patch tokens as follows:

[0040] / K9ds \ ,y Sim., = Softmax -7=- £P

[0041] where K, d, and N denote respectively the CLS token query vector, the patch token key matrix, the latent dimension, and the number of patch tokens. Here, pd and K ej>Nxd. The A patch tokens are split into two Groups: (1) an attentive group, consisting of the v • N tokens with the highest similarity score to the CLS token, and (2) an inattentive group, consisting of the remaining (1 - v) • tokens. Here, this is a hyperparameter between 0 and 1. We observe that the inattentive tokens mostly cover non-bitter facial regions (e.g., see [Fig. 2] in 222), such as the cheeks and forehead, as well as the background. From now on, it is assumed (without restriction) that the attentive tokens cover bitter and important facial regions, while the inattentive tokens correspond to unimportant non-bitter regions.

[0042] Aggregation of inattentive tokens. Since several inattentive tokens often correspond to the same facial region (e.g., cheek, forehead, etc.), the downstream matching objectives associated with them would likely be redundant. By applying an aggregation algorithm to inattentive tokens, many non-pear regions can be represented with only a handful of aggregation centers. A selective match can then be defined by rejecting all non-aggregation center tokens, ensuring that no match is established with them.

[0043] In one embodiment, a simple density peak aggregation algorithm is adopted (Long, Sifan, Zhen Zhao, Jimin Pi, Shengsheng Wang, and Jingdong Wang. “Beyond attentive tokens: Incorporating token importance and diversity for efficient vision transformers.” In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 10334–10343, 2023, which is incorporated herein in its entirety by reference), wherein two variables p and ô are defined for each inattentive token. Here, p measures the density of the l-th token and ô computes the minimum distance between the l-th token and any other inattentive token having a higher density. Mathematically, they are defined as suif- if 3 f S.tp; > (3) oihetwise

[0044] where t;, tj E TiW / , ctTiWz all denote inattentive tokens. Given that the The aggregation center must have a higher density than neighboring tokens and must also be far from other aggregation centers. The score of the aggregation center of the 'th token' is calculated by P, • ô;. The tokens among the highest-scoring aggregation centers are selected as aggregation centers, where Kc is a hyperparameter. The remaining inattentive tokens are discarded, and the aggregation center tokens subsequently act as representative proxies for them.

[0045] Selective matching using CARB. In the second stage 204 providing a matching approximation and refinement block (CARB), the discarded inattentive tokens are first substituted (for example, to the aggregate approximation block 224), along with their corresponding aggregate centers, and the relevant visual features are aggregated to obtain a complete 2D feature map (for example, 226) also visually represented in 222 in [Fig. 2]. With the backbone fixed, the feature map 226 is passed through a light projector 230, which is supervised by a (new) locality constraint repulsion loss (LCR) 232. Since the LCR loss 232 operates on the features of attentive tokens and inattentive aggregate centers, only the most important matches are directly refined, thus obtaining a selective match.LCR loss weakens existing erroneous matches in a weighted manner considering the proximity (locality) and match type (repulsion) constraints of the token pairs (e.g., the attentive mask 218).

[0046] Locality Constraint Repulsion Loss (LCR). The LCR 232 loss is designed and exploited to produce high-fidelity fine features by optimally refining local correspondences. Henceforth, and denote respectively the approximated attentive and inattentive tokens (aggregation centers), and define T = T0 / / u^inatt as the game of all the tokens considered.

[0047] The matching can be formally defined as the probability that a patch token fj matches h to a patch token in the image, which is expressed as follows: exp ((4¼ ÇxX . Cr)A) ) H) 1¼¼]¼ ~............................................................r,

[0048] où <e>h Q) is the final representation of projected peculiarity of the tp correction and ' is the temperature parameter.

[0049] It is observed that image patches that are spatially distant from each other often correspond to different facial regions. Consequently, it should follow that strong correspondences between distant patches are likely to be erroneous and should be discouraged. A locality constraint is calculated to formalize this idea using the following function:

[0050] where h, 0 £ T, and 11*11 calculates the spatial distance. The log function saturates the coefficient to discourage the network from excessively focusing on separating very distant matches. Although a similar constraint was introduced in Thewlis, James, Hakan Bilen, and Andrea Vedaldi, "Unsupervised learning of object frames by dense equivariant image labelling." Advances in Neural Information Processing Systems 30 (2017), which is incorporated here in its entirety, the main motive was to avoid collapse during equivalence learning.

[0051] Given the approximate attentive and inattentive token games and T^^), there are three types of correspondences: attentive-attentive (att - att), attentive-inattentive (att - inatt), and inattentive-inattentive (inatt - inatt). A repulsion coefficient is introduced to quantify the importance of each type of correspondence:

[0052] ^att-att if tp tj g Tw / / rep(ti, tj) = ^inutt-inatt SZ t^ tj G Tinatt ^att4natP otherwise

[0053] where each coefficient r is a hyperparameter. In a practical embodiment, -ait

[0054] and T„ / r inatt are set to be greater than Tinatt — inatt to prioritize facial bitter differentiation and bitter versus non-bitter disambiguation respectively over non-bitter differentiation. The attentive mask can be used to determine the repulsion coefficient for specific token pairings. In one embodiment, for an all-pair matrix determined from the mask M, an attentive-attentive (att - att) pairing can store Tatt-ath, an attentive-inattentive (att_ inatt) pairing can store Tinatt — inatt, and an attentive-inattentive (inatt _ inatt) pairing can store Tatt — inall.

[0055] By combining all the components defined above, the LCR loss is mathematically expressed as follows: "H

[0056] LCR loss aims to forge effective local peculiarities by systematically weakening erroneous, spatially distant matches among and between important and unimportant patch tokens.

[0057] Inference. After training, optimized representations are obtained for all image regions. From there, during inference, the aggregation and inattentive token approximation procedures are bypassed (e.g., disabled) and only the original features for downstream tasks are used.

[0058] Since the second stage training is complete, calculating the correspondence with a trained network is not necessary. During training, token type separation and aggregation, as well as token-pair correspondences, were calculated and refined to train the network to differentiate tokens from different parts of the face.

[0059] In one embodiment, a final network comprises the first and second stages as trained (i.e., the MAE-based encoder and the projector network), for example, to provide a landmark matching network. In another embodiment, the regressor components are coupled to the projector network as trained, and this network is further trained to perform landmark regression (for example, see [Fig. 6]).

[0060] Furthermore, in one embodiment such that, for the purpose of fair comparison with prior work, a spatial expansion is performed for the size of the feature input in the projector. A naive reduction in the patch size to expand the output not only quadratically increases computational and memory costs, but can also lead to the formation of inferior feature representations. Instead, a cover-and-stride technique is adopted to produce finer and richer enlarged representations.

[0061] Experiments. Datasets. The dorsal (floor 1) was pre-trained on the CelebA dataset (by Liu et al.), which contained 162,770 images. Facial landmark detection was evaluated on four datasets: MAFL (or Zhang et al.), 300W (by Sagonas et al.), and two variants of AFLW (by Koestinger et al.). MAFL consisted of 19,000 training images and 1,000 test images. 300W comprised 3,148 training images and 689 test images. The AFLWM dataset contained 10,122 training images and 2,995 test images, which were cropped from MTFL (by Zhang et al.). The AFLWR dataset contained tighter cropped facial images, with the training and test sets containing 10,122 and 2,991 images, respectively. It is worth noting that 300W provided 68 annotations per image, while the other three datasets provided only 5 annotations.

[0062] Re-annotation. Although AFLWR has been used in previous work, the accuracy of the annotations is questionable. These annotation errors include errors due to semantic mismatches, translations, and random shifts. For a more consistent and reliable evaluation, the AFLWR test set was re-annotated. In the following sections, AFLWRO and AFLWRC are used to refer to the original and corrected datasets, respectively.

[0063] Implementation details. In some embodiments, models were pre-trained on the CelebA dataset using MAE with three backbones: DeiT-T, DeiT-S, and DeiT-B. All models were trained for 400 epochs with a The batch size was 512, the learning rate was 3e-4, and the patch size was 8. The image was resized to 136 x 136, and the 96 x 96 center was cropped as input for both concordance and regression of landmarks. The attention rate was set to 0.25 for DeiT-B and 0.1 for DeiT-T and DeiT-S. Aggregation was applied after the third encoder layer, and the number of aggregates (Kc) was set to 4. For LCR loss, the three repulsion hyperparameters were set to Tatt - att = 5, Tatt - inatt = 5, and Tinutt - inatt = 2. Ablation studies were performed. Although DeiT-based backbones were used, other ViT frame types (e.g., BeiT) can be used and trained accordingly.

[0064] Landmark Concordance: Evaluation Protocol. 1000 pairs of reference and test images were generated from the MAFL test set for evaluation. The first 500 pairs served as a reference for landmark concordance between the same identities, which contained the original image and its thin-plate-spline (TPS) distorted counterpart. The remaining 500 pairs were of different identities. During the evaluation, all feature maps were upsampled bilinearly at the image resolution. Landmark representations of the reference image were used to query the test image. The location exhibiting the greatest cosine similarity was considered the concordant prediction. Finally, the mean pixel error was calculated between the prediction and the actual ground state.

[0065] Quantitative Results. The present framework (e.g., in three embodiments) was compared to existing EDA frameworks as shown in Table 1 by grouping the results based on the final feature size. The best and second-best results are shown in bold and underlined, respectively. Three different backbones were used (e.g., in the embodiments) to control the number of parameters for a fair comparison. In the first group, the model embodiment with DeiT-T, being a fraction of the size of previous work, already outperforms the EDA. In the second and third groups, the embodiments clearly surpass previous work by large margins of ~20% and ~44% for identical and different identities, respectively.This is attributed (without restriction) to the highly fertile initial features from MAE pretraining, which, when strategically refined by selective matching using CARB, generate distinctive final features that were vital for successful bitter match.

[0066] [Table 1] Table 1 Frame (Meth #Mil Parameter Dim. by Same Diff ode) lions specificity Average pixel error DVE 12.4 64 0.92 2.38 CL 23.8 64 0.92 2.62 LEAD 23.8 64 0.51 2.60 DeiT-T 5.4 64 0.47 1.99 CL 23.8 128 0.82 2.19 DeiT-S 21.4 128 0.31 1.69 CL 23.8 256 0.71 2.06 LEAD 23.8 256 0.48 2.50 DeiT-S 21.4 256 0.33 1.72 DeiT-B 85.3 256 0.27 1.61

[0067] Qualitative results. Concordance results for landmarks between different identities were visualized (not shown) and compared to existing EDA methods. Discrepancies on different landmarks when using CL and LEAD were shown in different column groups; for example, the first group of three columns contained eye-related discrepancies. The framework method of the present embodiments clearly achieved more accurate concordance performance on all landmarks, even on challenging examples such as those wearing glasses. It is acknowledged that the framework method of the present embodiments experienced some failures when the poses of the reference and test samples were greatly dissimilar or when landmark regions were severely occluded.

[0068] Landmark detection: Evaluation protocol. The pre-trained ridge and spotlight were frozen, and only a lightweight regressor was trained. The regressor according to the present embodiment comprises a convolution block (instead of a linear layer) and a linear layer during training on all annotated samples. The convolution block uses the spatial context to produce l intermediate heatmaps for each landmark, which are converted into l 2D coordinate pairs by a soft-argmax operation and delivered to a linear layer that outputs the final landmark prediction. For all experiments, 1 = 50. The concatenated first- and second-stage features are exploited as a more robust input to the regressor, expecting the ridge to provide rich and task-independent representations during the first stage and that the projector complements task-specific cues for the detection of landmarks during the second stage.

[0069] All samples annotated. The framework method of the embodiments was compared to previous work on landmark detection references as shown in Table 2 (e.g., Quantitative Evaluations on Landmark Detection with All Samples Annotated). Comparison with the existing EDA indicates the percentage error of interocular distance on four human face datasets: MAFL, AFLWM, AFLWR, and 300W. For AFLWR, the results are reported on both the original (AFLWRO) and corrected (AFLWRC) datasets. The framework embodiments presented here, despite using significantly smaller features and avoiding costly hypercolumns, outperform previous work on each of the four datasets, even with the embodiment having the smallest backbone, DeiT-T.Considering the best results of the implementations of the framework method, the framework method achieves a performance gain of ~9%-15% on different references.

[0070] [Table 2] Table 2 Framework (Method) #Param s. Millions Dim. of particularity Hypecol MAF AFLWAFLWjAFLWROW . Used L Interocular distance (%) DVE 12.6 64 N 2.76 6.96 6.33 5.58 4.58 CL 23.8 3840 O 2.76 6.17 5.69 5.06 4.84 LEAD 23.8 3840 O 2.44 6.05 5.71 5.11 4.87 DeiT-T 5.4 256 N 2.20 5.89 5.54 4.86 4.22 DeiT-S 21.4 512 N 2.08 5.33 5.40 4.69 3.94 DeiT-B 85.3 1024 N 2.07 5.23 5.33 4.60 3.95

[0071] Such convincing performance is a testament to the discriminatory capacity of the features, which provide complex disambiguation cues to the regressor for locating landmarks. It is also emphasized that all methods allow for a lower error with the re-annotated AFLWRG trial set, thus confirming the superior annotation quality.

[0072] Limited annotated samples. The embodiments of the present framework were compared to previous work on landmark detection under different annotation settings on the AFLWM dataset. Table 3 (e.g., Evaluations) Quantitative studies on the detection of landmarks with limited annotated samples) indicate the percentage error of the interocular distance. Embodiments of the framework have outperformed all existing SSL methods with a significant performance gain under all annotation and feature dimension settings. Specifically, a relative gain of 8.6% on average and up to 20.1% compared to the existing EDA was achieved. Furthermore, a lower standard deviation was observed with repeated experiments, indicating that the present framework produces optimal features more consistently, thus demonstrating its robustness.

[0073] [Table 3] Table 3 Method ode Particularity Dim. Number of annotated 1 5 10 20 50 100 DVE[ 46] 64 14.23 ± 1 .45 12.04 + 2 .03 12.25 + 2. 42 11.46 + 0 .83 12.76 + 0.53 11.88 + 0 .16 CL[9] 64 24.87 + 2 .67 15.15 + 0 .53 13.52+1, 08 11.77 + 0 .68 11.57 + 0.03 10.06 + 0 .45 LEAD

[19] 64 21.80 + 2 .54 13.34 + 0 .43 11.50 + 0, 34 10.13 + 0 .45 9.29 + 0 .41 9.11+0. 25 SCE- MAE 64 18.41 + 1 .21 11.79 ± 0 .44 10.57 ± 0. 24 9.65 ± 0. 14 8.60 ± 0 .17 831 ± 0. 06 CL[9] 128 27.31 + 1 .39 18.66 + 4 .59 13 39 + 0.30 11.77 + 0 .85 10.25 + 0.22 9.46 + 0. 05 LEAD

[19] 128 21.20+1 .67 13.22+1 ,43 10.83 + 0, 65 9.69 + 0.41 8.89 + 0 .20 8.83 + 0.33 SCE- MAE 128 20.14 ± 1 .76 11.99 ± 0 .71 10.40 ± 0. 22 9.25 ± 0. 14 8.49 ± 0 .19 7.96 ± 0, 21 CL[9] 256 28.00 + 1 .39 15.85 + 0 .86 12 98 + 0.16 11.18 + 0 .19 9.56 + 0 44 9.30 + 0. 20 LEAD

[19] 256 21.39 + 0 .74 12.38 + 1.28 11.01 + 0.48 10.06 + 0 .59 8.51+0.09 8.56 + 0.21 SCE- MAE 256 17.08 ± 1 .35 11.28 ± 0 .54 10.30 ± 0.09 8.95 ± 0.08 8.20 ± 0 .20 7.58 ± 0.09

[0074] Ablation studies: Importance of each component. To better understand the proposed SCE-MAE framework, the component-based ablation analysis of the landmark matching task is reported in Table 4. The first three rows The first two lines indicate the use of only the first-floor backbone features, while the last two lines, respectively, indicate the inclusion of aggregation and LCR loss in the proposed framework. CL and LEAD use hypercolumns, while the proposed framework leverages the features of the last vanilla layer of the pre-trained MAE, which is indicated as the baseline. Using the backbone alone, the baseline outperforms CL and LEAD, validating the rationale that the MIM is a pretext task better suited for learning to represent bitterness. Aggregation is observed to aid concordance between the same identity, while LCR loss enhances concordance performance between different identities.Overall, these trends converge with expectations: initially, first-stage MAE features at the regional level capture local subtleties but are too crude to generalize landmarks across different identities; aggregation disambiguates landmarks from non-important regions, improving the same identity concordance performance; and finally, LCR loss forges critical local correspondences between important facial regions, resulting in the best performance for both settings.

[0075] [Table 4] Table 4 Aggregate Method A LCR Identical Differences Mean Pixel Error CL - - 0.69 537 LEAD - - 2.35 6.22 Baseline NN 0.55 3.51 SCE-MAE ON 0.30 4.04 OO 0.27 1.61

[0076] Visualization of landmark representations. The t-SNE 300 plots of landmark representations corresponding to 1000 test images are visualized in [Fig. 3]. Since LEAD

[19] only performs knowledge distillation in its second stage, the first-stage hypercolumn representations were used as they can be considered the upper bound of the second-stage objective. For each method, the t-SNE is run 00 times, and the mean and standard deviation of the Silhouette Coefficient, a metric (higher, better) indicating the quality of the aggregation as a function of the mean intra-aggregate (higher, better) and inter-aggregate (higher, better) distance, are reported. For CL, the samples within the aggregate (e.g., 302) are more dispersed, resulting in a greater intra-aggregate distance. For LEAD, although the Although the aggregates are denser, the aggregates in the left / right corner of lip 304A / 304B and nose 306 are not clearly separated, resulting in a smaller inter-aggregate distance. The drawings from an SCE-MAE embodiment of this illustration depict both well-separated and dense clusters, reflected in a high Silhouette Coefficient, thus corroborating the superior quality of such landmark representations.

[0077] According to one embodiment, a trained bitterness detection network according to the framework proposed herein is configured as a component of an effects pipeline, for example, a makeup effects pipeline. The trained bitterness detection network can form a component of a face tracking engine. The pipeline can be provided as a component of a virtual try-on (VTO) application to simulate the application of an effect, for example, a makeup effect, to a face image. The image can be a still image or an image (for example, a frame) from a video. As noted earlier, a trained bitterness detection network can be configured as a component of other downstream applications, such as 3D facial reconstruction, facial recognition, facial emotion / expression recognition, and more contemporary applications such as facial beauty prediction. A VTO application is described in more detail.

[0078] Downstream applications. According to prior art embodiments, an effects rendering pipeline (for example, a workflow of operations of a computing device) can be configured to apply an effect to an image, such as an effect that simulates a product or service using facial features. In one embodiment, the effect simulates a product or service applied to the face to provide a virtual try-on experience (VTO).

[0079] The product may be a makeup product or a device product (e.g., eyeglasses, hearing aids, dental appliances including braces, etc.); and the service includes a cosmetic procedure or a surgical procedure or other procedure modifying the face (e.g., changing the shape, size, color, etc. of a facial feature).

[0080] In one embodiment, a facial landmark detection network is a component of, or communicates with, an application, and the facial landmarks determined by the network are provided for further use by the application. The application may include any of the following: a VTO application; a teleconsultation application; a video chat application; a videoconferencing application; or a facial recognition application, among others.

[0081] Figure 4 is a schematic diagram of the operations of a single threaded-effect pipeline according to a prior art embodiment. The pipeline comprises A single 402 thread. A current frame to be processed is designated frame t. An immediately preceding frame that has been processed is designated frame t - 1, and an immediately following frame to be processed is designated frame t + 1. Frame t is received at operation 404 to be processed sequentially at operation 406, such as by a landmark detection component (e.g., a face tracking engine). After that, an effects rendering component processes effects at operation 408 to produce frame t for output at operation 410, where, at output, one or more effects are applied to frame t, for example.

[0082] The face tracking organ includes, in one embodiment, one or more neural networks for such a purpose, for example, a network in accordance with frame 200 of [Fig.2].

[0083] In one embodiment, the face-tracking element is adapted to track (i.e., locate) classes of objects related to a face, including a face object itself. The output of such a face-tracking element includes a bounding box, a mask, or some other structure for inferring a cropped frame (e.g., a cropped face image). In a cropped frame, for example, any background in the frame, such as the t-frame, is minimized. Figure 5A shows a representative t-frame 500 (an example face image 502) including a face 504 and a background 506. The background may include other portions of the subject as well as non-subject portions. A bounding area 508 represented as a dotted area shows coordinates to define a cropped image, including a face 504 and reduced background content (part of the background 506).

[0084] In one embodiment, the landmark detection component or face tracking organ is configured to determine (e.g., predict) a face's location within a frame and locations within the face of one or more facial landmarks (e.g., semantic facial features). In one embodiment, such features include a face contour (e.g., portions of a jaw, a chin, or both), a nose, an inner mouth, an outer mouth, a left eye, a right eye, a left eyebrow, and a right eyebrow, as shown in [Fig. 5B].

[0085] Figure 5B is an illustration of the output of a face-tracking component of an effects pipeline showing a cropped face image 510 of a face 504, as described, and face points 512 according to one embodiment. Respective groups of face points 512 include, in one embodiment, face contour points 512A, eyebrow face points 512B, and nose face points 512C, etc., for each object for which face points are determined. The schematic representation is that of a cropped face image Figure 510 is annotated with face point groups for illustrative purposes. The output of the face tracking device does not necessarily have to include an annotated image and can be separate data. The face points in an individual group are numbered or indexed (e.g., 0, 1, 2...) and help define the outline of the detected object. The face tracking device assigns each point so that it is placed in locations consistent with the outline of the object it represents. For example, a particular point might always be at the right corner of the mouth. In one example, the face points are X,Y pixel coordinates relative to the cropped face image 510 and are associated with respective detected objects from the (e.g., one of the) arrays of a face tracking device.

[0086] Figure 6 illustrates a computing environment 600, according to one embodiment, such as for practicing one or more aspects of a method, for example, including, but not limited to, VTO operations. The computing environment 600 shows a user computing device 602, such as a smartphone, a communication network 604, a server 606, and a server 608. The communication network 604 includes wired and / or wireless networks, which may be public or private and may include, for example, the Internet. The server 606 includes a server computing device such as for providing a website. The server 608 includes a server computing device such as for providing e-commerce transaction services. Although shown separately, the servers 606 and 608 may comprise a single server device. The computing environment is simplified.For example, payment transaction gateways and other components such as those used to perform an e-commerce transaction are not represented.

[0087] The computing device 602 includes a storage device 610 (for example, a non-transient device such as memory and / or an integrated circuit disk, etc.) for storing instructions which, when executed by a processor (not shown), cause the computing device 602 to perform operations such as a computer-implemented method. The storage device 610 stores a virtual fitting application 612 comprising components such as software modules providing a user interface 614, a face-tracking component 616 with one or more neural networks 618 configured for face detection including face point determination, a VTO rendering pipeline component 620 with a stabilization component 622, a product recommendation component 624 with product data 626, and a shopping component 628 with a shopping cart 630 (for example, purchase data).One or more of the 618 neural networks includes a two-stage frame network, for example, in one embodiment, comprising a first stage having a MIM backbone, or in a . An embodiment comprising a first stage having a MIM backbone (such as MAE of the first stage 202) and a second stage (for example 204) implementing selective matching as described here and / or as trained as described here. In one embodiment, the second stage is a projector array configured to include or couple to a regressor 619.

[0088] In one embodiment, the VTO application is a web application as obtained from the 606 server. Although not shown, the 602 user device may store a web browser for running the web-based VTO 612 application. In one embodiment, the 612 VTO application is a native application conforming to an operating system (also not shown) and to the software development requirements that may be imposed by a hardware manufacturer, for example, of the 602 user device. The native application may be configured for web-based or similar communication to the 606 and 608 servers, as is known.

[0089] Figure 6 shows various input and output data or information associated with the use of the VTO application 612, for example. This includes an input image 640 of the user to be processed for a VTO experience, an output image 642 on which product effects are simulated providing a VTO experience, a VTO product selection 650 comprising user input selecting one or more product effects to be simulated, VTO product options 652 comprising options for products to be virtually tried, for example for selection by a user of the device 602, and purchase transaction information 660 comprising purchase information provided to and / or received from a user to purchase a product.

[0090] In one embodiment, via one or more user interfaces 614, VTO product options 652 are presented for selection to be virtually tried out by simulating effects on an input image 640. In one embodiment, the VTO product options 652 are derived from, or associated with, product data 626. In one embodiment, the product data 626 can be obtained from the server 606 and provided by the product recommendation component 624. Although not shown, user or other input can be received for use in determining product recommendations. The user can be prompted, for example via one of the interfaces 614, to provide input to determine product recommendations. In one embodiment, the product recommendation component 624 communicates with the server 606.In one embodiment, server 606 determines the recommendation based on the input received via component 614 (e.g., and 624) and provides product data accordingly. User interface 614 can then display the product choices. VTO, for example, by updating its display in response to data received when the user navigates or otherwise interacts with the user interface 614.

[0091] In one embodiment, one or more user interfaces 614 provide instructions and commands to obtain the input image 640 and the VTO product selection input 650, such as the identification of one or more recommended VTO product options 652 to try. In one embodiment, the input image 640 is a user's face image, which may be a still image or a frame from a video. In one embodiment, the input image 640 may be received from a camera (not shown) of the device 602 or from a stored image (not shown). The input image 640 is provided to the face-tracking organ 616, for example, for processing to detect objects in the face image using one or more trained deep neural networks 618. In one example, the network classifies, locates, or segments a face mask (or other occlusive object) in the image.In one embodiment, a face mask presence classification example is useful for outputting a request (e.g., an instruction to a user, for example via 614 user interfaces) to lower or remove a face mask. This can be applied to any occlusive object for which the face tracking engine is trained. In one embodiment, the occlusion can be manipulated during rendering, as described here, to avoid rendering on an inclusion.

[0092] In one embodiment, an output (not shown) from the face-tracking component 616, such as classification results, localization results, or segmentation results for one or more detected objects, is provided to the VTO rendering pipeline component 620. In one example, the output might include a bounding box, for example 508 of [Fig. 5A], and, as shown in [Fig. 5B], face points 512 (e.g., groups thereof) for detected objects, etc. The input image 640 is also provided to (e.g., made available to) the VTO rendering pipeline component 620. The VTO product selection 630 is also provided to the VTO rendering pipeline component 620 to determine the effects to be rendered.In an embodiment related to makeup simulation, one or more effects may be indicated such as for one or more of the product categories including: lips, eyeshadow, eyeliner, eyeshadow, etc.

[0093] The VTO 620 rendering pipeline component, in one embodiment, determines whether to render one or more product effects on the input image 640 to simulate a fitting. For example, in response to the facial mask classification output, the VTO 620 rendering pipeline component may decide not to render a product effect, for example, because a mask (an occlusion) is detected. When a When a face mask is detected, for example, the VTO rendering pipeline component 620 can optionally trigger the user interface 414 to prompt the user to remove the face mask. A new image (a new instance of image 640) can be received and processed by the face tracking component 616. In one embodiment, images are received continuously as a component of a live stream (for example, a selfie video). In one embodiment, occlusions are handled during rendering in such a way as to avoid rendering over an inclusion, as described here.

[0094] If the VTO rendering pipeline component 620 determines to render one or more product effects, in one embodiment, the VTO rendering pipeline component 620 renders effects on the input image 640, for example, by drawing (rendering) layered effects, one layer for each product effect, to produce the output image 642. Portions of the operations of the VTO rendering pipeline component 620 (for example, such as drawing the layers) may be performed by a graphics processing unit, in one embodiment. The rendering conforms to the product data 626 as selected by VTO product selection 650 and is responsive to the location of detected objects. For example, a VTO product selection of a lipstick, lip gloss, or other lip-related product calls for the application of an effect to one or more detected mouth- or lip-related objects at their respective locations.Similarly, a selection of eyebrow-related products calls for the application of a selected product effect to the detected eyebrow objects. Generally, for symmetrical looks, the same eyebrow effects are applied to each eyebrow, the same lip effect to each lip, or the same eye effect to each eye area, but this isn't always the case. In one example, the rendering is applied to an area that is relative to the detected objects, such as one or more adjacent detected objects. Some VTO product selections include a selection of more than one product (for example, defining an "appearance"), such as coordinated products for eyebrows and eyes, or other combinations of detected objects, including the entire face.Product data can define respective "appearances" that group related products, for example, and associate the appearance with a name to be displayed via the user interface, such as a command allowing the user to select an appearance from a group of appearances presented in a list, table, or other presentation format. The VTO 620 rendering pipeline component can render each effect, for example, one at a time until all effects are applied. The order of application can be defined by rules or in the product selection, for example, a lipstick before a lip gloss.

[0095] In an embodiment where an occlusive object is detected and its location is determined, for example, as represented in a segmentation mask, the rendering can respond to such a segmentation mask. An effect rendering can be applied to portions of the face that are not occluded. A segmentation mask can indicate which facial pixels are available to (for example, can) receive an effect such as a makeup effect, and which pixels are not available to receive an effect.

[0096] The user interface 614 provides the output image 642. The output image 642, in one embodiment, is presented as a portion of a live stream of successive output images (each being an example of image 642), such as when a selfie video is augmented to present an augmented reality experience. In one embodiment, the output image 642 is presented together with the input image 640, as in a side-by-side display for comparison. In one embodiment, the output image 642 can be saved (not shown), for example on the storage device 610, and / or shared (not shown) with another computing device.

[0097] In one embodiment, the input images (not shown) include input images from a videoconference session, and the output images include a video shared with another (or several other) participant(s) in a videoconference session. In one embodiment, the VTO application is a component or extension module of a teleconsultation application or a videoconferencing application (each not shown) that allows the user of device 602 to wear makeup during a teleconsultation or videoconference (respectively) with one or more other conference participants.

[0098] In one embodiment, the VTO 620 rendering pipeline component is configured to apply object stabilization (for example, using a 622 stabilization component) to stabilize respective locations of detected objects between, for example, successive frames of a video.

[0099] In one embodiment, the face-tracking device 616 locates facial features but without detecting the presence of a face mask (or other occlusive object). Consequently, in such an embodiment, the operations of the VTO rendering pipeline component 620 are configured without taking occlusions into account.

[0100] According to one embodiment (not shown), a teleconsultation, video chat, or videoconference application incorporates an integrated virtual try-on feature so that a user can appear to have a selected makeup effect during a chat or conference. This will be an environment similar to environment 600, which can be configured. In a video chat or videoconference environment, a user device offers a teleconsultation or videoconferencing application with integrated VTO functionalities. The application is stored on a non-transient storage device. Integrated VTO features are provided as by the components of a VTO application.

[0101] The user device is configured to communicate with a server providing video chat or conferencing services in order to communicate with one or more other user devices. Examples of platforms providing a video conferencing service, which are not exhaustive, include MICROSOFT TEAMS™ available from MICROSOFT Corporation of Redmond, WA; ZOOM ONE™ available from ZOOM Video Communications, Inc. of San Jose, CA; and GOOGLE MEET™, available from GOOGLE LLC of Mountain View Parkway, among others.

[0102] In short, teleconsultation or videoconferencing services allow the sharing of live video between two or more user devices communicating via an intermediary device, namely a server. A first user device obtains a video stream from a camera (either an internal camera or an external camera coupled to it) and provides it to the server for communication to other participating devices (e.g., members of a videoconference, a clinician, or a beauty consultant in a teleconsultation) who are participating in the conference as maintained by the conference or chat server. Such a server provides the respective video streams received from the respective user devices to another user device for the conference or chat.It is understood that the server can process (for example, perform video processing of) any of the video streams it receives and retransmits for a conference or teleconsultation.

[0103] The respective user teleconsultation or videoconferencing applications running on the respective devices can be configured to present the received video streams according to a layout or view selected in a user interface on a display device. A layout or view can display a member who is the active speaker, a pinned conference member, or all conference members, etc., as is known.

[0104] In one embodiment, the conference or chat application is configured to apply at least one effect to the images from the user's device, allowing a virtual try-on during the teleconsultation or video conference meeting, so that other members receive the output images as rendered using the integrated VTO application with the at least one effect applied.

[0105] An input image represents a frame of an input video stream from a camera local to the user device, while an output image represents a frame of an output video stream determined from one or more frames of the stream Input video. Each output image is presented according to the user interface or other application commands. Thus, sometimes during a teleconsultation or conference, the output image may not be displayed by the user's device, for example, when another participant has a focus and only that participant's feed is shown. However, the output image is sent to the server for retransmission for display (e.g., selective) by other user devices according to the respective commands of their local teleconsultation or video conferencing applications. It is understood that no VTO effect is applied if the camera command is "disabled" and no camera image is shared with the server.

[0106] In one embodiment, the teleconsultation, conferencing, or chat application is configured with user interfaces that have controls to allow a user to select whether to apply a VTO effect. In one embodiment, the user interface is activated to receive user input to select a preview of one or more effects, invoking the VTO components to process the input video stream and render an output video stream with the rendered effect(s) for display by the user's device. In one embodiment, during the preview, the output video stream is not shared with the server and is therefore not provided to other devices during the preview period. In one embodiment, the user interface is activated to present detailed information about each of the effect products and is further activated to allow the purchase of products.

[0107] In this work, a two-stage framework for addressing self-supervised estimation tasks of facial landmarks is disclosed. To learn locally efficient and distinct representations, structured improvements are targeted at both stages. For the first stage, regional MAE is used instead of instance-level SSL methods to infer more fertile initial representations. For the second stage, it is shown to be beneficial to identify important facial regions and directly refine only salient matches. The disclosed approaches (e.g., in embodiments) generate discriminative and high-quality landmark representations that result in superior performance compared to previous EDA work on landmark matching and detection tasks.

[0108] Aspects and features of the embodiments will be apparent to a person skilled in the art and will include those of the following variants:

[0109] A computer-implemented method for supervised self-learning (SSL) training of a network for detecting facial landmarks for a face in an input image, wherein the network includes encoder components encoding face features and decoder components determining local matches between features to determine estimates of landmarks, the method comprising: training a first network comprising a masked image modeling (MIM) network that processes non-overlapping patches determined from the input image with an SSL objective, wherein the encoder components include a portion of the MIM network to encode features of the input image for decoding; and training a second network comprising the encoder components, as trained, in series with the decoder components, the decoder components trained to determine local matches comprising respective relationships between the features of the input image provided by the encoder components.

[0110] Method according to the above, in which the MIM network comprises an MAE network.

[0111] A method according to the above, wherein the MIM network, as trained, is configured to: provide respective tokens for patches for processing by the decoder components; and combine information from tokens relating to non-peer regions of the input image to define approximate tokens, reducing the number of tokens for processing by the decoder components.

[0112] A method according to the above, wherein the MIM network, as trained, is configured to: define respective patch tokens for each patch and a class token (CLS) representing the image as an aggregation of information from the respective patch tokens; identify each patch token as an attentive token or an inattentive token according to a respective similarity to the CLS token determined for each patch token; and combine information from inattentive tokens to provide approximated inattentive tokens, reducing the number of inattentive tokens to be processed by the decoder components.

[0113] A method according to the above, wherein the MIM network, as trained, is configured to perform inattentive token aggregation to combine information from the inattentive tokens, defining aggregation centers to represent the information.

[0114] Method according to the above, wherein the MAE network is configured as a vision transformer (VIT) using self-attention mechanisms to process images.

[0115] A method according to the foregoing comprising training a final network for landmark detection, the final network comprising regressor components configured to determine landmark estimation features processed by the decoder components as trained, the regressor components configured in series with the decoder components as driven, and decoder components in series with encoder components as driven.

[0116] Method according to the above, wherein the second network comprises a projector network, the decoder components comprising a portion of the projector network.

[0117] Method according to the above, wherein the driving of the decoder components drives the projector network using a locality constraint repulsion loss (LCR).

[0118] A method according to the above, wherein the LCR operates on bitter region features and combined non-bitter region information that reduces the processing to achieve selective matching for local matches.

[0119] Method according to the foregoing, wherein CSF loss is defined in accordance with: ~ yy .WM / ) (7)

[0120] where floc : (tÿ tj ) defines a locality constraint; h-ep (tj, tj ) defines a repulsion coefficient, respectively prioritizing bitter differentiation and bitter versus non-bitter disambiguation over bitter differentiation; and tj; defines a matching as a probability that a patch token tj matches a patch token k in the image 7.

[0121] Method according to the foregoing, wherein: exp((¾.(x)f)) pfLUridLx) —............... -.....",

[0122] where <l>f is a final projected peculiarity representation of the correction, hct r is a temperature parameter.

[0123] Method according to the foregoing, wherein: - tj * 4

[0124] where tj, tj ≤ t, T is the game of all the tokens considered, and [| • || calculates the spatial distance.

[0125] Method according to the foregoing, wherein: I 11?» « « < o the rwise

[0126] where Tatt and Tinatt denote respectively attentive and inattentive tokens; T = Tatt u Tinatt , the game of all tokens considered; three types of correspondences are represented as: attentive-attentive (at - att), attentive-inattentive (att- inatt), and inattentive-inattentive (inatt - inatt), and each coefficient r is a hyperparameter.

[0127] A system comprising at least one processor, a non-transient storage device coupled to the at least one processor, the storage device storing instructions executable by the at least one processor to cause the system to: provide a facial landmark detection network for faces in the input images; and process, using the network, an input image including a face to determine and provide facial landmarks for this purpose; wherein the network comprises: encoder components configured to encode facial features, the encoder components comprising trained components of a masked image modeling (MIM) network configured to process non-overlapping patches determined from the input image, the MIM network trained with an SSL objective;and decoder components configured to determine local correspondences between features to determine estimates for facial landmarks, decoder components trained to determine local correspondences including respective relationships between features of the input image.

[0128] System according to the above, in which the MIM network comprises an MAE network.

[0129] System according to the foregoing, in which the MAE network is configured as a vision transformer (VT) using self-attention mechanisms to process images.

[0130] System according to the above, wherein the network includes regressor components configured to determine the estimates of the bit from features processed by the decoder components as driven, the regressor components configured in series with the decoder components as driven, and the decoder components in series with the encoder components as driven.

[0131] System according to the above, in which the decoder components are driven as components of a projector network, the decoder components in series with the encoder components as driven.

[0132] System according to the above, wherein the decoder components are driven using a locality constraint loss of repulsion (LCR).

[0133] System according to the above, wherein the LCR operates on particularities of bitter regions and combined information from non-bitter regions which reduce the processing to obtain selective matching processing for local matches.

[0134] System according to the foregoing, in which the CSF loss is defined in accordance with: (7)

[0135] where: ft (tÿ tj ) defines a locality constraint; hep (t? tj) defines a repulsion coefficient, respectively prioritizing bitter differentiation and bitter versus non-bitter disambiguation over non-bitter differentiation; and tj; defines a matching as a probability that a patch token ^ matches a patch token tj in the image X.

[0136] System according to the foregoing, in which:

[0137] where <l>is a final projected peculiarity representation of correction h, and is a temperature parameter.

[0138] System according to the foregoing, in which: ~ fog(h ~ MK î )'

[0139] where tj £ T, T is the game of all the tokens considered, and | . || calculates the spatial distance.

[0140] A system according to the foregoing, wherein: | E (6) V thetheiwise

[0141] where Tatt and Tinatt denote respectively attentive and inattentive tokens approximated; T = Tatt u Tinath the game of all tokens considered; three types of correspondences are represented as: attentive-attentive (att — att), attentive-inattentive (att — att), and inattentive-inattentive (inat - inatt), and each coefficient r is a hyperparameter.

[0142] System according to the above, wherein the instructions are executable to further cause the system to apply an effect to the input image using facial features.

[0143] System according to the above, wherein the effect simulates a product or service applied to the face to provide a virtual try-on experience.

[0144] System according to the above, wherein the product comprises a makeup product or a device product; and the service comprises a cosmetic procedure or a surgical procedure or other facial modification procedure.

[0145] System according to the above, wherein the network is a component of an application or communicates with an application and facial landmarks are provided for further use by the application, wherein the application includes any of a VTO application; a teleconsultation application, a video chat application, a video conferencing application or a facial recognition application.

[0146] A practical implementation may include all or part of the features described herein. These features, characteristics, and various combinations thereof, as well as others, may be expressed in the form of processes, apparatus, systems, means for performing functions, program products, and other means, combining the features described herein. A number of embodiments have been described. Nevertheless, it is understood that various modifications may be made without departing from the spirit and scope of the processes and techniques described herein. Furthermore, other steps may be proposed, or steps may be eliminated, from the described process, and other components may be added to or removed from the described systems. Accordingly, other embodiments fall within the scope of the following claims.

[0147] Throughout the description and claims of this document, the terms "include" and "contain" and their variations mean "including but not limited to" and are not intended to exclude (and do not exclude) other components, integers or steps.

[0148] The features, integers, characteristics, compounds, chemical fractions, or groups described in conjunction with a particular aspect, embodiment, or example of the invention shall be understood as applicable to any other aspect, embodiment, or example, except where there is incompatibility between them. All features disclosed herein (including the claims, abstract, and accompanying drawings), and / or all steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of these features and / or steps are mutually exclusive. The invention is not limited to the details of the preceding examples or embodiments.The invention extends to any new feature, or any new combination, of the features disclosed in this patent memorandum (including any attached claim, abstract and drawing) or to any new step, or any new combination, of the steps of any disclosed method or process.

[0149] It will be understood that aspects of a computer-implemented method and / or corresponding aspects of a computer program product are also disclosed. A computer program product, for example, includes a storage device storing computer-readable instructions which, when executed by at least one processor of a computing device, cause the computing device to perform operations of a computer-implemented method.< / l> < / l> < / e>

Claims

Demands

1. A computer-implemented method for the self-supervised learning (SSL) training of a network for the detection of facial landmarks for a face in an input image, wherein the network includes encoder components encoding facial features and decoder components determining local correspondences between features to determine landmark estimates, the method comprising: training a first network comprising a masked image modeling (MIM) network that processes non-overlapping patches determined from the input image with an SSL objective, wherein the encoder components include a portion of the MIM network to encode features of the input image for decoding;and the driving of a second network comprising the encoder components, as driven, in series with the decoder components, the decoder components driven to determine local correspondences comprising respective relations between the features of the input image provided by the encoder components.

2. Method according to claim 1, wherein the MIM network comprises an MAE network.

3. A method according to claim 2, wherein the MIM network, as trained, is configured to: provide respective tokens for patches for processing by the decoder components; and combine information from tokens relating to non-biter regions of the input image to define approximate tokens, reducing the number of tokens for processing by the decoder components.

4. A method according to claim 3, wherein the MIM network, as trained, is configured to: define respective patch tokens for each patch and a class token (CLS) representing the image as an aggregation of information from the respective patch tokens; identify each patch token as an attentive token or an inattentive token according to a respective similarity to the CLS token determined for each patch token; and combine information from the tokens inattentive to provide approximate inattentive tokens, reducing the number of inattentive tokens to be processed by the decoder components.

5. Method according to claim 4, wherein the MIM network, as trained, is configured to perform inattentive token aggregation to combine information from inattentive tokens, by defining aggregation centers to represent the information.

6. Method according to claim 2, wherein the MAE network is configured as a vision transformer (ViT) using self-attention mechanisms to process images.

7. Method according to claim 1, comprising driving a final network for marker detection, the final network comprising regressor components configured to determine marker estimation features processed by the decoder components as driven, the regressor components configured in series with the decoder components as driven, and the decoder components in series with the encoder components as driven.

8. Method according to claim 1, wherein the second network comprises a projector network, the decoder components comprising a portion of the projector network.

9. Method according to claim 8, wherein the drive of the decoder components drives the projector array using a locality constraint loss of repulsion (LCR).

10. Method according to claim 9, wherein the LCR operates on particularities of bitter regions and combined information of non-bitter regions which reduce the processing to achieve selective matching processing for local matches.

Citation Information

Cited By

  • ViT model adaptive rarefaction method and device based on packet token

    CN121835776A

  • Grouping token-based vit model self-adaptive sparsification method and device

    CN121835776B