Shielding robust transformation method for accurately detecting facial feature points, related device and related computer readable medium

The ORFormer transformer framework introduces a messenger marker Mi to detect and recover the features of occluded regions, solving the performance degradation problem of existing methods in partially invisible faces and extreme lighting conditions, and achieving higher detection accuracy and robustness.

CN120673451APending Publication Date: 2025-09-19MEDIATEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510306504.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-03-13
Filing Date
2025-03-14
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing facial landmark detection methods suffer from performance degradation on partially invisible faces or under extreme lighting conditions, and are unable to effectively detect and recover features in occluded areas.

Method used

Adopting the ORFormer transformer framework, we simulate occlusion by introducing a learnable messenger marker Mi, detect occluded regions using self-attention and cross-attention mechanisms, and recover missing features by aggregating features of non-occluded regions through messenger markers to generate high-quality occlusion-robust heatmaps.

Benefits of technology

The accuracy and robustness of facial landmark detection are improved under partial occlusion and extreme conditions, outperforming existing methods on challenging datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673451A_ABST
    Figure CN120673451A_ABST
Patent Text Reader

Abstract

The invention provides an occlusion robust transformation method for accurately detecting facial feature points, a related device and a related computer readable medium. The method comprises the following steps: operating a sheltering robust converter framework by utilizing a processing circuit, and starting to perform reasoning by utilizing a training model of the sheltering robust converter framework according to at least one input image so as to perform facial feature point detection and generate at least one output image; and during reasoning by using the training model, performing occlusion perception cross attention processing on any one input image in the at least one input image so as to obtain an occlusion image for feature recovery by combining two feature maps of two code sequences corresponding to the occluded and unoccluded partial image patches of any input image, therefore, facial feature point information about facial feature point detection is generated, and the facial feature point detection is related to a corresponding output image in the at least one output image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to facial feature point detection, and more particularly to an occlusion robust transformation method, a related device, and a related computer-readable medium for accurately detecting facial feature points. Background Art

[0002] The present invention relates to image processing, and more particularly to an occlusion robust transformation method for accurately detecting facial feature points, a related device, and a related computer-readable medium.

[0003] Despite significant progress in facial landmark detection (FLD), existing FLD methods still suffer from performance degradation on partially invisible faces, such as occluded faces or faces under extreme lighting conditions or poses. Therefore, a novel approach and related architecture are needed to address these issues without introducing any side effects or in a way that is less likely to introduce side effects. Summary of the Invention

[0004] The object of the present invention is to provide an occlusion robust transformation method, related apparatus and related computer-readable medium for accurately detecting facial feature points, so as to solve the above-mentioned problems.

[0005] At least one embodiment of the present invention provides an occlusion-robust transformation method for accurately detecting facial landmarks, which can be applied to processing circuitry within an electronic device. For example, the method may include: executing an occlusion-robust transformer framework on the processing circuitry, initiating inference using a trained model of the occlusion-robust transformer framework based on at least one input image to detect facial landmarks; performing occlusion-aware cross-attention processing on any one of the at least one input images during inference using the trained model, obtaining an occlusion map for feature recovery by merging two feature maps of two code sequences corresponding to occluded and unoccluded image patches of the input image, respectively, to generate facial landmark information related to facial landmark detection.

[0006] At least one embodiment of the present invention provides an apparatus operating according to the above method, wherein the apparatus may include at least a processing circuit within an electronic device. According to some embodiments, the apparatus may include the entire electronic device.

[0007] At least one embodiment of the present invention provides a computer-readable medium related to the above method, wherein the computer-readable medium may store program code, and when the program code is executed by a processing circuit, the program code causes the processing circuit to operate according to the method.

[0008] The present invention provides an advantage in that the method and related apparatus (e.g., processing circuitry and electronic devices) can easily perform accurate facial landmark detection regardless of whether any face displayed in at least one input image is partially invisible, thereby improving overall performance. To address the problems in the related art, the present invention provides ORFormer, a novel transformer-based method that can detect invisible regions and recover their missing features from visible portions. Specifically, ORFormer associates each image patch with an additional learnable tag called a messenger tag. The messenger tag aggregates the features of all patches other than its own. This allows the consistency of a patch with other patches to be assessed by comparing the similarity between its conventional embedding and the messenger embedding, thereby enabling invisible region identification. The method then utilizes the features aggregated by the messenger tags to recover occluded patches. Using the recovered features, ORFormer compiles high-quality heatmaps for the downstream FLD task. Extensive experiments demonstrate that the method generates heatmaps that are resilient to partial occlusion. By integrating the generated heatmaps into existing FLD methods, the method outperforms the state-of-the-art on challenging datasets such as WFLW and COFW.

[0009] These and other objects of the present invention will no doubt become apparent to those of ordinary skill in the art after reading the following detailed description of the preferred embodiment as illustrated in the various drawing figures. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 is a diagram showing an overview of ORFormer according to an embodiment of the present invention, wherein: (a) for each patch P i , the present invention can introduce patch mark X i and a learnable messenger marker M i for occlusion detection and handling; (b) the messenger tag uses the patch tag in addition to its corresponding patch tag to compute attention; and (c) if the patch P i If there is occlusion in the image, the proposed architecture can detect occlusion by evaluating the difference between the regular embedding X′i and the messenger embedding M′i, and then recover the occlusion features based on the messenger embedding aggregated from other image patches.

[0011] Figure 2is a diagram showing an overview of an occlusion robust transformer (ORFormer) method for accurate facial landmark detection according to an embodiment of the present invention, wherein: (a) the present invention's architecture may first train a quantized heatmap generator that takes an image I as input and generates its edge heatmap H, and after pre-training, prior knowledge of unoccluded faces is encoded in a codebook C and a decoder D; and (b) with a frozen codebook and decoder, the present invention's architecture may introduce an ORFormer to generate an occlusion map α and two code sequences S I and S M , thus generating the quantitative feature Z I and Z M , where the recovered feature Z rec By adding Z I and Z M is merged with the patch-specific weights given in α and used to generate the occlusion robust heatmap H rec .

[0012] Figure 3 is a diagram showing the network architecture of ORFormer according to an embodiment of the present invention, wherein: ORFormer takes an image patch P as input and generates two code sequences S through a codebook prediction head I and S M ; where S I Calculated by reference image patch marking, S M Computed via messenger labeling, the occlusion map α represents the occlusion probability of a particular patch and is inferred by the occlusion detection head.

[0013] Figure 4 is a diagram showing the integration of ORFormer into the existing FLD method, where: ORFormer is used for occlusion detection and feature recovery to produce high-quality heat maps; and the generated heat maps are used as additional inputs to the FLD method and provide recovered features to make the FLD method robust to occlusion.

[0014] Figure 5 FIG. 1 is a diagram illustrating an electronic device involved in a method according to an embodiment of the present invention.

[0015] Figure 6 is a flowchart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0016] The following description and claims use certain terms to refer to specific components. As those skilled in the art will appreciate, electronic device manufacturers may refer to a component by different names. This document does not intend to distinguish between components that differ in name but function. In the following description and claims, the terms "including" and "comprising" are used in an open-ended manner and should be interpreted as meaning "including, but not limited to..." Additionally, the term "coupled" is intended to mean either an indirect or direct electrical connection. Thus, if one device is coupled to another, that connection may be through a direct electrical connection or through an indirect electrical connection via other devices and connections.

[0017] 1. Introduction

[0018] Facial landmark detection (FLD) aims to locate specific feature points on a human face, such as the eyes, nose, and mouth. It is crucial for many downstream applications, such as face recognition, facial expression recognition, head pose estimation, and augmented reality. Recent advances in deep neural networks have significantly enhanced facial landmark detection. However, existing FLD methods perform poorly on partially invisible faces (caused by occlusion, extreme lighting conditions, or extreme head rotation) because the features extracted from the invisible regions are corrupted. A FLD method with both invisible region detection and reliable feature extraction is needed.

[0019] According to this paper, an occlusion-robust transformer, named ORFormer, is introduced, which can identify invisible regions and recover their missing features, as well as generate high-fidelity heatmaps that can adapt to challenging scenes. Figure 1 As shown, ORFormer is built on the basis of visual transformers (see Dosovitskiy et al., “An image is worth 16x16 words: Transformers for image recognition at scale”, arXiv preprint arXiv:2010.11929, 2020; hereafter referred to as “Dosovitskiy”), where image patch labels interact with each other through a self-attention mechanism10 (labeled as “attention” for brevity). For invisible part detection, each patch label X i With an additional learnable marker M i (called messenger tags) are associated.

[0020] Courier Mark M i The occlusion in patch i is simulated and the occlusions except X are aggregated. iThen, the occlusion detection module 11 (for simplicity, labeled as “occlusion detection”) accesses the difference between the regular embedding X′i and the messenger embedding M′i to determine whether there is occlusion in patch i. For occlusion handling, the feature recovery module 12 (for simplicity, labeled as “feature recovery”) recovers the features of all patches except for the one in the previous step by X′i. i and M′ i The missing features of the occluded patch are recovered by a convex combination of the convex combination of α and the combination coefficients predicted by the occlusion detection module 11. The generated features are then used to generate a heat map, and the proposed mechanism makes the output heat map robust under extreme conditions.

[0021] The high-quality heatmaps generated by ORFormer are integrated into existing landmark detection methods as supplementary information. The method achieves state-of-the-art performance on multiple benchmark datasets, including the large-scale dataset Wider Facial Landmarks in-the-wild (WFLW) (see Wu et al., "Look at Boundary: A Boundary-Aware Face Alignment Algorithm", CVPR, 2018) and the challenging facial feature dataset Caltech Occluded Faces in the Wild (COFW) (see Burgos-Artizzu et al., "Robust face Landmarkes indication under occlusion", ICCV, 2013), demonstrating the robustness of the method in handling partially invisible faces.

[0022] According to the present invention, first, a novel occlusion-robust transformer, ORFormer, is proposed, which leverages the proposed learnable messenger landmarks to model potential occlusions and recover missing features. ORFormer enables the transformer to detect and handle invisible landmarks in a general way. Second, ORFormer can be used to generate robust heatmaps, thereby enhancing the applicability of existing FLD methods to partially invisible faces. Third, our method outperforms state-of-the-art facial landmark detection methods on two challenging datasets, WFLW and COFW, demonstrating its robustness to extreme cases.

[0023] 2. Related Work

[0024] 2.1 Facial Feature Detection

[0025] Most FLD methods rely on coordinate regression and / or heatmap regression. The former directly estimates the coordinates of facial landmarks. The latter predicts a heatmap for each landmark and performs FLD through post-processing.

[0026] Coordinate regression

[0027] Some methods use linear layers as decoders to regress landmarks from convolutional neural network (CNN) features. A new loss function is designed to improve landmark supervision. Facial contours are used to constrain landmark supervision while providing a variety of extreme cases in the dataset. Fourier feature pooling can be used to handle the highly nonlinear relationship between images and facial shape. These methods provide end-to-end trainable solutions.

[0028] To exploit the self-attention mechanism in the Transformer for facial structure exploration, some studies have used the Transformer decoder to learn the mapping between CNN features and landmarks. A coarse-to-fine decoder can be used, focusing on sparse local patches. Another approach proposes learning landmark queries along pyramidal CNN features. However, these methods face challenges in handling partial occlusions, as linear layers in CNNs and global feature dependencies in the Transformer are sensitive to partial occlusions.

[0029] Heatmap regression

[0030] Inspired by advances in heatmap generation, some studies have integrated heatmap regression into facial landmark detection. They convert landmark annotations into heatmaps for model supervision. Some studies use heatmaps to estimate uncertainty and visibility probability for stable model convergence. Other studies employ stacked hourglass networks with intermediate heatmap supervision and utilize the Argmax operator to obtain landmarks. However, Argmax in heatmap regression limits direct landmark supervision due to its non-differentiable nature.

[0031] Recent research, such as replacing Argmax with other differential decoders, has enabled heatmap regression methods to be trained end-to-end and supervised by heatmaps and landmarks. For example, there is research that reduces heatmap regression to confidence scores and offset predictions to avoid a large number of upsampling layers and the use of Argmax. With the help of differential decoders, schemes (see Huang et al., "ADNet: Leveraging Error-Bias Towards Normal Direction in Face Alignment", ICCV, 2021) and schemes (see Zhou et al., "Reducing Semantic Ambiguity in Facial Landmark Detection", CVPR, 2023) designed new loss functions with landmark and heatmap supervision to alleviate the negative impact caused by landmark annotation ambiguity. There is also research that uses deep balanced models to calculate cascaded landmark refinement. The ability of heatmap regression methods to be supervised by landmarks and heatmaps while preserving facial structure has enabled them to reach the state-of-the-art status.

[0032] However, the above coordinate regression and heatmap regression methods are susceptible to partial facial occlusion, extreme lighting conditions, or extreme head rotations due to feature occlusion and corruption.

[0033] 2.2. Occlusion-Robust Facial Landmark Detection

[0034] The three main categories of methods for occlusion-robust facial landmark detection are discussed as follows:

[0035] The first category of methods estimates the probability of occlusion for each sign and mitigates the negative impact of corrupted features computed in occluded regions. For example, joint feature point, uncertainty, occlusion probability, and / or visibility predictions have been proposed. However, these methods rely on additional annotations indicating whether the sign is occluded, while the proposed method does not require such annotations.

[0036] The second category of methods explores the consensus between image patches to identify occluded image patches. For example, Burgos-Artizzu et al. proposed a method that forces regressors that focus on different image patches to reach a consensus and trust those regressors that use local features of non-occluded areas. Although this method has similar concepts to the method proposed in the present invention, it ignores occluded features and does not restore them, limiting the ability to detect occluded feature points. In contrast, the present invention proposes a messenger token, which aggregates information from non-occluded areas and supports feature recovery of occluded patches.

[0037] The third category utilizes global context to handle occlusion. Global context is directly incorporated into fully convolutional neural networks. A geometry-aware module is used to exploit the geometric relationships between different facial components. Zhu et al.'s paper, "Occlusion-Robust Face Alignment Using a Viewpoint-Invariant Hierarchical Network Architecture," CVPR 2022, models the hierarchical structure between facial components. However, these works do not explicitly detect occluded regions and therefore cannot recover features from these regions, resulting in suboptimal results for severe occlusions.

[0038] 2.3. Transformers for Feature Recovery

[0039] Transformers (e.g., Dosovitskiy

[15] ) have been widely used in vision tasks. Transformers utilize attention mechanisms to capture long-range dependencies between tokens, but are sensitive to feature corruption or partial occlusion.

[0040] To address this issue, cross-attention has been exploited to recover occluded features between different frames in the context of object re-identification. However, this method relies on multiple frames, while our method focuses on recovering occluded features within a single image. Park et al., “HandOccNet: Occlusion-Robust 3D Hand Mesh Estimation Network”, CVPR, 2022, proposed a 3D hand mesh estimation method that involves training CNN blocks to separate primary and secondary features and then recovering occluded features using cross-attention. In contrast, our method integrates occlusion detection and processing mechanisms into a single transformer, enabling adaptive detection and recovery of occluded features within a single frame.

[0041] Zhou et al., "Towards Robust Blind Face Restoration with CodebookLookupTransformer", NeurIPS, 2022 (hereafter referred to as "Zhou") pre-trained quantized autoencoders, adopted the ViT model (see Dosovitskiy), and used self-attention to recover damaged features for blind face restoration. Although their method has similarities with ours, relying on self-attention to recover partially damaged features may fail because the attention values ​​of occluded landmarks cannot be faithfully calculated. To alleviate this problem, we develop messenger tokens and propose a module to adaptively combine regular and messenger embeddings for feature recovery.

[0042] 3. Method proposed by the present invention

[0043] This section introduces ORFormer, a general method that can be integrated into regular transformers for occlusion detection and handling. Figure 2 Our approach is described. First, we adopt the concept of vector quantization (see Aaron van den Oord et al., “Neural discrete representation learning”, NeurIPS, 2017; hereafter referred to as “Vanden Oord”), similar to the approach in Codeformer (see Zhou), and a pre-trained quantized heatmap generator (Section 3.1 of this paper). Subsequently, the learned discrete codebook and decoder are used as prior knowledge for heatmap generation. Using this prior knowledge, we utilize ORFormer to perform code sequence prediction and feature recovery for partially occluded image patches (Section 3.2). Finally, with the help of ORFormer, we integrate the heatmaps generated from the recovered features into the existing FLD method (Section 3.3).

[0044] 3.1. Quantization Heatmap Generator

[0045] To enhance the robustness of heatmap generation to occlusions, we incorporate the training of a quantized heatmap generator. By training this generator on unoccluded faces, we learn a high-dimensional latent space specifically tailored for heatmap generation under ideal conditions. The learned codebook reduces the uncertainty of recovering occluded features, as the code terms are learned from unoccluded faces.

[0046] like Figure 2 (a) shows an unobstructed face image I∈R h×w×3 Encoded by encoder E into the latent space Z∈R m×n×d Following the principle of Vector Quantised-Variational AutoEncoder (VQVAE) (see Van den Oord), each patch Z in the encoded feature Z is i,j Replaced with a learnable codebook The closest codebook item of the N codes in , i.e., code, is used to obtain the quantized feature ZQ∈R m×n×d and its corresponding code index sequence S∈{0, 1, ..., N-1} h×w ,Right now

[0047] as well as

[0048] S i,j =arg min s ||Z i,j -c s||2 (1)

[0049] Subsequently, the decoder D generates an edge heatmap based on the quantized features ZQ where N E is the number of edges (facial contours) of each face. In this paper, we adopt the same edge heatmap definition as Wu et al.

[0050] Loss Function

[0051] To train the quantized heatmap generator using a learnable codebook, the image-level loss L img is used. In addition, the intermediate latent space loss L latent is used to minimize the distance between the codebook C and the encoded features Z. These loss functions are defined as:

[0052]

[0053] in is the true edge heatmap, sg(·) represents the stopping gradient operator, and β is a hyperparameter for loss balancing. Learning codebook heatmap generator L codebook The complete loss function of is given by:

[0054] L codebook =L img +λ latent ·L latent , (3)

[0055] where λ latent is a hyperparameter used for loss balancing.

[0056] 3.2ORFormer

[0057] Given an occluded or partially invisible face as input, the traditional nearest neighbor search described in Equation (1) may fail to find the occluded patch due to the feature corruption of the occluded patch. However, relying solely on self-attention in the transformer (such as described in Zhou) is not sufficient to generate heatmaps, because the attention map computed using the corrupted features no longer faithfully captures the relationship between patches. To this end, we propose ORFormer to detect occluded patches and recover their features.

[0058] Figure 2 As shown in (b), the proposed ORFormer 100 is introduced after training the heatmap generator. ORFormer 100 obtains image patches from the features Z′ extracted by the encoder E. As input, ORFormer 100 uses the regular and messenger markers to compute patch features. It generates an occlusion map α∈R for a specific patch m×n and two code sequences, S I∈{0, 1, ..., N-1} m×n and S M ∈{0, 1, ..., N-1} m×n While SI is calculated based on regular markings and obtains information from all patches, S M It is derived from messenger markers and has occlusion awareness. I and S M The encoding in the reference codebook C generates the quantized feature Z I and Z M Based on the occlusion map α, Z is transformed into I and Z M Merge to form the restored feature Z rec Finally, Z rec Used together with the pre-trained decoder to generate the heatmap H rec .

[0059] After the pre-training phase, the codebook C and decoder D are frozen, while the encoder E is fine-tuned to facilitate heatmap generation under feature occlusion. Figure 3 The proposed ORFormer 100 is illustrated in detail as follows: ORFormer 100 comprises a self-attention module 310 , a cross-attention module 320 , an occlusion detection head module 330 , and a codebook prediction head module 340 .

[0060] like Figure 3 As shown, ORFormer 100 is a transformer with L layers. At each layer l, it computes self-attention between regular image patch labels in the following way:

[0061]

[0062] The query key Sum From patch labels X via linear embedding 1 Here, residual learning and feed-forward network (FFN) are used.

[0063] In addition to the traditional self-attention between image patch labels, a messenger label is introduced, denoted as M 1 , each patch corresponds to a messenger marker. The messenger marker is designed to simulate feature occlusion. Figure 3 As shown, only their queries is computed, each query is used to aggregate features from all tags except its corresponding patch tag via cross-attention:

[0064]

[0065] in

[0066]

[0067] Equation (6) computes the cross-attention score between the i-th messenger tag and the j-th image patch tag. By excluding features from the corresponding patch, the generated messenger tag M l+1 Encode features borrowed from other image markers and simulate feature occlusion.

[0068] Following the attention mechanism and feed-forward network, the occlusion detection head module 330 is introduced to embed X by referencing image patches. l +1 and Messenger Embedded M l+1 The occluded patch is detected by the difference between them. The occlusion map of a specific patch is obtained

[0069]

[0070] where the function dist(·,·) computes the element-wise difference between two embeddings. 1+1 is a fully connected layer that converts the embedding returned by dist into a scalar. σ(*) is a sigmoid function that ensures Between 0 and 1. The higher indicates that patch k is more likely to be occluded.

[0071] Obtain the occlusion map α at the (l-1)th layer 1 ∈R m×n After that, the messenger tag of layer l is allowed to suppress the feature aggregation of occluded patches. Specifically, the cross attention adopted by the messenger tag in formula (6) is modified as follows:

[0072]

[0073] because gives the probability of occlusion occurring in the jth image patch, so the coefficient in equation (8) Prevents the courier token from having a larger Aggregate features in patch j of value. In the first layer, the initial occlusion map α 1 is set to 0. In the last layer, i.e., the Lth layer, the generated occlusion map α L+1 It will be used for feature recovery in the next step. Figure 2 In (b), for simplicity, it is denoted as α.

[0074] exist Figure 3 In the last layer of ORFormer, the image embedding X is generated L+1 and Messenger Embedded ML+1 is input to the codebook prediction head module 340. The module 340 embeds the image X L+1 Prediction code sequence SI∈{0, 1, ..., N-1} m×n , where S I Each entry in is searched for its corresponding patch by searching the code index through formula (1). The quantized feature Z is generated by retrieving the corresponding m×n code entry from the codebook C. I ∈R m×n×d Similarly, another code sequence S M and quantitative features Z M Based on Messenger Embedding M L+1 generate.

[0075] Quantitative feature Z I and Z M Store complementary information. I considers all patches, but is sensitive to corrupted features, while Z M Then we focus on the non-occluded patches, but ignore the original patch features P, such as Figure 3 As shown. Predicted occlusion map α∈R m×n Used to merge Z in a specific patch I and Z M , to reassemble the final recovered features Z rec ∈R m×n×d ,Right now

[0076]

[0077] in represents the element-wise multiplication of A and B along the third dimension of B.

[0078] After the pre-training phase, we learn the ORFormer and fine-tune the encoder E while keeping the codebook C and decoder D fixed. We do this on S I and S M The code sequence prediction L code Using cross entropy loss:

[0079]

[0080] in The ground truth of the code sequence S comes from the pre-trained heatmap generator mentioned in Section 3.1. We also rec and The image-level loss L given in formula (2) is used between img . Learn ORFormer L ORFormer The complete loss function is:

[0081] LORFormer =L code (S I )+L code (S M )+λ img ·L img , (11)

[0082] where λ img is a hyperparameter used for loss balancing.

[0083] Integration with FLD methods

[0084] With the help of ORFormer 100 for occlusion detection and feature recovery, the quantitative heatmap generator can produce high-quality heatmaps. To evaluate the effectiveness of the output heatmaps, they are integrated into the existing FLD method as additional structural guidance. Figure 4 As shown, the integration involves merging heatmaps generated by a heatmap generator with feature maps produced by existing FLD methods. Specifically, the heatmaps are concatenated with earlier feature maps and then merged with a single lightweight CNN block. Leveraging the merged features, the proposed method can model more robust facial structure and improve the performance of existing FLD methods, especially on occluded or partially invisible faces.

[0085] Figure 5 Schematic diagram of an electronic device 500 involved in the method of the present invention. Examples of the electronic device 500 include, but are not limited to: personal computers (PCs), such as desktop computers and laptops, servers, all-in-one computers (AIOs), tablet computers and multi-function mobile phones, and wearable devices.

[0086] The electronic device 500 may include a processing circuit 510 capable of running an ORFormer framework 511 (referred to herein for simplicity). Figure 5 tagged as “ORFormer” in ), e.g. Figure 4 The entire framework shown, where Figure 2 (b) shows the entire architecture (for brevity, Figure 4 The key components are shown in the lower left portion of FIG, such as ORFormer 100) are integrated into Figure 2(a) shows the FLD architecture (i.e., the architecture of the FLD method). Processing circuit 510 may be configured to control the operation of electronic device 500. More specifically, a computer-readable medium (e.g., storage device 501) may be used to store program code 502 to be loaded onto processing circuit 510 as an ORFormer framework 511 running on processing circuit 510. When executed by processing circuit 510, program code 502 may cause processing circuit 510 to operate according to the method to perform the relevant operations of ORFormer framework 511. For example, multiple program modules may run on processing circuit 510 to control the operation of electronic device 500, wherein ORFormer framework 511 may be one of the multiple program modules, but the present invention is not limited thereto. In addition, the image input device 505 can be used to input or receive multiple input images, the RAM 520 can be used to temporarily store the multiple input images, and the ORFormer framework 511 running on the processing circuit 510 can be used to process the multiple input images. More specifically, occulusion-aware cross-attension processing is performed on any one of the multiple input images, by merging the encoded sequence S of the occluded and non-occluded image patches of the input image. I and S M The corresponding feature maps (or quantitative features Z I and Z M) obtains an occlusion map α for feature recovery to generate facial feature information related to facial feature detection, which is used to generate corresponding output images from multiple output images. The image output device 530 can be used to output or display the multiple output images, but the present invention is not limited to this. For example, the RAM 520 can be configured to temporarily store multiple input images and multiple output images, and / or the storage device 501 can be configured to store multiple input images and multiple output images. In the above embodiment, the storage device 501 can be implemented using a non-volatile memory such as a hard disk drive (HDD), a solid state drive (SSD), and a flash memory. The image input device 505 can be implemented using a camera. The processing circuit 510 can be implemented using at least one processor. The RAM 520 can be implemented using a dynamic random access memory (DRAM). The image output device 530 can be implemented using a display device such as a liquid crystal display (LCD) panel or an organic light-emitting diode (OLED) panel. The display device can be implemented using a touch panel, but the present invention is not limited to this. According to some embodiments, the architecture of electronic device 500 and / or the components therein may vary.

[0087] Figure 6 4 is a flow chart of a method according to an embodiment of the present invention. The method can be applied to an electronic device 500 and a processing circuit 510 in the electronic device 500.

[0088] In step S11, the electronic device 500 may utilize the processing circuit 510 to run the ORFormer framework 511 to perform inference using the trained model of the ORFormer framework 511 based on at least one input image (e.g., at least one image among multiple input images) to perform facial feature point detection.

[0089] More particularly, the inference comprises generating a prediction based on at least one input image (e.g., Figure 2 (b) generates at least one occlusion robust heatmap (eg, occlusion robust heatmap H rec ) for indicating at least one facial feature point of at least one face shown in the aforementioned at least one input image. For example, a multi-stage processing circuit (or processing circuit stage) for generating the aforementioned at least one occlusion robust heat map based on the aforementioned at least one input image, such as Figure 2The series of stages {210, 220, 230, 240, 250} from the encoder to the decoder shown in (b) may include an encoding stage 210 (e.g., a stage of the encoder E) and a decoding stage 250 (e.g., a stage of the decoder D), which are respectively the first and the last stages of the plurality of stages for encoding the aforementioned at least one input image into a predetermined space for reasoning (e.g., a latent space Z∈R m×n×d ), and for decoding from a predetermined space.

[0090] In step S12, during the inference process using the trained model, the processing circuit 510 (or the ORFormer framework 511 running thereon) may perform occlusion-aware cross-attention processing on any one of the at least one input image (e.g., Figure 3 Occlusion-aware cross-attention processing in the network architecture of ORFormer 100 shown in FIG, by merging two code sequences (e.g., code sequence S I and S M ) respectively correspond to the feature maps (or quantitative features Z I and Z M ) obtains the occlusion map α for feature recovery to generate facial feature point information about facial feature point detection.

[0091] More specifically, the two code sequences include a first code sequence and a second code sequence, e.g. Figure 2 (b) The code sequence S shown I and S M , the two feature maps include a first feature map and a second feature map corresponding to the first code sequence and the second code sequence respectively, wherein the first feature map and the second feature map represent a first set of quantized features and a second set of quantized features respectively, e.g. Figure 2 (b) The quantitative feature Z shown I and Z M In addition, the first code sequence, for example, the code sequence S I , is calculated from the regular marker and is arranged to bring information of multiple encoded patches (e.g., all patches) of any of the above input images, while the second code sequence, such as the code sequence SM, is derived from the messenger marker and has occlusion awareness capabilities.

[0092] based on Figure 2 (b) shows two code sequences (for example, code sequence S I and S M ) by referring to the pre-learned codebook in the training model (e.g. Figure 2 The codebook C) shown in (b) generates Figure 2(b) shows two feature maps (e.g., quantized feature Z I and Z M ). In addition, based on the occlusion map α, the occlusion map is merged in a specific patch manner. Figure 2 (b) shows two feature maps (e.g., quantized feature Z I and Z M ), and form the recovery feature Z rec , used to utilize the recovered features and the pre-trained decoder in the training model (e.g. Figure 2 (b)) generates the aforementioned at least one occlusion robust heat map (e.g., occlusion robust heat map H rec ). For example, the recovered feature Z rec By combining two feature maps (e.g. Figure 2 The quantitative feature Z shown in (b) I and Z M ) is combined with the specific patch weight given in the occlusion map α, and is used to generate at least one occlusion robust heat map mentioned above (e.g., occlusion robust heat map H rec ).

[0093] by Figure 4 The input and output images shown on the far left and far right are examples of any of the aforementioned input images and corresponding output images, respectively. The corresponding output image can be changed or modified from any of the aforementioned input images based on multiple facial feature point detection results of accurate facial feature point detection, wherein the aforementioned at least one occlusion robust heat map, such as the occlusion robust heat map H rec , which may include multiple facial feature point detection results. The aforementioned multi-stage processing circuit (or processing circuit stage) for inference, such as Figure 2 The series of stages {210, 220, 230, 240, 250} from encoder to decoder shown in (b) can be arranged to act as a heatmap generator in the ORFormer framework 511, e.g. Figure 4 The entire framework shown in is used for occlusion detection and feature recovery, thereby generating at least one occlusion robust heat map, such as the occlusion robust heat map H rec By means of a heatmap generator integrated in the ORFormer framework 511, the ORFormer framework 511 may be arranged to refer to the aforementioned at least one input image and the aforementioned at least one occlusion robust heatmap, e.g. the occlusion robust heatmap H rec , to generate the at least one facial landmark of the at least one face shown in the at least one input image.

[0094] Despite the challenge of corrupted feature maps of non-visible regions, the method of the present invention can utilize ORFormer100 to identify non-visible regions and restore their missing features through messenger labels Mi, as described above. Figure 1 (a), for each patch P i , the processing circuit 510 (or the ORFormer framework 511 running thereon) may introduce a patch mark X i and a learnable messenger marker M i For occlusion detection and processing; refer to Figure 1 (b), the messenger tags may be arranged to compute attention using patch tags in addition to their corresponding patch tags; and reference Figure 1 (c), the processing circuit 510 (or the ORFormer framework 511 running thereon) can detect occlusions by evaluating the dissimilarity between the regular embedding X′i and the messenger embedding M′i, and then if i If there is occlusion in , the occluded features are recovered based on the messenger embedding aggregated from other image patches.

[0095] In the process of generating the at least one occlusion robust heat map based on the at least one input image, the processing circuit 510 (or the ORFormer framework 511 running thereon) may perform related operations with the help of the ORFormer 100, for example, the ORFormer framework 511 (or the ORFormer 100 therein) may use at least one messenger marker M i To simulate the occlusion in at least one patch i, and to aggregate at least one patch mark X corresponding to the aforementioned at least one patch i i For detection, the ORFormer framework 511 (or the ORFormer 100 therein, which can be viewed as the ORFormer stage 220, as at least one next / subsequent stage of the encoding stage 210) takes the image patch P as input and generates an occlusion map α and two code sequences S I and S M , and the ORFormer framework 511 (or the combination of the codebook C and its access control circuit, collectively referred to as the codebook circuit, which can be regarded as the first subsequent stage 230, as the next stage of the ORFormer stage 220) from two code sequences S I and S M Decode two feature maps, such as the quantized feature Z I and Z M For restoration, the ORFormer framework 511 (or the feature restoration module 12 therein, which can be regarded as the second subsequent stage 240, serving as the next stage of the first subsequent stage 230) converts two feature maps (e.g., quantized feature ZI and Z M ) is combined with the patch-specific weights given in the occlusion map α to produce the recovered features Z rec , and the ORFormer framework 511 (or the decoder D therein, which can be viewed as the decoder stage 250, serving as the next stage of the second subsequent stage 240) recovers the feature Z from rec Generate the aforementioned at least one occlusion robust heat map (eg, occlusion robust heat map H rec ).

[0096] like Figure 6 As shown, the loop of steps S11 and S12 can be executed multiple times, respectively for processing each input image to generate a corresponding output image. For the sake of brevity, similar descriptions of this embodiment are not repeated here.

[0097] For ease of understanding, this method can be Figure 6 The workflow shown in FIG. 1 is used as an example for explanation, but the present invention is not limited thereto. According to some embodiments, Figure 6 Add, delete, or change one or more steps in the workflow shown. For example, after completing Figure 6 After completing steps S11 and S12 in the current iteration of the illustrated workflow, when re-entering step S11 in another iteration, the processing circuit 510 can selectively change the image source of the input images to obtain another set of input images as the input images for another iteration, and begin inference with the trained model based on the latest input images to generate another set of output images as the latest output images for another iteration. For the sake of brevity, similar descriptions of these embodiments will not be repeated in detail here.

[0098] According to some embodiments, since the quantization heat map generator needs to be trained first, e.g. Figure 2 (a), so the training / pre-training steps of the trained model can be inserted into Figure 6 In the workflow shown, as part of the workflow before steps S11 and S12. More specifically, the method may include: before executing steps S11 and S12 (or a loop thereof), pre-training a quantized heat map generator so as to learn prior knowledge from the original data set to train the codebook C and the decoder D, wherein the quantized heat map generator can take the image I as input through the encoder E and generate its edge heat map H, and after pre-training, the prior knowledge of the unoccluded face is encoded into the codebook C and the decoder D. For the occluded face, the trained model can be further trained with artificial occlusion to generate the occlusion map α and the two code sequences S I and S M , see Figure 2(b) Using the frozen codebook C and decoder D (respectively marked with snowflake symbols to indicate their frozen states for ease of understanding), the processing circuit 510 (or the ORFormer framework 511 running thereon) can use the ORFormer 100 to generate the occlusion map α and the two code sequences S I and S M , thus obtaining the quantitative feature Z I and Z M , where the restored feature Z rec By taking the quantized feature Z I and Z M is combined with the weights of the specific patch given in the occlusion map α, and used to recover the feature Z from rec Generate occlusion robust heatmap H rec See also Figure 3 , the processing circuit 510 (or the ORFormer framework 511 running thereon) may utilize the ORFormer 100 to take the image patch P as input to the ORFormer 100 and generate two code sequences S via the codebook prediction head 340 I and S M , where the code sequence S I The code sequence S can be calculated by referring to the image patch mark. M The occlusion map α can be calculated from the messenger markers and can represent the occlusion probability of a particular patch and can be inferred by the occlusion detection head module 330. For example, the number of layers L in the ORFormer 100 can be a positive integer greater than (especially much greater than) one, and the layer index l of each layer l can be an integer falling within the interval [0, (L-1)]. For the input of layer l in the case of l=0, X l =X 0 =P,M l =M 0 , where M 0 Can be random, α l =α 0 =0 (may indicate “visible”). In the above reasoning process, the processing circuit 510 (or the ORFormer framework 511 running thereon) may perform codebook retrieval, using two code sequences S I and S M Select quantized feature Z from codebook C I and Z M , to perform feature recovery, as shown in formula (9). Figure 4The processing circuit 510 (or the ORFormer framework 511 running thereon) can utilize the ORFormer 100 to perform occlusion detection and feature recovery, thereby obtaining a high-quality heat map, wherein the generated heat map serves as an additional input to the FLD architecture and provides the recovered features to make the FLD architecture robust to occlusion. For the sake of brevity, similar descriptions of these embodiments will not be repeated in detail here.

[0099] Those skilled in the art will readily observe that many modifications and variations can be made to the apparatus and method while retaining the teachings of the present invention.Accordingly, the above disclosure should be construed as being limited only by the metes and bounds of the appended claims.

Claims

1. An occlusion-robust transformation method for accurately detecting facial feature points, the method being applied to a processing circuit within an electronic device, the method comprising: Running an occlusion robust transformer framework using the processing circuitry, and starting inference using a trained model of the occlusion robust transformer framework based on at least one input image to perform facial landmark detection; as well as During inference using the trained model, occlusion-aware cross-attention processing is performed on any input image of the at least one input image, and an occlusion map for feature recovery is obtained by merging two feature maps of two code sequences corresponding to occluded and non-occluded image patches of any input image, thereby generating facial feature point information for facial feature point detection.

2. The occlusion robust transformation method for accurately detecting facial feature points according to claim 1, wherein: The inference includes inference for generating at least one occlusion robust heatmap from the at least one input image, indicating at least one facial landmark of at least one face shown in the at least one input image.

3. The occlusion robust transformation method for accurately detecting facial feature points according to claim 2, wherein: The multiple stage processing circuitry for the inference for generating the at least one occlusion robust heat map based on the at least one input image includes an encoding stage and a decoding stage, wherein the encoding stage and the decoding stage are respectively the first and last stages of the multiple stages, and are respectively used to encode the at least one input image into a predetermined space for the inference and to decode from the predetermined space.

4. The occlusion robust transformation method for accurately detecting facial feature points according to claim 1, wherein: The two code sequences include a first code sequence and a second code sequence, and the two feature maps include a first feature map and a second feature map corresponding to the first code sequence and the second code sequence, respectively.

5. The occlusion robust transformation method for accurately detecting facial feature points according to claim 4, wherein: The first feature map and the second feature map represent a first set of quantized features and a second set of quantized features, respectively.

6. The occlusion robust transformation method for accurately detecting facial feature points according to claim 1, wherein: The two code sequences include a first code sequence and a second code sequence, wherein the first code sequence is calculated by conventional labeling and is arranged to bring information of multiple encoded patches of any input image, and the second code sequence is derived by messenger labeling and has occlusion perception capability.

7. The occlusion robust transformation method for accurately detecting facial feature points according to claim 1, wherein: Based on the codes in the two code sequences, two feature maps are generated with reference to the pre-learned codebooks in the training model.

8. The occlusion robust transformation method for accurately detecting facial feature points according to claim 7, wherein: The reasoning includes generating at least one occlusion robust heat map based on the at least one input image; and merging the two feature maps in a patch-specific manner based on the occlusion map to form a restored feature for generating at least one occlusion robust heat map using the restored feature and a pre-trained decoder in the training model.

9. The occlusion robust transformation method for accurately detecting facial feature points according to claim 8, wherein: The recovered features are generated by merging the two feature maps with patch-specific weights given in the occlusion map, and are used to generate at least one occlusion robust heatmap.

10. The occlusion robust transformation method for accurately detecting facial feature points according to claim 1, wherein: The multi-stage processing circuitry with respect to said inference is arranged to act as a heatmap generator within said occlusion robust transformer framework for occlusion detection and feature recovery, thereby generating at least one occlusion robust heatmap; And, by means of the heat map generator integrated into the occlusion robust transformer framework, the occlusion robust transformer framework is arranged to generate at least one facial landmark of at least one face shown in at least one input image with reference to the at least one input image and the at least one occlusion robust heat map.

11. The occlusion robust transformation method for accurately detecting facial feature points according to claim 1, wherein: The inference comprises inference for generating at least one occlusion robust heatmap based on the at least one input image; the two code sequences comprise a first code sequence and a second code sequence corresponding to a conventional marker and a messenger marker, respectively; And the method further comprises: During the inference process of generating the at least one occlusion robust heatmap based on the at least one input image, at least one messenger marker is used to model occlusion present in at least one patch, and features of all patch markers except at least one patch marker corresponding to the at least one patch are aggregated.

12. The occlusion robust transformation method for accurately detecting facial feature points according to claim 1, wherein: The reasoning includes generating at least one occlusion robust heatmap based on the at least one input image; the method further includes: In the process of generating the at least one occlusion robust heat map based on the at least one input image by reasoning, taking a plurality of image patches as input, generating the occlusion map and the two code sequences; In the process of inferring and generating at least one occlusion robust heat map based on the at least one input image, the two feature maps are decoded from the two code sequences.

13. The occlusion robust transformation method for accurately detecting facial feature points according to claim 1, wherein: The inference includes inferring and generating at least one occlusion robust heatmap based on the at least one input image; And the method further comprises: In the process of generating the at least one occlusion robust heatmap by reasoning based on the at least one input image, merging the two feature maps with the specific patch weights given in the occlusion map to produce a recovered feature; as well as In the process of generating at least one occlusion robust heatmap by reasoning based on the at least one input image, the at least one occlusion robust heatmap is produced from the restored features.

14. An apparatus according to claim 1, wherein: The apparatus includes at least a processing circuit within the electronic device.

15. The device according to claim 14, wherein The apparatus includes the electronic device.

16. A computer readable medium associated with the method of claim 1, wherein: The computer-readable medium stores program code that, when executed by the processing circuit, causes the processing circuit to operate according to the method.