Method and system for using specific local refinement of facial components in facial landmark detection

The method enhances facial landmark detection accuracy by refining specific facial components through cascaded regression and component-specific feature extraction, addressing inefficiencies in existing methods and deep learning approaches.

CN113924603BActive Publication Date: 2025-07-15GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080041024.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-11
Filing Date
2020-05-21
Publication Date
2025-07-15
Estimated Expiration
2040-05-21

AI Technical Summary

Technical Problem

The prior art is difficult to effectively refine facial components in facial landmark detection, resulting in insufficient detection accuracy and accuracy, especially when facial image posture changes.

Method used

The specific local refinement method of facial components is adopted. By defining multiple facial components specific local areas, the cascade regression method is used to extract local features and organize facial landmark sets, and the position of facial landmarks is gradually refined to reduce interference to other facial components.

Benefits of technology

The accuracy and robustness of facial landmark detection are improved, especially when facial posture changes, and the detection speed is improved without increasing calculation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113924603B_ABST
    Figure CN113924603B_ABST
Patent Text Reader

Abstract

A method includes: receiving a facial image (204); obtaining a facial shape (206) using the facial image (204); defining a plurality of face-component specific local regions using the facial image (204) and the facial shape (206), wherein each of the plurality of face-component specific local regions includes a respective separately considered face component from a plurality of separately considered face components of the facial image (204), and the respective separately considered face component of the plurality of separately considered face components corresponds to a respective first facial landmark set (208) of a plurality of first facial landmark sets in the facial shape (206); for each of the plurality of face-component specific local regions, performing a cascade regression method using each of the plurality of face-component specific local regions and a respective facial landmark set (208) of the plurality of first facial landmark sets to obtain a respective facial landmark set (210) of a plurality of second facial landmark sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of facial landmark detection, and more particularly, to a method and system for facial landmark detection using specific local refinement of facial components. Background Art

[0002] Facial landmark detection plays a crucial role in facial recognition, facial animation, 3D facial reconstruction, virtual makeup, etc. The goal of facial landmark detection is to locate a plurality of fiducial facial key points around a plurality of facial components and a plurality of facial contours in a plurality of facial images. Summary of the Invention

[0003] An object of the present disclosure is to propose a method and system for facial landmark detection using specific local refinement of facial components.

[0004] In a first aspect of the present disclosure, a computer-implemented method includes: a method for an inference stage, where the method for the inference stage includes: receiving a first facial image; obtaining a first facial shape using the first facial image; defining a plurality of face component-specific local regions using the first facial image and the first facial shape, where each of the plurality of face component-specific local regions includes a corresponding separately considered face component from a plurality of separately considered face components in the first facial image, and the corresponding separately considered face component in the plurality of separately considered face components corresponds to a corresponding first facial landmark set in a plurality of first facial landmark sets in the first facial shape, where the corresponding first facial landmark set in the plurality of first facial landmark sets includes a plurality of facial landmarks; for each of the plurality of face component-specific local regions, performing a cascade regression method using each of the plurality of face component-specific local regions and a corresponding facial landmark set in the plurality of first facial landmark sets to obtain a corresponding facial landmark set in a plurality of second facial landmark sets. Each stage of the cascade regression method includes: extracting a plurality of local features using each of the plurality of face component-specific local regions and a corresponding facial landmark set in a plurality of previous stage facial landmark sets, where the extracting step includes extracting each of the plurality of local features from a face landmark-specific local region around a corresponding facial landmark in the corresponding facial landmark set in the plurality of previous stage facial landmark sets, where the face landmark-specific local region is in each of the plurality of face component-specific local regions; and the corresponding facial landmark set in the plurality of previous stage facial landmark sets corresponding to the initial stage of the cascade regression method is the corresponding facial landmark set in the plurality of first facial landmark sets; and organizing the plurality of local features based on a plurality of correlations between the plurality of local features to obtain a corresponding facial landmark set in a plurality of current stage facial landmark sets, where the corresponding facial landmark set in the plurality of current stage facial landmark sets corresponding to the last stage of the cascade regression method is the corresponding facial landmark set in the plurality of second facial landmark sets.

[0005] In a second aspect of the present disclosure, a system includes at least one memory and at least one processor. The at least one memory is configured to store a plurality of program instructions. The at least one processor is configured to execute the plurality of program instructions, and the plurality of program instructions cause the at least one processor to perform a plurality of steps, including: performing a method of an inference stage, wherein the method of the inference stage includes: receiving a first facial image; obtaining a first facial shape using the first facial image; defining a plurality of face component-specific local regions using the first facial image and the first facial shape, wherein each of the plurality of face component-specific local regions includes a corresponding separately considered face component from a plurality of separately considered face components in the first facial image, and the corresponding separately considered face component in the plurality of separately considered face components corresponds to a corresponding first facial landmark set in a plurality of first facial landmark sets in the first facial shape, wherein the corresponding first facial landmark set in the plurality of first facial landmark sets includes a plurality of facial landmarks; for each of the plurality of face component-specific local regions, performing a cascaded regression method using each of the plurality of face component-specific local regions and a corresponding facial landmark set in the plurality of first facial landmark sets to obtain a corresponding facial landmark set in a plurality of second facial landmark sets. Each stage of the cascaded regression method includes: extracting a plurality of local features using each of the plurality of face component-specific local regions and a corresponding facial landmark set in a plurality of previous stage facial landmark sets, wherein the extracting step includes extracting each of the plurality of local features from a facial landmark-specific local region around a corresponding facial landmark in the corresponding facial landmark set in the plurality of previous stage facial landmark sets, wherein the facial landmark-specific local region is in each of the plurality of face component-specific local regions; and the corresponding facial landmark set in the plurality of previous stage facial landmark sets corresponding to the initial stage of the cascaded regression method is the corresponding facial landmark set in the plurality of first facial landmark sets; and organizing the plurality of local features based on a plurality of correlations between the plurality of local features to obtain a corresponding facial landmark set in a plurality of current stage facial landmark sets, wherein the corresponding facial landmark set in the plurality of current stage facial landmark sets corresponding to the last stage of the cascaded regression method is the corresponding facial landmark set in the plurality of second facial landmark sets. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] In order to more clearly illustrate the embodiments of the present invention or related technologies, the following drawings will be described when briefly introducing the embodiments. Obviously, the drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without paying on the premise of these drawings.

[0007] Figure 1It is a block diagram illustrating input, processing, and output hardware modules in a terminal according to an embodiment of the present disclosure.

[0008] Figure 2 It is a block diagram illustrating a facial landmark detector according to an embodiment of the present disclosure.

[0009] Figure 3 It is a diagram illustrating 68 numbered facial landmarks among many examples to be referred to in the present disclosure.

[0010] Figure 4 It is illustrated by a block diagram of a full-face landmark acquisition module in the Figure 2 facial landmark detector according to an embodiment of the present disclosure.

[0011] Figure 5 It is illustrated by a block diagram of a cropping module in the Figure 2 facial landmark detector according to an embodiment of the present disclosure.

[0012] Figure 6 It is illustrated by a block diagram of a plurality of face component-specific local refinement modules in the Figure 2 facial landmark detector according to an embodiment of the present disclosure.

[0013] Figure 7 It is illustrated by a block diagram of a merging module in the Figure 2 facial landmark detector according to an embodiment of the present disclosure.

[0014] Figure 8 It is illustrated by a block diagram of a cropping module in the Figure 2 facial landmark detector according to another embodiment of the present disclosure.

[0015] Figure 9 It is illustrated by a block diagram of a cropping module in the Figure 2 facial landmark detector according to an embodiment of the present disclosure.

[0016] Figure 10 It is illustrated by a block diagram of a plurality of cascade regression stages in one of the Figure 6 plurality of face component-specific local refinement modules according to an embodiment of the present disclosure.

[0017] Figure 11 It is illustrated by a block diagram of a local feature extraction module and a local feature organization module in each of the Figure 10 plurality of cascade regression stages according to an embodiment of the present disclosure.

[0018] Figure 12A FIG. illustrates a block diagram of a plurality of face landmark - specific local feature mapping functions used in the local feature extraction module at the initial stage in the plurality of cascaded regression stages (in Figure 10 ). Figure 11 ).

[0019] Figure 12B FIG. illustrates a block diagram of one of the plurality of face landmark - specific local feature mapping functions implemented by a random forest in Figure 12A according to an embodiment of the present disclosure.

[0020] Figure 13 FIG. illustrates a block diagram of a local feature concatenation module, a face component - specific projection module, and a face landmark set increment module in the local feature organization module in Figure 11 according to an embodiment of the present disclosure.

[0021] Figure 14 FIG. illustrates a block diagram of a plurality of cascaded regression stages (a plurality of cascaded training stages) in Figure 10 according to an embodiment of the present disclosure.

[0022] Figure 15 FIG. illustrates a block diagram of a face landmark - specific local feature mapping function training module and a face component - specific projection matrix training module in one of the plurality of cascaded training stages in Figure 14 according to an embodiment of the present disclosure.

[0023] Figure 16 FIG. illustrates a block diagram of a joint detection module implementing the global face landmark acquisition module in Figure 4 according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Hereinafter, various technical matters, various structural features, various implementation purposes, and various effects of the present disclosure in many embodiments will be described in detail with reference to the accompanying drawings. Specifically, the terms in the embodiments of the present disclosure are only used for the purpose of explaining a certain embodiment and do not limit the invention.

[0025] The same reference numerals in different drawings indicate substantially the same elements, and the description of one element is applicable to other elements.

[0026] As used herein, an apparatus, an element, a method, or a step that is employed, as described by terms such as "use" or "from", means that the apparatus, the element, the method, or the step is directly utilized or indirectly utilized through an intervening apparatus, an intervening element, an intervening method, or an intervening step.

[0027] As used herein, a term "acquire" used in cases such as "acquire A" means receiving "A" or outputting "A" after an operation.

[0028] Figure 1 FIG. is a block diagram illustrating input, processing, and output hardware modules in a terminal 100 according to an embodiment of the present disclosure. Referring to Figure 1 , the terminal 100 includes a camera module 102, a processor module 104, a memory module 106, a display module 108, a storage module 110, a wired or wireless communication module 112, and a plurality of buses 114. In one embodiment, the terminal 100 may be a mobile phone, a smartphone, a tablet computer, a laptop computer, a desktop computer, or any electronic device having sufficient computing power for facial landmark detection.

[0029] The camera module 102 is an input hardware module for capturing a facial image 204 (marked in Figure 2 ), which is transmitted to the processor module 104 through the plurality of buses 114. In one embodiment, the camera module 102 includes an RGB camera or a grayscale camera. In another embodiment, the facial image 204 may be acquired using another input hardware module, such as the storage module 110, or the wired or wireless communication module 112. The storage module 110 is used to store the facial image 204 transmitted to the processor module 104 through the plurality of buses 114. The wired or wireless communication module 112 is used to receive the facial image 204 from a network through wired or wireless communication, and the facial image 204 is transmitted to the processor module 104 through the plurality of buses 114.

[0030] The memory module 106 stores a plurality of inference stage program instructions, which are executed by the processor module 104, enabling the processor module 104 to execute a method of an inference stage for facial landmark detection using the specific local refinement of the facial components to generate a facial shape 206 (marked in Figure 2 ), and the facial shape 206 will be described with reference to Figures 2 to 13 . In one embodiment, the memory module 106 may be a transient or non-transient computer-readable medium, and the transient or non-transient computer-readable medium includes at least one memory. The processor module 104 includes at least one processor, and the processor directly or indirectly sends a plurality of signals to the digital camera module 102, the memory module 106, the display module 108, the storage module 110, and the wired or wireless communication module 112 via the plurality of buses 114 and / or directly or indirectly receives a plurality of signals from the digital camera module 102, the memory module 106, the display module 108, the storage module 110, and the wired or wireless communication module 112. The at least one processor may be one or more central processing units (CPUs), one or more graphics processing units (GPUs), and / or one or more digital signal processors (DSPs). The one or more CPUs may send the content image 204, some of the plurality of program instructions, and other data or instructions to the one or more GPUs and / or the one or more DSPs via the plurality of buses 114.

[0031] The display module 108 is an output hardware module and is configured to display the facial shape 206 on the facial image 204, or an application result obtained using the facial shape 206 on the facial image 204 received from the processor module 104 via the plurality of buses 114. The application result can come from, for example, facial recognition, facial animation, 3D facial reconstruction, and applying virtual makeup. In another embodiment, the facial shape 206 on the facial image 204, or the application result obtained using the facial shape 206 on the facial image 204, can be output using other output hardware modules such as the storage module 110 or the wired or wireless communication module 112. The storage module 110 is configured to store the facial shape 206 on the facial image 204, or the application result obtained using the facial shape 206 on the facial image 204 received from the processor module 104 via the plurality of buses 114. The wired or wireless communication module 112 is configured to transmit the facial shape 206 on the facial image 204 or the application result obtained using the facial shape 206 on the facial image 204 to the network via the wired or wireless communication, where the facial shape 206 on the facial image 204 or the application result obtained using the facial shape 206 on the facial image 204 is received from the processor module 104 via the plurality of buses 114.

[0032] In one embodiment, the memory module 106 also stores training stage program instructions, which are executed by the processor module 104, enabling the processor module 104 to use facial component specific local refinement for a method of the training stage of a facial landmark detection, which will be described with reference to Figures 14 to 15 being described.

[0033] In the above embodiment, the terminal 100 is a computing system, and all components of the computing system are integrated together via the plurality of buses 114. Other types of computing systems, such as a computing system having a remote camera module instead of the camera module 102, are also within the scope of this disclosure. In the above embodiment, the memory module 106 and the processor module 104 of the terminal 100 store and execute a plurality of inference stage program instructions and a plurality of training stage program instructions respectively. Other types of computing systems, such as a computing system that includes different terminals, which are respectively used for a plurality of inference stage program instructions and a plurality of training stage program instructions, are also within the scope of this disclosure.

[0034] Figure 2FIG. 0 is a block diagram illustrating a facial landmark detector 202 according to an embodiment of the present disclosure. The facial landmark detector 202 is configured to receive a facial image 204, perform a method of an inference phase of facial landmark detection using facial component-specific local refinement, and output a facial shape 206. The facial shape 206 includes a plurality of facial landmarks. The facial shape 206 is shown on the facial image 204 to indicate the plurality of positions of the plurality of facial landmarks relative to the plurality of facial components and a facial contour in the facial image 204. Throughout the present disclosure, for similar reasons, a plurality of facial landmarks are shown on a plurality of facial images. In one example, a number of the facial landmarks is sixty-eight. Figure 3 FIG. 1 is a diagram illustrating sixty-eight numbered facial landmarks of a plurality of facial landmarks to be referred to in many examples to be referred to in the present disclosure. Refer to Figure 2 and 3 , a facial landmark 208 among the plurality of facial landmarks is the facial landmark (17) of the facial shape 206, and the facial landmark 210 among the plurality of facial landmarks is the facial landmark (24) of the facial shape 206. The plurality of facial landmarks are separated into a first set obtained by a global facial landmark obtaining module 402 in Figure 4 and a second set obtained by a plurality of facial component-specific local refining modules 602 to 608 in Figure 6 . Each facial landmark in the first set is indicated by a dot style used by the facial landmark 208, while each facial landmark in the second set is indicated by a dot style used by the facial landmark 210.

[0035] The facial landmark detector 202 includes the global facial landmark obtaining module 402 described with reference to Figure 4 , a cropping module 502 described with reference to Figure 5 , the plurality of facial component-specific local refining modules 602 to 608 described with reference to Figure 6 , and a merging module 702 described with reference to Figure 7 .

[0036] Figure 4 FIG. 25 is a diagram illustrating, according to an embodiment of the present disclosure, in Figure 2A block diagram of the global facial landmark acquisition module 402 in the facial landmark detector 202 in []. The global facial landmark acquisition module 402 is configured to receive the facial image 204 and use the facial image 204 to obtain a facial shape 406. Refer to Figure 3 and Figure 4 , in one embodiment, the facial shape 406 includes a plurality of facial landmarks (1) to (68) globally used for a face (i.e., for the entire face) in the facial image 204. The plurality of facial landmarks (1) to (68) in the facial shape 406 are the plurality of facial landmarks (1) to (17) for the facial contour in the facial image 204, the plurality of facial landmarks (18) to (27) for the plurality of eyebrows in the facial image 204, the plurality of facial landmarks (37) to (48) for the plurality of eyes in the facial image 204, the plurality of facial landmarks (28) to (36) for a nose in the facial image 204, and the plurality of facial landmarks (49) to (68) for a mouth in the facial image 204.

[0037] Figure 5 is a diagram illustrating in accordance with an embodiment of the present disclosure in Figure 2 A block diagram of the cropping module 502 in the facial landmark detector 202 in []. The cropping module 502 is configured to use the facial image 204 and the facial shape 406 to define a plurality of face component-specific local regions 504 to 510. Each of the plurality of face component-specific local regions 504 to 510 includes a corresponding separately considered face component 520, 524, 528, or 532 from the plurality of separately considered face components 520, 524, 528, and 532 in the facial image 204. In one embodiment, the plurality of separately considered face components 520, 524, 528, and 532 are separated according to a plurality of facial features 522, 526, 530, and 534. In Figure 5 the embodiment of [], the plurality of facial features 522, 526, 530, and 534 are functionally grouped. The facial feature 522 is two eyebrows in the plurality of face component-specific local regions 504. The facial feature 526 is two eyes in the plurality of face component-specific local regions 506. The facial feature 530 is a nose in the plurality of face component-specific local regions 508. The facial feature 534 is a mouth in the plurality of face component-specific local regions 504. The two eyebrows are functionally grouped because, for example, they both have the function of keeping rain and sweat out of the two eyes. The two eyes are functionally grouped because, for example, they work together to provide vision.

[0038] The respective separately considered facial components 520, 524, 528, and 532 among the multiple separately considered facial components correspond to a respective facial landmark set 512, 514, 516, or 518 among the multiple facial landmark sets 512 to 518 in the facial shape 406. The respective facial landmark sets 512, 514, 516, or 518 among the multiple facial landmark sets 512 to 518 include multiple facial landmarks. For example, referring to Figure 3 and Figure 5 as shown, the facial landmark set 512 among the multiple facial landmark sets 512 to 518 includes the multiple facial landmarks (18) to (27) of the facial shape 406. The facial landmark set 514 among the multiple facial landmark sets 512 to 518 includes the multiple facial landmarks (37) to (48) of the facial shape 406. The facial landmark set 516 among the multiple facial landmark sets 512 to 518 includes the multiple facial landmarks (28) to (36) of the facial shape 406. The facial landmark set 518 among the multiple facial landmark sets 512 to 518 includes the facial landmarks (49) to (68) of the facial shape 406.

[0039] After the full-face landmark acquisition module 402 outputs the facial shape 406 including the multiple facial landmarks (18) to (27) known for identifying the eyebrow position in the facial image 204, the multiple facial landmarks (37) to (48) known for identifying the eye position in the facial image 204, the multiple facial landmarks (28) to (36) known for identifying the nose position in the facial image 204, and the multiple facial landmarks (49) to (68) known for identifying the mouth position in the facial image 204, the cropping module 502 can use the facial shape 406 to define the multiple facial component specific local regions 504 to 510.

[0040] In one embodiment as shown in Figure 5 the defining step includes defining each of the multiple facial component specific local regions 504 to 510 by cropping such that multiple separately considered facial components (524, 528, 532), (520, 528, 532), (520, 524, 532), or (520, 524, 528) other than the respective separately considered facial components 520, 524, 528, or 532 among the multiple separately considered facial components are at least partially removed. The multiple facial landmark sets 512 to 518 are correspondingly located on the separated multiple facial component specific local regions 504 to 510.

[0041] In the above embodiment, the defining step includes defining each of the plurality of specific local regions 504 to 510 of the facial components by cropping. Accordingly, the plurality of sets of facial landmarks 512 to 518 are located on the separated plurality of specific local regions 504 to 510 of the facial components. Other ways of defining each of the plurality of specific local regions of the facial components, such as using the coordinates of a plurality of corresponding corners of each of the plurality of specific local regions of the facial components in a facial image to define a corresponding boundary of each of the plurality of specific local regions of the facial components in the facial image, are within the scope of the present disclosure. Accordingly, the plurality of sets of facial landmarks are located on the plurality of specific local regions of the facial components, all of which are in the facial image. In the above embodiment, the shape of each of the plurality of specific local regions 504 to 510 of the facial components is a rectangle. Other shapes of any specific local region of the facial components, such as a circle, are within the scope of the present disclosure.

[0042] Figure 6 is a diagram illustrating Figure 2 a block diagram of the plurality of specific local refinement modules 602 to 608 in the facial landmark detector 202 according to an embodiment of the present disclosure. For each of the plurality of specific local regions 504 to 510 of the facial components, a corresponding specific local refinement module 602, 604, 606, or 608 of the plurality of specific local refinement modules 602 to 608 is configured to receive each of the plurality of specific local regions 504 to 510 of the facial components and perform a cascaded regression method using each of the plurality of specific local regions 504 to 510 of the facial components and a corresponding set of facial landmarks 512, 514, 516, or 518 of the plurality of sets of facial landmarks 512 to 518 to obtain a corresponding set of facial landmarks 618, 620, 622, or 624 of the plurality of sets of facial landmarks 618 to 624. Many details of an exemplary one of the plurality of specific local refinement modules 602 to 608 will be described with reference to Figures 10 to 13 be described.

[0043] Figure 7 is a diagram illustrating Figure 2A block diagram of the merging module 702 in the facial landmark detector 202 in. The merging module 702 is configured to receive the multiple sets of facial landmarks 618 to 624 and a set of facial landmarks 704 in the facial shape 406, and merge the multiple sets of facial landmarks 618 to 624 located on the separate multiple facial component specific local regions 504 to 510 and the set of facial landmarks 704 in the facial shape 406 into a facial shape 206. The set of facial landmarks 704 corresponds to the facial contour in the facial image 204 and includes the multiple facial landmarks (1) to (17) in the facial shape 406.

[0044] In the above embodiment, the defining step includes defining each of the multiple facial component specific local regions 504 to 510 by cropping. The merging step includes merging the multiple sets of facial landmarks 618 to 624 located on the separate multiple facial component specific local regions 504 to 510. For other ways of defining each of the multiple facial component specific local regions by defining the corresponding boundaries of the multiple facial component specific local regions in the facial image, the multiple sets of facial landmarks are located on the multiple facial component specific local regions in the facial image accordingly. Therefore, the merging step may not be necessary.

[0045] Figure 8 is a diagram illustrating in accordance with another embodiment of the present disclosure in Figure 2 a block diagram of a cropping module 802 in the facial landmark detector 202 in. Compared with Figure 5 the cropping module 502 in, the cropping module 802 is configured to use the facial image 204 and the facial shape 406 to define multiple facial component specific local regions 804 to 814. Each of the multiple facial component specific local regions 804 to 814 includes a corresponding separately considered facial component 828, 832, 836, 840, 844, or 848 from the multiple separately considered facial components 828, 832, 836, 840, 844, and 848 in the facial image 204. In one embodiment, the multiple separately considered facial components 828, 832, 836, 840, 844, and 848 are separated according to the multiple facial features 830, 834, 838, 842, 846, and 850. In Figure 8In an embodiment, the multiple facial features 828, 832, 836, 840, 844, and 848 are grouped non - functionally. The facial feature 830 is a left eyebrow in one of the multiple facial component - specific local regions 804. The facial feature 834 is a right eyebrow in one of the multiple facial component - specific local regions 806. The facial feature 838 is a left eye in one of the multiple facial component - specific local regions 808. The facial feature 842 is a right eye in one of the multiple facial component - specific local regions 810. The facial feature 846 is a nose in one of the multiple facial component - specific local regions 812. The facial feature 850 is a mouth in one of the multiple facial component - specific local regions 814.

[0046] The respective separately - considered facial components 828, 832, 836, 840, 844, or 848 among the multiple separately - considered facial components 828, 832, 836, 840, 844, and 848 correspond to a respective facial landmark set 816, 818, 820, 822, 824, or 826 among the multiple facial landmark sets 816 to 826 in the facial shape 406. The respective facial landmark set 816, 818, 820, 822, 824, or 826 among the multiple facial landmark sets 816 to 826 includes multiple facial landmarks. Refer to Figure 3 and Figure 8 , for example, the facial landmark set 816 among the multiple facial landmark sets 816 to 826 includes the multiple facial landmarks (18) to (22) of the facial shape 406. The facial landmark set 818 among the multiple facial landmark sets 816 to 826 includes the multiple facial landmarks (23) to (27) of the facial shape 406. The facial landmark set 820 among the multiple facial landmark sets 816 to 826 includes the multiple facial landmarks (37) to (40) of the facial shape 406. The facial landmark set 822 among the multiple facial landmark sets 816 to 826 includes the multiple facial landmarks (43) to (46) of the facial shape 406. The facial landmark set 824 among the multiple facial landmark sets 816 to 826 includes the multiple facial landmarks (28) to (36) of the facial shape 406. The facial landmark set 826 among the multiple facial landmark sets 816 to 826 includes the multiple facial landmarks (49) to (68) of the facial shape 406. The remaining description of the facial landmark detector 202 including the cropping module 502 can be applied mutatis mutandis to the facial landmark detector 202 including the cropping module 802.

[0047] Figure 9 is illustrated in accordance with an embodiment of the present disclosure in Figure 2A block diagram of the cropping module 902 in the facial landmark detector 202 in. Compared with Figure 5 the cropping module 502 in, the cropping module 902 is configured to use the facial image 204 and the facial shape 406 to define a plurality of face component-specific local regions 904 to 908. Each of the plurality of face component-specific local regions 904 to 908 includes a corresponding separately considered face component 916, 920, or 924 from among the plurality of separately considered face components 916, 920, and 924 in the facial image 204. In Figure 9 one embodiment of, the plurality of separately considered face components 916, 920, and 924 are separated according to a plurality of senses. The separately considered face component 916 is a vision-related sensory component 918 and is the two eyebrows and two eyes in the plurality of face component-specific local regions 904. The separately considered face component 920 is a smell-related sensory component 922 and is a nose in the plurality of face component-specific local regions 906. The separately considered face component 924 is a taste-related sensory component 926 and is a mouth in the plurality of face component-specific local regions 908.

[0048] The corresponding separately considered face component 916, 920, or 924 among the plurality of separately considered face components 916, 920, and 924 corresponds to a corresponding facial landmark set 910 to 914 among the plurality of facial landmark sets 910 to 914 in the facial shape 406. The corresponding facial landmark set 910, 912, or 914 among the plurality of facial landmark sets 910 to 914 includes a plurality of facial landmarks. For example, referring to Figure 3 and Figure 5 as shown, the facial landmark set 910 among the plurality of facial landmark sets 910 to 914 includes the plurality of facial landmarks (18) to (27) and the plurality of facial landmarks (37) to (48) of the facial shape 406. The facial landmark set 912 among the plurality of facial landmark sets 910 to 914 includes the plurality of facial landmarks (28) to (36) of the facial shape 406. The facial landmark set 914 among the plurality of facial landmark sets 910 to 914 includes the plurality of facial landmarks (49) to (68) of the facial shape 406. The remaining description of the facial landmark detector 202 including the cropping module 502 may apply mutatis mutandis to the facial landmark detector 202 including the cropping module 902.

[0049] Figure 10 is to illustrate according to an embodiment of the present disclosure in Figure 6the cascaded regression stages R1 to R in one of the multiple face component-specific local refinement modules 602 to 608 in M a block diagram. Hereinafter, the description for each of the multiple face component-specific local refinement modules 602 to 608 is first described without reference to the drawings. Then, taking the face component-specific local refinement module 604 as an example, and with reference to Figure 10 for illustration. For simplicity, the description with reference to Figures 11 to 13 is only given taking the face component-specific local refinement module 604 as an example. The description of the face component-specific local refinement module 604 is converted into the description of each of the face component-specific local refinement modules 604 so that the appended claims can use the description with reference to Figure 10 as an example.

[0050] For each of the multiple face component-specific local regions, a corresponding face component-specific local refinement module among the multiple face component-specific local refinement modules is configured to receive each of the multiple face component-specific local regions, and perform a cascaded regression method using each of the multiple face component-specific local regions and a corresponding first face landmark set among the multiple first face landmark sets to obtain a corresponding second face landmark set among the multiple second face landmark sets. The corresponding face component-specific local refinement module among the multiple face component-specific local refinement modules includes multiple cascaded regression stages. Each of the multiple cascaded regression stages is configured to receive each of the multiple face component-specific local regions and a face landmark set among the multiple previous stage face landmark sets corresponding to each of the multiple face component-specific local regions, perform the cascaded regression method, and output a face landmark set among the multiple current stage face landmark sets corresponding to each of the multiple face component-specific local regions. The face landmark set among the multiple previous stage face landmark sets corresponding to the starting stage of the multiple cascaded regression stages is the corresponding face landmark set among the multiple first face landmark sets. The face landmark set among the multiple current stage face landmark sets for one stage of the multiple cascaded regression stages becomes the face landmark set among the multiple previous stage face landmark sets for the other stage immediately following the stage. The face landmark set among the multiple current stage face landmark sets corresponding to the last stage of the multiple cascaded regression stages is the corresponding face landmark set among the multiple second face landmark sets.

[0051] For example, the facial component specific local refinement module 604 is configured to receive the facial component specific local region 506, and use the facial component specific local region 506 and the facial landmark set 514 to perform the cascade regression method to obtain the facial landmark set 620. The facial component specific local refinement module 604 includes a plurality of cascade regression stages R1 to R2. M The multiple cascade regression stages R1 to R M Each of the facial component specific local area 506 and a previous stage facial landmark set 1106 (in Figure 11 ), perform multiple steps in one stage of the cascade regression method, and output a current stage facial landmark set 1110 (in Figure 11 The multiple cascade regression stages R1 to R M The facial landmark set 1106 corresponding to the first stage R1 is the facial landmark set 514. M Phase R t (exist Figure 11 The current stage facial landmark set 1110 (marked in FIG. 1 ) becomes the previous stage facial landmark set 1106 for use in the next stage R t Another stage after R t+1 Corresponding to the plurality of cascaded regression stages R1 to R M The final stage of M The current stage facial landmark set 1110 is the facial landmark set 620 .

[0052] Figure 11 FIG. 1 is a diagram illustrating an embodiment of the present disclosure. Figure 10 The multiple cascade regression stages R1 to R M Each stage R t FIG. 1 is a block diagram of a local feature extracting module 1102 and a local feature organizing module 1104 in FIG. 1 . The plurality of cascaded regression stages R1 to R M Each stage R t The system comprises a local feature extraction module 1102 and a local feature organization module 1104. The local feature extraction module 1102 is configured to receive the facial component specific local region 506 and the previous stage facial landmark set 1106, extract a plurality of local features 1108 using the facial component specific local region 506 and the previous stage facial landmark set 1106, and output the plurality of local features 1108. Figure 12A andFigure 12B Among them, the multiple cascaded regression stages R1 to R M Among them (such as Figure 10 shown), the local feature extraction module 1102 of the start stage R1 in M will be taken as an example for description. The description of the local feature extraction module 1102 of the start stage R1 in the multiple cascaded regression stages R1 to R M can be applied mutatis mutandis to the local feature extraction module 1102 of any other stage in the multiple cascaded regression stages R1 to R Figure 3 and Figure 11 and Figure 12A and Figure 12B , the extraction step includes extracting each of the multiple local features (such as 1204) (such as 1210) from a corresponding facial landmark (such as a specific local area of a facial landmark around the facial landmark (37) (such as 1206)) of the facial landmark set of the previous stage (such as 1202). The specific local area of the facial landmark (such as 1206) is in the specific local area of the facial component (such as 506). Referring to Figure 11 , the local feature organization module 1104 is configured to receive the facial landmark set 1106 of the previous stage and the multiple local features 1108, and organize the multiple local features 1108 based on multiple correlations between the multiple local features 1108, so as to use the multiple local features 1108 and the facial landmark set 1106 of the previous stage to obtain the facial landmark set 1110 of this stage. Then Figure 12A and Figure 12B In the example of, referring to Figure 11 and Figure 13 , the organization step is to organize the multiple local features (such as 1204) based on multiple correlations between the multiple local features (such as 1204) to use the multiple local features (such as 1204) and the facial landmark set of the previous stage (such as 1202) to obtain the facial landmark set of the current stage (such as 1312).

[0053] Figure 12A is a block diagram illustrating multiple facial landmark specific local feature mapping functions used in the local feature extraction module 1102 of the start stage R1 in the multiple cascaded regression stages R1 to R M (in Figure 10 ) according to an embodiment of the present disclosure. Referring to Figure 11 and and . Referring to Figure 12A and Figure 12B, in the starting stage R1, the local feature extraction module 1102 performs a plurality of operations, and the plurality of operations include a plurality of face landmark specific local feature mapping functions and in a corresponding face landmark specific local feature mapping function (such as ), maps the face landmark specific local area (such as 1206) around the corresponding face landmark (such as face landmark (37)) of the previous stage face landmark set 1202 to each of the plurality of local features 1204 (such as 1210). The plurality of face landmark specific local feature mapping functions and are independent. Each of the plurality of face landmark specific local feature mapping functions and is represented by an expression (1) as follows.

[0054]

[0055] Wherein, l represents the l-th face landmark, as Figure 3 shown, t represents the t-th stage in the plurality of cascaded regression stages R1 to R M . Each of the plurality of local features 1204 (such as 1210) is represented by an expression (2) as follows.

[0056]

[0057] Wherein, l c identifies a face component specific local area having a separately considered face component c, such as the face component specific local area 506 having two eyes, and identifies a corresponding previous stage face landmark set of the separately considered face component c, such as the previous stage face landmark set 1202 corresponding to two eyes.

[0058] In the above embodiment, the plurality of local features 1204 are extracted using the plurality of independent face landmark specific local feature mapping functions and . Other methods for extracting a plurality of local features, such as using local binary patterns (LBP) or scale-invariant feature transform (SIFT), are within the scope of the present disclosure.

[0059] Figure 12B is to illustrate, according to an embodiment of the present disclosure, in Figure 12AThe multiple face landmark specific local feature mapping functions implemented by a random forest 1208 and A block diagram of one of them. Refer to Figure 12A and Figure 12B , in one embodiment, the multiple face landmark specific local feature mapping functions and Each of them is implemented by a corresponding random forest. Taking the multiple face landmark specific local feature mapping functions implemented by the random forest 1208 as an example for illustration. The description of the multiple face landmark specific local feature mapping functions can be analogously applied to other multiple face landmark specific local feature mapping functions and The random forest 1208 includes multiple decision trees 1212 and 1214. Each of the multiple decision trees 1212 and 1214 includes at least one splitting node 1216 and at least one leaf node 1218. Each of the at least one splitting nodes 1216 determines whether to branch left or right. During training, each of the at least one leaf nodes 1218 is associated with a continuous prediction for a regression target. The face landmark specific local region 1206 around the face landmark (37) of the previous stage face landmark set 1202 traverses the multiple decision trees 1212 and 1214 until reaching a leaf node 1218 of each decision tree 1212 and 1214. In one embodiment, the face landmark specific local region 1206 is a circular region with a radius of r and centered at the position of the face landmark (37). The local feature 1210 is a vector including multiple bits, and each bit corresponds to a corresponding leaf node 1218 of the random forest 1208. The one leaf node 1218 of each of the multiple decision trees 1212 and 1214 reached by the face landmark specific local region 1206 corresponds to a bit of the local feature 1210 having a value of "1". Each of the other bits of the local feature 1210 has a value of "0".

[0060] In the above embodiment, the multiple face landmark specific local feature mapping functions and Each of the facial landmark specific local feature mapping functions is implemented by the random forest 1208. Other ways of implementing each of the multiple facial landmark specific local feature mapping functions, such as using a convolutional neural network, are within the intended scope of the present disclosure. In the above embodiment, the facial landmark specific local area 1206 is a circular shape. Other shapes of a facial landmark specific local area, such as a square, a rectangle and a triangle, are within the intended scope of the present disclosure.

[0061] Figure 13 FIG. 1 is a diagram illustrating an embodiment of the present disclosure. Figure 11 1 is a block diagram of a local feature concatenating module 1302, a facial component-specific projecting module 1304, and a facial landmark set incrementing module 1306 in the local feature organizing module 1104. The local feature organizing module 1104 includes the local feature concatenating module 1302, the facial component-specific projection module 1304, and the facial landmark set incrementing module 1306. The local feature concatenating module 1302 is configured to receive the plurality of local features 1204 and concatenate the plurality of local features 1204 into a facial component-specific feature 1308. The facial component-specific projection module 1304 is used to receive the facial component-specific feature 1308, project the facial component-specific local region 506 (such as the facial component-specific projection matrix) according to a facial component-specific projection matrix. Figure 12A The facial component specific features 1308 corresponding to the facial component are projected in a facial component specific manner, and a facial landmark set increment 1310 is output. The facial landmark set increment 1310 is obtained by equation (3), as shown below.

[0062]

[0063] in Indicates a facial landmark set increment corresponding to a separately considered facial component c at stage t, such as the facial landmark set increment 1310, denotes a facial component specific projection matrix corresponding to the separately considered facial component c at stage t, A facial component specific feature corresponding to a separately considered facial component c at stage t is indicated, such as the facial component specific feature 1308. In one embodiment, the facial component specific projection matrix is a linear projection matrix. The facial landmark set increment module 1306 receives the facial landmark set increment 1310 and the previous stage facial landmark set 1202, and applies the facial landmark set increment 1310 to the previous stage facial landmark set 1202 to obtain the current stage facial landmark set 1312.

[0064] Figure 14 is to illustrate, in accordance with an embodiment of the present disclosure, the Figure 10 among the multiple cascaded regression stages R1 to R M of the multiple cascaded training stages T1 to T P of a block diagram. Each of the cascaded training stages T1 to T P is configured to receive a plurality of training sample face component specific local regions 1402, a plurality of corresponding ground truth facial landmark sets 1404 of the plurality of training sample face component specific local regions 1402, and a plurality of previous stage facial landmark sets 1506 corresponding to the plurality of training sample face component specific local regions 1402 (marked in Figure 15 ). Each of the plurality of training sample face component specific local regions 1402 is defined using a training sample face image and includes a plurality of separately considered face components of the same type. Each of the multiple cascaded training stages T1 to T P is further configured to use the plurality of training sample face component specific local regions 1402, the plurality of ground truth facial landmark sets 1404, and the plurality of previous stage facial landmark sets 1506 to train a plurality of facial landmark specific local feature mapping functions 1408 and a face component specific projection matrix 1410. The facial landmark specific local feature mapping functions 1408 are correspondingly used, for example, as Figure 12A the plurality of facial landmark specific local feature mapping functions in and the face component specific projection matrix 1410 is used, for example, as Figure 12B the face component specific projection matrix in where the separately considered face component c is two eyes. Each of the multiple cascaded training stages T1 to T P-1 is further configured to output a plurality of current stage facial landmark sets 1514 corresponding to the training sample face component specific local regions 1402 (marked in Figure 15 ). The previous stage facial landmark sets 1506 corresponding to the initial stage T1 of the multiple cascaded regression stages T1 to T P are a plurality of facial landmark sets 1406. Each of the plurality of facial landmark sets 1406 can be similarly referred to Figure 4 andFigure 5 The described set of facial landmarks 514 is obtained. The multiple cascaded regression stages T1 to T P-1 One stage T in t (marked in Figure 15 ) the current stage facial landmark set 1514 of becomes the previous stage facial landmark set 1506, for the stage T immediately t After another stage T t+1 .

[0065] Figure 15 Is an illustration of the multiple cascaded training stages T1 to T in accordance with an embodiment of the present disclosure Figure 14 Each stage T in P In the facial landmark specific local feature mapping function training module (facial landmark-specific local feature mapping function training module) 1502 and a facial component-specific projection matrix training module (facial component-specific projection matrix training module) 1504 of a block diagram. The multiple cascaded training stages T1 to T t Each stage T in P Includes a facial landmark specific local feature mapping function training module 1502 and a facial component-specific projection matrix training module 1504. t

[0066] The facial landmark specific local feature mapping function training module 1502 is configured to receive the multiple training sample facial component-specific local regions 1402, the ground truth facial landmark set 1404, and the previous stage facial landmark set 1506, and train each of the multiple facial landmark specific local feature mapping functions 1408 independently of each other and output the multiple local feature sets 1512 corresponding to the multiple training sample facial component-specific local regions 1402, using the multiple training sample facial component-specific local regions 1402, the ground truth facial landmark set 1404, and the previous stage facial landmark set 1506. In one embodiment, each of the multiple facial landmark specific local feature mapping functions 1408 is obtained by minimizing an objective function (4) shown below.

[0067]

[0068] Figure 14 Where t represents Figure 14 The multiple cascaded training stages T1 to T in PIn a t-th stage, for i iterating all the training sample face component specific local regions 1402, l represents the l-th facial landmark as shown in Figure 3 wherein the l-th facial landmark is shown, is an increment of a ground truth facial landmark set corresponding to the i-th training sample face component specific local region at the t-th stage, π l is obtained from the increment of the ground truth facial landmark set Two elements (2l, 2l - 1) are extracted from the increment of the ground truth facial landmark set is a 2D offset of the l-th facial landmark in the l-th training sample face component specific local region, l i is the i-th training sample face component specific local region, is the local region of the previous stage facial landmark set corresponding to the i-th training sample face component specific local region, such as one of the multiple previous stage facial landmark sets 1506, is a facial landmark specific local feature mapping function corresponding to the l-th facial landmark at the t-th stage, such as one of the multiple facial landmark specific local feature mapping functions 1408, is a local feature corresponding to the l-th facial landmark at the t-th stage and the i-th training sample face component specific local region, such as a local feature of a local feature set of one of the multiple local feature sets 1512, is used to map the local feature to a 2D offset. The increment of the ground truth facial landmark set is obtained by an equation (5) as follows.

[0069]

[0070] Wherein, is the ground truth facial landmark set of the i-th training sample corresponding to the i-th training sample face component specific local region such as one of the multiple ground truth facial landmark sets 1404, and is a previous stage facial landmark set corresponding to the i-th training sample face component specific local region, such as one of the multiple previous stage facial landmark sets 1506. The local linear projection matrix is a 2×D matrix (2-by-D matrix), where D is the dimension of the local feature The standard regression random forest is used to learn each facial landmark specific local feature mapping function

[0071] An example of the random forest corresponding to a learned facial landmark specific local feature mapping function is referred to Figure 12B ​The described facial landmark specific local feature mapping function The corresponding random forest 1208. Multiple split nodes in the random forest are trained using the pixel difference features. To train each split node in the random forest, 500 randomly sampled pixel features are selected from a facial landmark specific local area around a facial landmark, and the feature that produces the maximum variance reduction is picked. The facial landmark specific local area is similar to the reference Figure 12B The described facial landmark specific local area 1206. After training, each leaf node stores a two-dimensional offset vector, which is the average of all the training sample facial component specific local areas 1402 in each leaf node. During the testing process, each of the multiple training sample facial component specific local areas 1402 traverses the random forest, and the pixel difference features of each of the multiple training sample facial component specific local areas 1402 are compared with each node until each of the multiple training sample facial component specific local areas 1402 reaches a leaf node. For each dimension of the local feature If the i-th training sample facial component specific local area reaches a corresponding leaf node, the value of each dimension is "1", otherwise it is "0".

[0072] The facial component specific projection matrix training module 1504 is configured to receive multiple ground truth facial landmark set increments 1510 and the multiple local feature sets 1512, and train the facial component specific projection matrix 1410 and output the current stage facial landmark set 1514, using the ground truth facial landmark set increments 1510 and the multiple local feature set increments 1512. Each of the multiple ground truth facial landmark set increments 1510 is the ground truth facial landmark set increment in the objective function (4) The facial component specific projection matrix 1410 is trained using the multiple local feature sets 1512 corresponding to the training sample facial component specific local areas 1402, the training sample facial component specific local areas 1402 include multiple facial components of the same type considered separately, but not the multiple local features corresponding to the multiple training sample facial component specific local areas, the multiple training sample facial component specific local areas include other types of facial components considered separately. In one embodiment, the facial component specific projection matrix 1410 is obtained by minimizing an objective function (5) shown as follows.

[0073]

[0074] Where the first term is the regression target, is a face component specific feature corresponding to a specific local region of the i-th training sample face component in the t-th stage, is a face component specific projection matrix, such as the face component specific projection matrix 1410, and the second term is for an L1 regularization of, where λ controls the regularization strength. The face component specific feature is the plurality of concatenated local features, where each local feature in the plurality of concatenated local features is the local feature described with reference to the objective function (4) Any optimization technique can be used, such as singular value decomposition (SVD), gradient descent or dual coordinate descent. After obtaining the face component specific projection matrix each of the current stage face landmark sets 1514 is

[0075] Figure 16 is to illustrate the implementation of a joint detection module 1602 according to an embodiment of the present disclosure in Figure 4A block diagram of the overall facial landmark acquisition module 402 in [the above]. In one embodiment, the overall facial landmark acquisition module 402 is implemented using a joint detection module 1602. The joint detection module 1602 is configured to receive the facial image 204 and perform joint detection using the facial image 204 to obtain a facial shape 406. The joint detection method obtains multiple facial landmarks corresponding to multiple facial components in a facial image together. For example, the joint detection method obtains the multiple facial landmarks (1) to (17) corresponding to the facial contour in the facial image 204, the multiple facial landmarks (18) to (27) corresponding to the eyebrows in the facial image 204, the multiple facial landmarks (37) to (48) of the eyes in the facial image 204 together, the multiple facial landmarks (28) to (36) of the nose in the facial image 204, and the facial landmarks (49) to (68) of the mouth in the facial image 204. In one embodiment, the joint detection method is a cascaded regression method. The cascaded regression method extracts multiple local features using the facial image 204, concatenates the multiple local features into a global feature, and performs a joint projection on the global feature to obtain a facial shape for a current stage. A joint projection matrix used during the joint projection is trained using a regression target that involves multiple facial landmarks of multiple facial components such as the facial contour, eyebrows, eyes, nose, and mouth. In another embodiment, the joint detection method is a deep learning facial landmark detection method. The deep learning facial landmark detection method includes a convolutional neural network having multiple layers, where at least one layer obtains multiple facial landmarks corresponding to multiple facial components in a facial image together.

[0076] In the above embodiment, the overall facial landmark acquisition module 402 is implemented using the joint detection method. Other ways of implementing the overall facial landmark acquisition module 402, such as using a random guess or an average facial shape obtained from multiple training samples, are within the scope of this disclosure.

[0077] Some embodiments have one or a combination of the following features and / or advantages. In the related art, a cascaded regression method is also a joint detection method that extracts multiple local features using a facial image, concatenates the multiple local features into a global feature, performs a joint projection on the global feature, and obtains a facial shape for a current stage. A joint projection matrix used during the joint projection is trained using a regression target that relates to multiple facial landmarks of multiple facial components, such as a facial contour, two eyebrows, two eyes, a nose, and a mouth. Therefore, the optimization of the joint projection matrix involves all facial components. Thus, for example, during the optimization process, a change in the facial landmark of the nose affects the changes in the facial landmarks of the facial contour, the eyebrows, the eyes, and the mouth. When the nose is abnormal, it has an adverse effect on the training of the joint projection matrix, resulting in the joint projection matrix not being optimal not only for a nose but also for a facial contour, eyebrows, eyes, and a mouth during an inference stage. Compared with the related art, some embodiments of the present disclosure define multiple facial component-specific local regions using a facial image and perform a cascaded regression method on each of the multiple facial component-specific local regions. The cascaded regression method of some embodiments of the present disclosure extracts multiple local features using each of the multiple facial component-specific local regions, concatenates the multiple local features into a facial component-specific feature, and performs a facial component-specific projection on the facial component-specific feature to obtain a corresponding set of facial landmarks in a set of multiple facial landmarks for a current stage. A facial component-specific projection matrix used during the facial component-specific projection is trained using a regression target that only relates to the multiple facial landmarks of a separately considered facial component, such as an eye. Therefore, the optimization of the facial component-specific projection matrix involves the separately considered facial component. Thus, for example, during the optimization process, a change in the facial landmark of the eye does not affect the changes in the multiple facial landmarks of the eyebrows, a nose, and a mouth. When the eye is abnormal, it does not have an adverse effect on the training of the multiple facial component-specific projection matrices of other facial components, so that the facial component-specific projection matrix is preferably used for the eyebrows, the nose, and the mouth during an inference stage. In addition, the complexity of optimizing the joint projection matrix is higher than the complexity of optimizing each of the multiple facial component-specific projection matrices.

[0078] In the related art, a cascaded regression method such as the cascaded regression method performing joint detection uses a random guess or an average face shape as an initialization (i.e., a previous stage face shape for the initial stage of the cascaded regression method). Since the cascaded regression method largely depends on the initialization, when the head pose of a face image used for face landmark detection deviates significantly from the head pose of the random guess or the average face shape, the performance of face landmark detection is poor. Compared with the related art, some embodiments of the present disclosure perform a joint detection method to roughly detect a face shape, and use the face shape as an initialization for a cascaded regression method, and the cascaded regression method performs face component specific local refinement for each of a plurality of face landmark sets in the face shape. The plurality of face landmark sets correspond to a plurality of face components considered separately. Therefore, a coarse-to-fine face landmark detection is performed, thereby improving the accuracy of a detected face shape. In addition, since the face component specific local refinement is performed specifically for a face component locally, the accuracy of the detected face shape can be obtained without sacrificing speed. Table 1 below illustrates the experimental results for comparing the accuracy and speed of a supervised descent method (SDM), which is a cascaded regression method using a random guess or an average face shape as an initialization, and some embodiments of the present disclosure perform a coarse-to-fine face landmark detection. The SDM is described in "Supervised Descent Method and Its Application in Face Alignment", Xiong, X., De la Torre Frade, F., in: IEEE International Conference on Computer Vision and Pattern Recognition, 2013. As shown, compared with the SDM, the coarse-to-fine face landmark detection in some embodiments of the present disclosure has a significant improvement in a normalized mean error (NME) without sacrificing speed.

[0079] Table 1

[0080]

[0081] In a related art, a deep learning facial landmark detection method uses a complex / deep architecture to improve the accuracy of a detected facial shape. Compared with the deep learning facial landmark detection method, the coarse-to-fine facial landmark detection in some embodiments of the present disclosure uses another deep learning facial landmark detection method, employing a shallower or narrower architecture for coarse detection and face component-specific local refinement for fine detection. Therefore, the accuracy of a detected facial shape can be improved without significantly increasing the computational cost.

[0082] Those of ordinary skill in the art can understand that each of the many units, many modules, many layers, many blocks, algorithms, and many steps of the systems or computer-implemented methods described and disclosed in the embodiments of the present invention is implemented by hardware, firmware, software, or a combination thereof. Whether the many functions run in a hardware, firmware, or software manner depends on the application conditions and the design requirements of the technical solution. Those of ordinary skill in the art can use different ways to implement the functions for each specific application, but such implementation manners should not exceed the scope of the present disclosure.

[0083] It should be understood that the systems and computer-implemented methods disclosed in the embodiments of the present disclosure can be implemented in other ways. The above embodiments are merely exemplary. The division of the many modules is only based on logical functions, and other divisions exist in implementation. The many modules may or may not be physical modules. It is possible to combine or integrate the many modules into one physical module. It is also possible to divide any module into multiple physical modules. It is also possible to omit or skip some features. On the other hand, the mutually coupled, directly coupled, or communicatively coupled shown or discussed operates indirectly or communicatively through some ports, devices, or modules in an electrical, mechanical, or other form.

[0084] The modules as separate components for illustration are physically separate and may also not be separate. The many modules are located in one place or distributed on multiple network modules. Some or all of the modules are used according to the purposes of the embodiments.

[0085] If the software functional modules are implemented and sold as products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution proposed by the present invention can be substantially or partially implemented in the form of a software product. Or, a part of the technical solution that is beneficial to the prior art can be implemented in the form of a software product. The software product is stored in a computer-readable storage medium and includes multiple commands for at least one processor of a system to run all or part of the steps disclosed in the multiple embodiments of the present disclosure. The storage medium includes a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a floppy disk, or other media capable of storing multiple program instructions.

[0086] While the present disclosure has been described in connection with what are considered to be the most practical and preferred embodiments, it is to be understood that the present disclosure is not limited to the disclosed embodiments, but is intended to cover various arrangements made without departing from the broadest scope of interpretation of the appended claims.

Claims

1. A computer-implemented method, characterized in that: Comprising: A method for performing an inference stage, wherein the method of the inference stage comprises: Receiving a first facial image; Obtaining a first facial shape using the first facial image; Defining a plurality of face component specific local regions using the first facial image and the first facial shape, wherein each of the plurality of face component specific local regions comprises a corresponding separately considered face component among a plurality of separately considered face components from the first facial image, and the corresponding separately considered face component among the plurality of separately considered face components corresponds to a corresponding first facial landmark set among a plurality of first facial landmark sets in the first facial shape, wherein the corresponding first facial landmark set among the plurality of first facial landmark sets comprises a plurality of facial landmarks; For each of the plurality of face component specific local regions, performing a cascaded regression method using each of the plurality of face component specific local regions and a corresponding facial landmark set among a plurality of previous stage facial landmark sets to obtain a corresponding facial landmark set among a plurality of second facial landmark sets, wherein each stage of the cascaded regression method comprises: Extracting a plurality of local features using each of the plurality of face component specific local regions and a corresponding facial landmark set among a plurality of previous stage facial landmark sets, Wherein: The extracting step comprises extracting each of the plurality of local features from a face landmark specific local region around a corresponding face landmark of the corresponding face landmark set among the plurality of previous stage facial landmark sets, wherein the face landmark specific local region is in each of the plurality of face component specific local regions; and The corresponding face landmark set among the plurality of previous stage facial landmark sets corresponding to the initial stage of the cascaded regression method is the corresponding face landmark set among the plurality of first facial landmark sets; and Organizing the plurality of local features based on a plurality of correlations between the plurality of local features to obtain a corresponding facial landmark set among a plurality of current stage facial landmark sets, wherein the corresponding facial landmark set among the plurality of current stage facial landmark sets corresponding to the final stage of the cascaded regression method is the corresponding facial landmark set among the plurality of second facial landmark sets; Wherein the defining step comprises: defining each of the plurality of face component specific local regions by cropping such that a plurality of separately considered face components other than the corresponding separately considered face component among the plurality of separately considered face components are at least partially removed, wherein the plurality of second facial landmark sets are correspondingly located on the separated plurality of face component specific local regions; the first facial shape further comprises a third facial landmark set corresponding to a facial contour from the first facial image; and The method of the inference stage further comprises: Combining the plurality of second facial landmark sets and the third facial landmark set correspondingly located on the separated plurality of face component specific local regions into a second facial shape.

2. The computer-implemented method according to claim 1, characterized in that: The plurality of separately considered face components are separated according to a plurality of facial features.

3. The computer-implemented method according to claim 2, wherein: The plurality of facial features are functionally grouped.

4. The computer-implemented method according to claim 2, characterized in that: The multiple facial features are grouped non - functionally.

5. The computer-implemented method according to claim 1, wherein: The step of extracting each of the multiple local features includes mapping a facial landmark - specific local region around the corresponding facial landmark of the corresponding facial landmark set in the multiple previous - stage facial landmark sets to each of the multiple local features according to a corresponding facial landmark - specific local feature mapping function among multiple facial landmark - specific local feature mapping functions.

6. The computer-implemented method according to claim 5, wherein: Further included is: A method for performing a training stage, where the method of the training stage includes: Training each of the multiple facial landmark - specific local feature mapping functions independently of each other.

7. The computer - implemented method according to claim 6, wherein: The step of organizing includes: Concatenating the multiple local features into a facial - component - specific feature; and Performing a facial - component - specific projection on the facial - component - specific feature corresponding to each of the multiple facial - component - specific local regions according to a corresponding facial - component - specific projection matrix among multiple facial - component - specific projection matrices; and The method of the training stage further includes: Using the multiple facial landmark - specific local feature mapping functions corresponding to each of the multiple facial - component - specific local regions to train the corresponding facial - component - specific projection matrix among the multiple facial - component - specific projection matrices, rather than using the multiple facial landmark - specific local feature mapping functions corresponding to the multiple facial - component - specific local regions other than each of the multiple facial - component - specific local regions.

8. The computer-implemented method according to claim 1, wherein: The step of organizing includes: concatenating the multiple local features into a facial - component - specific feature; and performing a facial - component - specific projection on the facial - component - specific feature corresponding to each of the multiple facial - component - specific local regions according to a corresponding facial - component - specific projection matrix among multiple facial - component - specific projection matrices.

9. A system, characterized in that: Including: At least one memory configured to store multiple program instructions; At least one processor configured to execute the multiple program instructions, the multiple program instructions causing the at least one processor to execute multiple steps, including: A method for performing an inference stage, where the method of the inference stage includes: Receiving a first facial image; Obtaining a first facial shape using the first facial image; Defining multiple facial - component - specific local regions using the first facial image and the first facial shape, where each of the multiple facial - component - specific local regions includes a corresponding separately - considered facial component from multiple separately - considered facial components of the first facial image, and the corresponding separately - considered facial component among the multiple separately - considered facial components corresponds to a corresponding first facial landmark set among multiple first facial landmark sets in the first facial shape, and the corresponding first facial landmark set among the multiple first facial landmark sets includes multiple facial landmarks; For each of the multiple facial component-specific local regions, a cascaded regression method is performed using each of the multiple facial component-specific local regions and a corresponding facial landmark set from the multiple first facial landmark sets to obtain a corresponding facial landmark set from the multiple second facial landmark sets, where each stage of the cascaded regression method includes: extracting multiple local features using each of the multiple facial component-specific local regions and a corresponding facial landmark set from multiple previous-stage facial landmark sets; wherein: the extracting step includes extracting each of the multiple local features from a facial landmark-specific local region around a corresponding facial landmark of the corresponding facial landmark set from the multiple previous-stage facial landmark sets, where the facial landmark-specific local region is within each of the multiple facial component-specific local regions; and the corresponding facial landmark set from the multiple previous-stage facial landmark sets corresponding to an initial stage of the cascaded regression method is the corresponding facial landmark set from the multiple first facial landmark sets; and organizing the multiple local features based on multiple correlations between the multiple local features to obtain a corresponding facial landmark set from multiple current-stage facial landmark sets, where the corresponding facial landmark set from the multiple current-stage facial landmark sets corresponding to a final stage of the cascaded regression method is the corresponding facial landmark set from the multiple second facial landmark sets; where the defining step includes: defining each of the multiple facial component-specific local regions by cropping such that multiple separately considered facial components other than the corresponding separately considered facial component among the multiple separately considered facial components are at least partially removed, where the multiple second facial landmark sets are correspondingly located on the separated multiple facial component-specific local regions; the first facial shape further includes a third facial landmark set corresponding to a facial contour from the first facial image; and the method of the inference stage further includes: merging the multiple second facial landmark sets correspondingly located on the separated multiple facial component-specific local regions and the third facial landmark set into a second facial shape.

10. The system according to claim 9, wherein: The multiple separately considered facial components are separated according to multiple facial features.

11. The system according to claim 10, wherein: The multiple facial features are functionally grouped.

12. The system according to claim 10, wherein, The facial features are non-functionally grouped.

13. The system according to claim 9, characterized in that: The step of extracting each of the multiple local features includes mapping the facial landmark-specific local region around the corresponding facial landmark of the corresponding facial landmark set from the multiple previous-stage facial landmark sets into each of the multiple local features according to a corresponding facial landmark-specific local feature mapping function among multiple facial landmark-specific local feature mapping functions.

14. The system according to claim 13, characterized in that: Further included is: performing a method of a training stage, where the method of the training stage includes: training each of the multiple facial landmark-specific local feature mapping functions independently of each other.

Citation Information

Patent Citations

  • Face alignment with shape regression

    US20160055368A1