A high-resolution artistic human face landmark detection method based on deep learning

By combining the global encoder-decoder and regional encoder-decoder networks and optimizing the loss function, accurate detection of facial landmarks in high-resolution artworks is achieved, which solves the problem of incomplete detection in artworks by existing methods and improves the detection effect.

CN118196850BActive Publication Date: 2025-10-14HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410136122.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-10-14
Estimated Expiration
2044-01-31

AI Technical Summary

Technical Problem

Existing methods have difficulty accurately detecting facial landmarks in high-resolution images of artworks, especially prints and paintings, resulting in incomplete detection or unrecognizable features.

Method used

A global encoder-decoder network and a regional encoder-decoder network are combined to perform coarse detection through the global encoder-decoder network and refine it using the regional encoder-decoder network. Combined with mean square error loss optimization, accurate detection of high-resolution artistic face landmarks is achieved.

Benefits of technology

It significantly improves the accuracy and completeness of facial landmark detection in artworks, especially in paintings and prints, and shows better detection results than existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118196850B_ABST
    Figure CN118196850B_ABST
Patent Text Reader

Abstract

The application discloses a high-resolution artistic face landmark detection method based on deep learning, first, an artistic face data set and corresponding landmark annotations used for model training and evaluation are prepared; a global low-resolution predicted landmark detection map and a regional high-resolution predicted landmark detection map are obtained respectively; the global low-resolution and regional high-resolution predicted landmark detection maps are optimized through a mean square deviation loss; finally, the high-resolution predicted landmark detection map of each region is restored to a global coordinate system, that is, face landmark detection is realized. The application shows excellent effect in face landmark detection of paintings and prints and other artworks. The application proposes a combination of a global encoding-decoding network and a regional encoding-decoding network to realize rough and refined marking of artistic face landmarks, and the marking effect is more excellent than that of existing methods.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a high-resolution artistic face landmark detection method based on deep learning, and is suitable for the fields of computer graphics and face feature detection. BACKGROUND

[0002] Face landmark detection aims to automatically identify and locate specific face key points or landmarks in a face image. These key points are usually important features of a face, such as eyes, eyebrows, nose, mouth, chin, etc. Through landmark detection, the structural information of the face can be obtained, so as to perform a series of applications such as face analysis, expression recognition, pose estimation, face transformation and feature extraction. Face landmark detection is a widely explored field in computer vision tasks, but it is not directly applicable to artworks. Artworks have great differences in image texture and geometric structure of facial features compared with natural face images, and most artwork detection is based on high-resolution or macro photography images, so a method for accurately detecting face landmarks in high-resolution artwork images is needed. A machine learning-based face landmark detection method dlib uses level regression to gradually optimize and refine the position of the landmark point through a series of regression trees, but the dlib method cannot well detect face landmarks in high-resolution face images of artworks. FOA proposes a method for detecting artistic face image landmarks using deep learning, which first estimates the global face landmark position, and then corrects the face landmark position using a pre-trained point distribution model. Due to the lack of a large number of artistic face high-resolution image training and the refinement of the global face landmark position, the face landmark detection result of this method is also not ideal. The present application expands the artistic face dataset and proposes a region encoder-decoder network for landmark refinement of part of the face. SUMMARY

[0003] The present application proposes a high-resolution artistic face landmark detection method based on deep learning, which realizes rough and fine artistic face landmark detection through a global encoder-decoder network and multiple regional encoder-decoder networks. The present application solves the defects of incomplete or even unrecognizable face landmark detection in existing methods in engravings and paintings and other artworks. The present application uses a ResNet-based encoder-decoder and uses a global feature map as an additional input of the regional encoder-decoder network, and makes some improvements to the traditional architecture and concept.

[0004] The technical scheme adopted by the present application to achieve the above-mentioned purposes is a high-resolution artistic face landmark detection method based on deep learning, which is implemented according to the following steps:

[0005] Step 1: Prepare an artistic face dataset and corresponding landmark annotations for model training and evaluation;

[0006] Step 2: Downsample high-resolution images in the artistic face dataset, and predict output heat maps and extract global low-resolution predicted landmark detection maps through a global encoder-decoder network for the downsampled low-resolution images;

[0007] Step 3: Crop the high-resolution images based on the global low-resolution predicted landmark detection maps, and connect the obtained regional images with the corresponding regional heat maps in the output heat maps in Step 2 to input into a regional encoder-decoder network to generate refined regional high-resolution predicted landmark detection maps;

[0008] Step 4: Optimize the global low-resolution and regional high-resolution predicted landmark detection maps through mean square error loss;

[0009] Step 5: Restore the high-resolution predicted landmark detection maps of each region to the global coordinate system to realize face landmark detection.

[0010] The present application has the following advantages:

[0011] 1) The present application applies style transformation technology to natural face images and creates a larger artistic face landmark dataset through semi-automatic landmark annotation. Compared with existing methods, since the present application trains artistic faces, the model shows excellent performance in face landmark detection of paintings and engravings and other artworks.

[0012] 2) The present application proposes a combination of a global encoder-decoder network and a regional encoder-decoder network to realize rough and refined marking of artistic face landmarks, and the marking effect is more excellent than existing methods. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 is the overall flowchart of the present application;

[0014] Figure 2 is the artistic face landmark detection flowchart of the present application;

[0015] Figure 3 is the performance comparison chart of the present application and other existing methods. DETAILED DESCRIPTION

[0016] The present application will be further described in detail below in combination with the drawings and specific embodiments.

[0017] As shown in Figure 1 , the high-resolution artistic face landmark detection method based on deep learning includes the following steps:

[0018] Step 1: Prepare artistic face dataset and corresponding landmark annotations for model training and evaluation.

[0019] The task of the present invention is related to artistic face landmark detection, so a large high-resolution dataset containing artistic style faces such as prints or paintings is needed to ensure sufficient training of the model. The present invention first collects high-resolution face images of different artistic styles from different museums and institutes as subset one, including the German National Museum, the Metropolitan Museum of Art, and the National Art Gallery. Then all artistic images containing faces are selected from the Wikiart dataset as subset two. In order to automatically detect and crop all face regions in subset one and subset two, the present invention first detects the face regions in subset one and subset two by the target detector YOLOv5, and then performs cropping operation on the regions. In addition, the present invention uses style transfer technology on the public dataset 300-W for face landmark detection, scales the natural face images in the 300-W dataset to 1024x1024x3 size and converts them into artistic style face images as subset three.

[0020] The processed subset one, subset two and subset three are used as artistic face dataset for model training, and the present invention annotates the face landmarks in the artistic face dataset in a semi-automatic manner to obtain the corresponding real landmark detection map, uses the face landmark detector dlib based on random forest to annotate the landmark points of the artistic face dataset, and manually corrects the landmark points or annotates the landmark points from scratch in the case of dlib failure.

[0021] Step 2: Downsample the high-resolution images in the artistic face dataset, and use the global encoder-decoder network to predict the output heat map and extract the global low-resolution predicted landmark detection map.

[0022] As Figure 2The figure is an artistic face landmark detection flowchart of the present application. First, the 1024x1024 resolution image in the artistic face dataset is down-sampled to obtain a low-resolution image of 256x256 size, and the low-resolution image is input into the global encoder-decoder network. The specific process of the network is as follows: by two down-sampling blocks composed of a convolution layer with a kernel of 3 and a step of 2, a BatchNorm normalization layer and a ReLu activation layer, the channel number of the input feature map is doubled and its size is reduced, which helps to extract high-level features of the image; then, the down-sampled image is processed through 6 ResNet blocks to enhance the modeling ability of the model to the image; then, through two up-sampling blocks composed of a transposed convolution layer with a kernel of 3 and a step of 2, a BatchNorm normalization layer and a ReLu activation layer, the channel number of the feature map is halved and its size is increased; finally, the up-sampled feature map is processed through a reflection padding layer and a convolution layer to map the feature map to output a heat map with 68 channels and the same width and height as the input image. The heat map marks the positions of each part of the face in the form of peak value, and uses spatial softmax to extract these landmarks from the heat map in a differentiable manner to obtain a global low-resolution predicted landmark detection map. The marked parts include eyes, nose, mouth and chin line.

[0023] Step 3: Based on the global low-resolution predicted landmark detection map, the high-resolution image is cropped to obtain a region image, and the region image is connected with the corresponding region heat map in the output heat map in step 2 to input into the region encoder-decoder network to generate a refined high-resolution predicted landmark detection map for each region.

[0024] Firstly, the low-resolution facial landmark point image obtained in step 2 is up-sampled by 4 times to match the high-resolution image. The rough landmark point positions in the global low-resolution predicted landmark point detection map obtained according to the global encoder-decoder network are used to crop the part regions of the high-resolution image, and the parts corresponding to the part regions include eyes, nose and mouth; in order to adapt to the resolution of 256*256, the size of the cropped region is filled according to a random value between 0.25-0.5 of an original region. Then, the region image corresponding to the automatically cropped part region is extracted from the heat map output by the global encoder-decoder network. The automatically cropped part region image and the corresponding region image extracted from the heat map are channel fused, and the fused different part region images are respectively input into the corresponding region encoder-decoder network, which has the same structure as the global encoder-decoder network. Through channel fusion, the region encoder-decoder network can obtain the global position of the landmark point as prior information, thereby supporting the refinement task. There is no weight sharing between different region encoder-decoder networks, so that each network can learn the specific features of its facial sub-region. The heat map output by the region encoder-decoder network has 17 fewer channels because it does not need to refine the lower jaw line part of the face. The heat map marks the positions of each part of the face in the form of peaks, and uses spatial softmax to extract these landmark points from the heat map in a differentiable manner to obtain the refined high-resolution predicted landmark point detection map of each region.

[0025] Step 4: optimizing the global low-resolution and regional high-resolution predicted landmark point detection maps through mean square error loss;

[0026] For the global low-resolution and regional high-resolution landmark point detection task, the present application uses the mean square error between the predicted landmark point detection map and the real landmark point detection map as the loss function, which is expressed as follows:

[0027]

[0028] wherein λ is a weighting factor, x respectively represents the global low-resolution predicted landmark point detection map and the global real landmark point detection map; y respectively represents the regional high-resolution predicted landmark point detection map and the regional real landmark point detection map; N G N represents the number of channels of the heat map output by the global encoder-decoder network, which is 68; r N represents the number of channels of the heat map output by the region encoder-decoder network, which is 51.

[0029] Step 5: restoring the high-resolution predicted landmark point detection map of each region back to the global coordinate system.

[0030] To facilitate model evaluation, the present application needs to transfer the high resolution face landmarks from each region back to the global coordinate system. First, based on the region high resolution predicted landmark point detection map, the bounding box coordinates and the original region size of the region are extracted, then the local coordinates of the extracted region are scaled by the original region size, and the offset of the bounding box is added, to obtain the position of the high resolution predicted landmark point of each region relative to the global coordinate system. For the chin line, the present application only uses the predicted landmark point position of the global predicted landmark point detection map. Therefore, the complete face landmark point prediction result is the combination of the chin line position of the global predicted landmark point detection map and the refined region high resolution predicted landmark point relative to the global coordinate system position.

[0031] As shown in Figure 3 The present application and the prior art method are compared in the performance of artistic face landmark point detection, as shown in the figure, the dlib method does not mark the eye and mouth area of the artistic face well, and the nose area also appears different degrees of tilt; in the enlarged image of the eye area, HR-Net and FOA miss some corner eye boundaries, while the landmark point detection method of the present application is more accurate in marking each face area.

[0032] The above is a further detailed description of the present application in combination with specific / preferred embodiments, and cannot be regarded as limiting the specific implementation of the present application to these descriptions. For ordinary skilled persons in the technical field to which the present application belongs, without departing from the concept of the present application, they can make several alternatives or modifications to the described embodiments, and these alternatives or modifications shall be regarded as belonging to the protection scope of the present application.

[0033] The part of the present application not described in detail belongs to the technology known to those skilled in the art.

Claims

1. A high-resolution artistic face landmark detection method based on deep learning, characterized by: Please follow the steps below to implement: Step 1: Prepare an artistic face dataset and corresponding landmark annotations for model training and evaluation; Step 2: Downsample the high-resolution images in the artistic face dataset, pass the downsampled low-resolution images through the global encoder-decoder network to predict the output heat map and extract the global low-resolution predicted landmark detection map; the specific method is as follows: First, the 1024×1024 resolution images in the Art Face Dataset are downsampled to 256×256 low-resolution images, which are then fed into the global encoder-decoder network. The network proceeds as follows: The input feature map has its number of channels doubled and its size reduced by two downsampling blocks consisting of a convolutional layer with a kernel size of 3 and a stride of 2, a BatchNorm normalization layer, and a Relu activation layer. The downsampled image is then processed through six ResNet blocks to enhance the model's image modeling capabilities. Two upsampling blocks, consisting of a transposed convolutional layer with a kernel of 3 and a stride of 2, a BatchNorm normalization layer, and a ReLu activation layer, halve the number of channels in the feature map and increase its size. Finally, the upsampled feature map is processed through a reflection padding layer and a convolutional layer to output a heat map with 68 channels and the same width and height as the input image. The heat map marks the location of each facial part in the form of peaks, and spatial softmax is used to extract these markers from the heat map in a differentiable manner to obtain a global low-resolution predicted marker detection map. The marked parts include the eyes, nose, mouth, and jawline. Step 3: Crop the high-resolution image based on the global low-resolution predicted landmark detection map, connect the obtained regional image with the corresponding regional heat map in the output heat map in step 2, and input it into the regional encoder-decoder network to generate the refined high-resolution predicted landmark detection map for each region; the specific method is as follows: First, the low-resolution facial landmark image obtained in step 2 is upsampled by a factor of 4 to match the high-resolution image. Regions of the high-resolution image are cropped based on the coarse landmark positions in the global low-resolution predicted landmark detection map obtained by the global encoder-decoder network. These regions correspond to the eyes, nose, and mouth. To fit the 256×256 resolution, the cropped regions are padded with random values ​​between 0.25 and 0.5 of the original region. Then, the regional image corresponding to the automatically cropped part area is extracted from the heat map output by the global encoder-decoder network; the automatically cropped part area image and the corresponding regional image extracted from the heat map are channel-fused, and the fused different part area images are respectively input into the corresponding regional encoder-decoder network, which has the same structure as the global encoder-decoder network; through channel fusion, the regional encoder-decoder network can obtain the global position of the marker point as prior information, thereby supporting the refinement task; there is no weight sharing between different regional encoder-decoder networks, so that each network can learn the specific features of its facial sub-region; the heat map output by the regional encoder-decoder network does not need to refine the jaw line part of the face, so the number of channels is reduced by 17; the heat map marks the position of each part of the face in the form of peaks, and uses spatial softmax to extract these markers from the heat map in a differentiable manner to obtain a high-resolution predicted marker point detection map for each refined region; Step 4: Optimize the global low-resolution and regional high-resolution predicted marker detection maps through mean square error loss; Step 5: Restore the high-resolution predicted landmark detection map of each region back to the global coordinate system to achieve face landmark detection.

2. The high-resolution artistic facial landmark detection method based on deep learning according to claim 1, characterized in that: Step 1: First, we collected high-resolution facial images of different artistic styles from various museums and research institutes as Subset 1. Then, we selected all artistic images containing faces from the Wikiart dataset as Subset 2. To automatically detect and crop all face regions in Subsets 1 and 2, we first detected the face regions in Subsets 1 and 2 using the object detector YOLOv5, and then performed a cropping operation on these regions. In addition, we used style transfer technology on the public dataset 300-W for facial landmark detection, scaling the natural face images in the 300-W dataset to 1024×1024×3 and converting them into artistic style face images as Subset 3. The processed subsets 1, 2, and 3 are used as the artistic face datasets for model training. The facial landmarks in the artistic face datasets are annotated semi-automatically to obtain the corresponding real landmark detection maps. The random forest-based face landmark detector dlib is used to annotate the landmarks of the artistic face dataset. The landmarks are manually corrected or re-annotated when dlib fails.

3. The high-resolution artistic facial landmark detection method based on deep learning according to claim 2, characterized in that: Step 4: For global low-resolution and regional high-resolution landmark detection tasks, the mean square error between the predicted landmark detection map and the true landmark detection map is used as the loss function, which is expressed as follows: Where λ is the weighting factor, x are the global low-resolution predicted marker detection map and the global true marker detection map; y are the regional high-resolution predicted marker detection map and the regional real marker detection map; N G The number of heatmap channels representing the output of the global encoder-decoder network is 68; N r The number of channels of the heatmap representing the output of the region encoder-decoder network is 51.

4. The high-resolution artistic face landmark detection method based on deep learning according to claim 3, characterized in that: Step 5: Restore the high-resolution predicted marker detection map of each region back to the global coordinate system; The high-resolution face landmarks are transferred from each region back to the global coordinate system. First, the bounding box coordinates and original region size of the extracted region are tracked based on the regional high-resolution predicted landmark detection map. Then, the local coordinates of the extracted region are scaled by the original region size and the offset of the bounding box is added to obtain the position of the high-resolution predicted landmark points of each region relative to the global coordinate system. For the jaw line, only the predicted landmark positions of the global predicted landmark detection map are used. Therefore, the complete face landmark prediction result is a combination of the jaw line position of the global predicted landmark detection map and the refined regional high-resolution predicted landmark positions relative to the global coordinate system.