Intelligent Generation Method from Photos of Facial Shapes and Facial Features to Exaggerated Comics

Through the StyleGAN-based network and layered exaggeration module, the problem of insufficient exaggeration of facial features in the existing technology is solved, and high-quality exaggeration of facial features and personalized comic generation is achieved.

CN116386115BActive Publication Date: 2025-07-08TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310360128.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-06
Publication Date
2025-07-08
Estimated Expiration
2043-04-06

AI Technical Summary

Technical Problem

The existing comic generation methods are difficult to exaggerate the high-quality facial features of the photo while maintaining their identity, especially the partial facial features, and the reference comic selection lacks personalized matching.

Method used

A network based on StyleGAN is adopted, through global and local exaggeration modules, combined with the hierarchical structure design of global face shape and local five-feature characteristics, the exaggeration network in the facial contour, eyes, nose, mouth and other areas is trained, and local features are matched and exaggerated by reference comics.

Benefits of technology

It achieves high-quality exaggeration of the facial features of the photo, maintains the identity information of the photo, and can adjust local characteristics as needed to generate more personalized exaggerated comics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116386115B_ABST
    Figure CN116386115B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent generation method for exaggerating photos from facial shapes to facial features into exaggerated comics. A network based on StyleGAN is established, and the local facial features of the photo are significantly exaggerated with reference to the reference comics, including the following steps: in the selection of reference comics, first, the global facial shape is matched, then the significant local facial attributes are matched, and finally, the rotation angle is screened to find a comic that better conforms to the personalized facial features for reference; a hierarchical structure design from global to local is adopted, and the exaggerated global network for the entire facial contour and the exaggerated local networks for local areas such as eyes, nose, and mouth are trained respectively, and different local areas are exaggerated with reference to different reference comics. The present invention can select different reference comics to exaggerate the facial shape and facial features respectively according to the facial feature of the photo and the user's opinion, significantly exaggerating the local facial features while avoiding the distortion phenomenon caused by stretching traditional feature points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of comic face generation, and particularly to an intelligent generation method for exaggerating comics from photos of face shapes to facial features. Background Art

[0002] Drawing a comic is a process of using sketches, pencil drawings or other artistic paintings to simplify or exaggerate the attributes of the subject. The exaggerated facial features convey satire and humor, while allowing people to immediately recognize themselves. However, only a few celebrities can get their own comics, because it takes a professional artist hours or even days to create a comic. Automatic generation methods can easily generate comics, but they are often limited to changing the overall face shape of the comic and cannot design exaggerated facial features as meticulously as a painter. Therefore, it is very necessary to develop a comic generation method that combines the whole and the part.

[0003] Early comic generation methods mainly included methods based on graphics and image-to-image conversion. The graphic-based method only focused on exaggeration, while the image-to-image conversion method only paid attention to texture rendering. Although subsequent works based on generative adversarial networks (GANs) achieved both functions simultaneously, the pictures they generated had poor quality. The recent work StyleCariGAN fine-tuned StyleGAN and proposed a method for generating vivid comics. This method can generate high-quality comics but cannot achieve multiple exaggerations by referring to comics. Subsequently, DualStyleGAN proposed a reference style map generation method based on StyleGAN, which can not only generate high-quality comics but also refer to different comic exaggeration styles.

[0004] However, although these StyleGAN-based methods can generate high-quality comics, it is challenging to maintain identity when the photo features are complex. On the other hand, all of the above methods mainly focus on the exaggeration of facial contours and it is difficult to significantly exaggerate the facial features. In addition, the current StyleGAN-based methods can only randomly select reference comics and do not summarize and match the face shape and facial feature characteristics of the photo to a suitable reference comic. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent generation method for exaggerating comics from photos of face shapes to facial features, and a new technical route for assisting in the creation of exaggerated comics by using global and local exaggeration modules is given. This method is simple and convenient to operate, and can arbitrarily deform the local facial features to produce various expression effects.

[0006] To achieve the above purpose, the present invention provides the following solutions:

[0007] An intelligent generation method for exaggerating photos from facial shapes to facial features into exaggerated comics, which establishes a network based on StyleGAN and significantly exaggerates the local facial features of the photo with reference to the reference comics, including the following steps:

[0008] S1. In the selection of reference comics, first match the global facial shape, then match the local significant facial attributes, and finally screen the rotation angle to find a comic that better conforms to the personalized facial features for reference;

[0009] S2. Adopt a hierarchical structure design from global to local, and train the exaggerated global network of the entire facial contour and the exaggerated local networks of local areas such as eyes, nose, and mouth respectively, and exaggerate different local areas with reference to different reference comics.

[0010] Further, in step S1, in the selection of reference comics, first match the global facial shape, then match the local significant facial attributes, and finally screen the rotation angle to find a comic that better conforms to the personalized facial features for reference, which specifically includes:

[0011] S101. Classification of the facial shapes of comics:

[0012] Inverse map the comic to a fixed geometric shape, and then match the geometric shapes of the photo and the comic respectively to achieve a suitable reference comic for the photo; wherein, the geometric shapes include five types: triangle, inverted triangle, rectangle, rhombus, and hourglass.

[0013] S102. Mathematical representation of the facial shape:

[0014] Determine the facial contour set, calculate the probabilities of the photo and the comic belonging to each geometric shape in the facial contour set, and perform a mathematical representation of the facial shape;

[0015] S103. Match the comic with the closest facial shape to the photo:

[0016] According to the criteria determined in S102, calculate the similarity between the photo and the th geometric shape to calculate the distance between the photo and the comic, and screen out the comic dataset with a facial shape similar to the photo;

[0017] S104. Match the comic with the closest facial feature to the photo:

[0018] To address the problem of matching local reference comics, an attribute feature matching method is proposed. First, an attribute classifier is used to detect the probabilities of 20 cartoon attributes related to the face, eyes, nose, and mouth. Then, the attribute a with the highest probability is found, and all comics with the probability of attribute a greater than the threshold M are used as local reference comics in the comic dataset, and the most prominent feature photos in the local reference comics are matched. Among them, by changing the required matching attribute a, other parts of the photo are exaggerated.

[0019] Furthermore, in step S2, a hierarchical structure design from global to local is adopted. An exaggerated global network for the entire facial contour and exaggerated local networks for local regions such as eyes, nose, and mouth are trained respectively. Different local regions are exaggerated with reference to different reference comics, which specifically includes:

[0020] S201, Production of the reference comic dataset:

[0021] Using Webcaricature as the reference comic dataset, through the de - stylization method in DualStyleGAN, the comics are converted into corresponding photos to obtain a corresponding photo dataset. 5905 comics are used as the training set and 132 as the test set.

[0022] S202, Processing of color texture:

[0023] For low - resolution two - dimensional feature maps, a color control module is designed. The low - resolution part of the style map of the comic passes through the color control module to generate a residual, and then the residual is added to the low - resolution part of the style map of the photo to obtain a two - dimensional feature map of the texture style of the generated comic.

[0024] S203, Processing of shape:

[0025] For high - resolution two - dimensional feature maps, a hierarchical exaggeration module is designed, corresponding to two - dimensional feature maps of different resolutions. The style map of the comic passes through these exaggeration modules to generate a set of exaggerated feature maps. In the processing of shape, a segmentation map is used to achieve exaggeration of different regions. First, for the global generator, the facial region segmentation map is multiplied by the exaggerated feature map and then added to the feature map of the photo to generate a comic with an exaggerated overall face shape. For the local facial feature generators of eyes, nose, and mouth, the segmentation maps of these three parts are multiplied by the exaggerated feature map and added to the photo feature map respectively to achieve exaggeration of local regions.

[0026] The intelligent generation method of exaggerated comics from face shapes to facial features in photos provided by the present invention can exaggerate the whole and local parts of the photo based on a reference comic. The present invention adopts a hierarchical network structure, including a global exaggeration module and a local exaggeration module, which exaggerate the facial shape and facial features of the photo respectively. In addition, for the design of the exaggeration module, a hierarchical network structure is designed for feature maps of different resolutions respectively to control features from rough to fine. In general, the technical effects of the present invention can be summarized as follows:

[0027] 1) An exaggeration module from the whole to the local is proposed, which can better exaggerate the facial features and achieves an effect beyond the existing methods;

[0028] 2) By using two-dimensional feature maps as comic features, the identity information of the photo can be better maintained;

[0029] 3) A reference comic matching method from the whole to the local is proposed, which can select a reference comic that matches the overall and local features of the photo. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0031] Figure 1 It is a flowchart of the intelligent generation method of exaggerated comics from face shapes to facial features in photos of the present invention;

[0032] Figure 2 It is a schematic diagram of the matching of face shapes and facial features in the present invention;

[0033] Figure 3 It is a schematic diagram of the exaggeration network in the present invention;

[0034] Figure 4 It is a schematic diagram of partial generation results of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0036] The object of the present invention is to provide a method for intelligently generating exaggerated comics from photos of facial shapes to facial features, and a new technical route for assisting in the creation of exaggerated comics by using global and local exaggeration modules is given. This method is simple and convenient to operate, and can arbitrarily deform local facial features to produce various expression effects.

[0037] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0038] As Figure 1 shown, the present invention provides a method for intelligently generating exaggerated comics from photos of facial shapes to facial features, establishing a network based on StyleGAN, and significantly exaggerating the local facial features of the photo with reference to a reference comic, including the following steps:

[0039] S1. In the selection of the reference comic, first match the global facial shape, then match the local significant attributes of the face, and finally screen the rotation angle, so as to find a more suitable reference comic for the individual facial features;

[0040] S2. In the design of the method for intelligently generating exaggerated comics from photos of facial shapes to facial features, a hierarchical structure design from global to local is adopted, and the exaggerated global network of the entire facial contour and the exaggerated local networks of local areas such as eyes, nose, and mouth are trained respectively. This method can not only better highlight the local features of the photo, but also exaggerate different local areas with reference to different reference comics.

[0041] Among them, in step S1, in the selection of the reference comic, first match the global facial shape, then match the local significant attributes of the face, and finally screen the rotation angle, which specifically includes:

[0042] S101. Classification of the facial shape of the comic:

[0043] The overall contour of the face determines the facial features of a person. For artists, the first step in comic creation is to determine the overall contour of the face. Master artists often convert complex human faces into geometric shapes they are familiar with as the basic contour of comic works. The present invention refers to this process, inversely maps the comic to a fixed geometric shape. Then, the geometric contours of the photo and the comic are matched respectively, realizing that the photo corresponds to a suitable reference comic.

[0044] However, converting a human face into complex geometric shapes in the eyes of an artist is a challenge that is difficult for a computer to achieve. Therefore, the present invention maximizes the complex geometric shapes in the eyes of an artist and generalizes them into a set of simple geometric shapes. That is, the mapping from a comic contour to a single complex shape is transformed into a mapping from a comic to a set of simple geometric shapes. Since there is currently no classification standard for the contours of a comic dataset, the present invention refers to the face shape classification of real human faces and finally generalizes the comic face shapes into five types: triangle, inverted triangle, rectangle, rhombus, and hourglass.

[0045] S102, perform a mathematical representation of the facial shape:

[0046] According to the standard determined in S101, determine the facial contour set. Calculate the probabilities that the photo and the comic belong to each basic shape in the facial contour set, perform a mathematical representation of the facial shape, and achieve a global face shape pre-matching between the comic and the photo. For any geometric shape S, divide it into three equal parts along the direction perpendicular to the central axis, and record the areas of the three parts as a vector , which is used to describe the characteristics of this geometric shape. For the facial contour set , belongs to [0, 5], and the feature vector matrix of the basic shape can be calculated. The area ratios of each basic shape in the facial contour set are specified as follows: (triangle), (inverted triangle), (rectangle), (rhombus), (hourglass). In addition, for the input photo , For the comic in the comic set, . Thus, the facial shape has been mathematically represented, which is convenient for further calculating the combination representation of any geometric shape by each basic shape.

[0047] S103, match the comic with the face shape of the photo that is closest:

[0048] According to the standard determined in S102, determine the facial contour set. The similarity between the photo and the th basic geometric shape can be calculated. Therefore, the probability that the photo belongs to the th basic geometric shape. The shape probability vector composed of these 5 probabilities is called . Similarly, the shape probability vector of the comic is called .

[0049] From this, the distance between the photo and the comic can be calculated . Assume that the distance between the two is less than When this is the case, this comic can be used as a reference for this photo. For each photo, enumerate each comic in the dataset and calculate this distance, and a comic dataset with a facial shape similar to that of the photo can be filtered out for the user to make further selections.

[0050] S104, match the comic with the most similar facial features of the photo:

[0051] Regarding the matching problem of local reference comics, an attribute feature matching method is proposed. First, use an attribute classifier to detect the probabilities of 20 cartoon attributes related to the face, eyes, nose, and mouth. Then, find the attribute a with the highest probability, and all comics with the probability of attribute a greater than the threshold M are used as local reference comics in the comic dataset. This method can match the most prominent feature photos in the local reference comics. It can also, according to the needs of the user, change the required matching attribute a to exaggerate other parts of the photo.

[0052] Figure 2 The pre-matching process of the global facial shape described above is given. For the reference comic and the input photo, map them from the regional space (defined by 3 regions) to the basic structure space (defined by the probabilities of 5 basic shapes), and then calculate the similarity between them. Below the dotted line are five basic shapes. The numbers in the geometric figures represent the proportions of these three regions.

[0053] Among them, in step S2, a hierarchical structure design from global to local is adopted, and an exaggerated global network for the entire facial contour and exaggerated local networks for local regions such as the eyes, nose, and mouth are trained respectively. The method of the present invention can not only better highlight the local features of the photo, but also exaggerate different local regions with reference to different reference comics, specifically including:

[0054] S201, production of the reference comic dataset:

[0055] Currently, there is no large-scale dataset for pairing photos and comics. Therefore, use Webcaricature as the reference comic dataset, and through the de-stylization method in DualStyleGAN, convert the comics into corresponding photos to obtain the corresponding photo dataset. For example, 5905 comics are used as the training set and 132 as the test set.

[0056] S202, processing of color texture:

[0057] For the low-resolution two-dimensional feature map, design a color control module. Pass the low-resolution part through the color control module to generate . Then add to For the low-resolution part, a two-dimensional feature map of the texture style for generating a comic is obtained.

[0058] S203, Processing of the shape:

[0059] For the high-resolution two-dimensional feature map , a hierarchical exaggeration module is designed, corresponding to two-dimensional feature maps of different resolutions respectively. Through these exaggeration modules, a set of exaggerated feature maps is generated. In the processing of the shape, a segmentation map is adopted to achieve the exaggeration of different regions. First, for the global generator, the facial region segmentation map is multiplied by the exaggerated feature map and then added to the feature map of the photo to generate a comic with an exaggerated overall face shape. For the local facial feature generator, the segmentation maps of the three regions of the eyes, nose, and mouth are respectively multiplied by the exaggerated feature map and added to the photo feature map to achieve the exaggeration of the local region.

[0060] Adding a local exaggeration module is a challenging task because the feature w+ extracted from the two-dimensional feature map is for the entire photo, and the facial region occupies most of the space in the photo. Based on the following considerations, the present invention uses the two-dimensional feature map as the comic feature: (a) It can achieve local exaggeration; (b) It can retain the position information of the facial features in the photo; (c) It can better maintain the identity of the photo. A new problem is also found in this process: when the features of the photo do not match those of the reference comic, the personal characteristics of the photo cannot be better highlighted. For example, when the face shape of the photo is an inverted triangular face with a large forehead, while the face shape of the reference comic is an extremely long face, it is difficult to highlight the face shape characteristics of the photo. The same is true for local features. When the photo has the attribute of a large nose, a reference comic with a large nose should also be used to better highlight the personal characteristics. To this end, the present invention proposes a new global-to-local reference comic matching algorithm, which can not only find a global reference comic with similar global face shape features, but also find a local reference comic with matching local attribute characteristics.

[0061] Figure 3 The network structure of the present invention is given. The photo p and the reference comic c are respectively input into the trained photo encoder and comic encoder to obtain the style map of the photo ( ) and the style map of the comic ( ). The texture module (orange) and the global module (blue) take as inputs to generate residuals, and the residuals are added to (multiplied by the global mask of the global module) to generate the target style map, and the globally comic is generated from the pre-trained StyleMapGAN. The 3 local modules (green) are trained in the same way with local masks to exaggerate the facial features.

[0062] Design description of loss function:

[0063] The segmentation map of the defined region, where . The and 's loss is used as the reconstruction loss: wherein,

[0064]

[0065] where, , . In this article, is set to 0.3. and are the multiplication and addition of matrix elements respectively.

[0066] Define a masked version of the LPIPS loss between and to ensure that still retains the facial feature area of the photo:

[0067] ;

[0068] where is the mask of the facial feature , is the generated comic by , is the generated comic by .

[0069] To make the generated comic maintain the prominent attributes of the input photo, use the attribute CLIP loss as part of the identity loss. Use a face attribute classifier trained on the WebCariA dataset to classify the attributes of the input photo. This dataset annotates 50 attributes describing facial shape features, such as the size of facial components. The most important (most likely) attribute atrr is found as the text description input for the CLIP loss, which guides the generator to minimize the cosine distance between the photo and the output comic in the CLIP space. Finally, the identity loss function is defined as:

[0070] ;

[0071] where, is the face recognition identity loss defined using the pre-trained ArcFace network in StyleClip.

[0072] Finally, the total loss of the network is:

[0073] ;

[0074] wherein , and balance multiple loss functions.

[0075] The present invention proposes a method for intelligently generating exaggerated comics from photos of facial shapes to facial features. By extracting comic features into a two-dimensional feature map, the exaggeration of facial shapes and facial features is combined to better highlight the facial features of the photos. In order to match the facial structure and facial features of comics and photos, a two-stage matching method is proposed, which greatly enhances the effect of comic generation. A large number of experimental results show that the method of the present invention can produce comics at a higher level than early works. Some experimental results are as Figure 4 shown. The method for generating comics from facial shapes to facial features provided by the present invention can potentially be applied to other tasks, such as more general image-to-image translation.

[0076] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0077] In this application, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for intelligently generating exaggerated comics from photos of facial shapes to facial features, characterized in that, Build a network based on StyleGAN to significantly exaggerate the local facial features of a photo with reference to a reference comic, including the following steps: S1. In the selection of the reference comic, first match the global face shape, then match the significant local facial attributes, and finally screen the rotation angle to find a comic that better conforms to the facial personalized features for reference; S2. Adopt a hierarchical structure design from global to local, and train the exaggerated global network for the entire facial contour and the exaggerated local networks for local areas such as eyes, nose, and mouth respectively, and exaggerate different local areas with reference to different reference comics; specifically including: S201. Production of the reference comic dataset: Use Webcaricature as the reference comic dataset, convert the comic into the corresponding photo through the de-stylization method in DualStyleGAN to obtain the corresponding photo dataset, and use 5905 comics among them as the training set and 132 as the test set; S202. Processing of color texture: For the low-resolution two-dimensional feature map, a color control module is designed to process the low-resolution part of the style map of the comic. The low-resolution part passes through the color control module to generate a residual, and then the residual is added to the low-resolution part of the style map of the photo to obtain a two-dimensional feature map of the texture style for generating the comic. S203. Processing of shape: For high-resolution two-dimensional feature maps, hierarchical exaggeration modules are designed, corresponding to two-dimensional feature maps of different resolutions, and the style map of the comic. Through these exaggeration modules, a set of exaggerated feature maps is generated; in the processing of shapes, a segmentation map is used to achieve exaggeration of different regions. First, for the global generator, the facial region segmentation map is multiplied by the exaggerated feature map and then added to the feature map of the photo to generate a comic with an exaggerated overall face shape; for the local facial feature generator, the segmentation maps of the three regions of the eyes, nose, and mouth are respectively multiplied by the exaggerated feature map and added to the photo feature map to achieve exaggeration of the local regions.

2. The intelligent generation method of an exaggerated comic from a photo of a face shape to facial features according to claim 1, characterized in that In the above step S1, in the selection of the reference comic, first match the global face shape, then match the significant local facial attributes, and finally screen the rotation angle to find a comic that better conforms to the facial personalized features for reference, specifically including: S101. Classification of the comic facial shape: Inverse-map the comic to match it to a fixed geometric shape, and then match the geometric shapes of the photo and the comic respectively to make the photo correspond to a suitable reference comic; among them, the geometric shape includes five types: triangle, inverted triangle, rectangle, rhombus, and hourglass; S102. Mathematical representation of the facial shape: Determine the facial contour set, calculate the probabilities of the photo and the comic belonging to each geometric shape in the facial contour set, and perform a mathematical representation of the facial shape; S103. Match the comic with the closest face shape to the photo: Calculate the photo according to the criteria determined in S102 and the th geometric shape to calculate the similarity, thereby calculating the distance between the photo and the comic, and screening out the comic data set with a face shape similar to that of the photo; S104. Match the comic with the closest facial feature to the photo: For the matching problem of local reference comics, an attribute feature matching method is proposed. First, use an attribute classifier to detect the probabilities of 20 cartoon attributes related to the face, eyes, nose, and mouth; then, find the attribute a with the highest probability, and all comics with the probability of attribute a greater than the threshold M are used as local reference comics in the comic dataset, and match the most prominent feature photo in the local reference comics; among them, other parts of the photo are exaggerated by changing the required matching attribute a.

Citation Information

Patent Citations

  • Intelligent generation method of exaggerated cartoon of facial photo with editable expressions

    CN115239549A

  • Caricature exaggeration

    US20050212821A1