A Personalized Shadow Puppet Face Generation Method Based on Constrained Generative Adversarial Networks
By using a constraint-based generative adversarial network (GAN) approach, the problem of texture and geometric structure loss when converting real human faces into shadow puppet faces is solved. The generated shadow puppet face images retain the features of real human faces while reflecting the style of shadow puppets, thus achieving personalized image conversion.
Patent Information
- Application Number
- CN202211411195.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-11
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2042-11-11
AI Technical Summary
Existing image conversion technologies cannot simultaneously preserve texture and geometric structure when converting real human faces into shadow puppet faces, resulting in converted images that fail to reflect the stylistic characteristics of shadow puppet faces or achieve personalization.
A constraint-based generative adversarial network (GAN) approach is adopted, which introduces a local geometric feature constraint function. Through the training process of the GAN, the generator generates a shadow puppet face image with some features consistent with the real human face. The loss is calculated using the local geometric feature constraint function, and the weight parameters are adjusted to control the style and personalization of the transformation result.
It achieves the preservation of the main features of a real human face during the conversion process, while reflecting the stylistic characteristics of the shadow puppet face. The generated shadow puppet face image has recognizable features of a real human face, and the degree of conversion is controllable. It can generate synthetic images of different styles according to user needs.
Smart Images

Figure CN115841418B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision, specifically to a personalized shadow puppet face generation method based on constrained generative adversarial networks. Background Technology
[0002] Shadow puppetry is an ancient traditional Chinese folk art with a wide reach, having spread abroad as early as the Yuan Dynasty. It is an important intangible cultural heritage of my country. Shadow puppet props are handmade, and performers manipulate the puppets by hand, accompanied by local tunes to tell stories, often with percussion or string instruments. It is a comprehensive art form with diverse expressive means and rich content. However, the advent of the new media era has shifted public attention to digital media, impacting shadow puppetry, which relies on oral instruction and handcrafted performance, and hindering its transmission. How to regain public attention for shadow puppetry and ensure its effective inheritance is a pressing issue that needs to be addressed in the new era.
[0003] Image-to-image conversion based on deep learning technology provides an effective means for the production and inheritance of shadow puppetry. Creating shadow puppet props with realistic characteristics based on live-action images, and then producing performance content that allows for immersive character portrayal, has become a potential form of promoting shadow puppetry, as well as a new means of its protection and inheritance.
[0004] Current image-to-image conversion technologies involve both texture and geometric structure conversion. However, typical image conversion techniques usually only involve one of these conversions to ensure that the source image of the converted image can still be identified through certain features before and after conversion; that is, the converted image needs to maintain the consistency of certain features. Considering the significant differences between real human faces and shadow puppet faces—the former having rich textures and varied features, while the latter having simple textures and a single pattern—converting a real human face to a shadow puppet face requires simultaneous conversion of both texture and geometric structure. However, the loss of texture or geometric structure information during this process may result in the converted image lacking certain identifiable features of the source image. Specifically, if the converted result retains the texture features of the source image, it cannot fully reflect the stylistic characteristics of the shadow puppet face; if the conversion process lacks constraints on the geometric structure of the real human face, partial conversion cannot be achieved, and complete conversion cannot reflect the individuality of the shadow puppet face. Therefore, current image conversion technologies still have shortcomings in the field of generating shadow puppet faces. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of current image-to-image conversion methods in handling transformations that simultaneously involve texture and geometric structure. This invention provides a constrained conversion method that utilizes a constrained generative adversarial network model. While ensuring that the conversion result reflects the stylistic characteristics of shadow puppet faces, it constrains the geometric deformation of different conversion parts as needed, thereby preserving typical facial geometric features and reflecting the individuality of shadow puppet faces.
[0006] This application is achieved through the following technical measures: a personalized shadow puppet face generation method based on constrained generative adversarial networks (GANs). A GAN is established to convert real human face images into shadow puppet face images. This GAN introduces local geometric feature constraint functions, calculates the loss through these functions to generate the required image domain, and obtains a GAN model that meets the requirements by constraining the geometric deformation of typical facial features during the training process. A real human face image is input into the generator of this GAN model to generate a shadow puppet face image whose features are consistent with those of the real human face.
[0007] Preferably, the typical facial features include eyes, nose, mouth, and overall facial contour.
[0008] Preferably, the three parts, eyes, nose, and mouth, are determined by dividing the image into different blocks and remain fixed during training, while the overall contour uses the entire image as the accumulation area.
[0009] Preferably, the loss function of the generative adversarial network is:
[0010]
[0011] In equation (1), G represents the generator and D represents the discriminator. That is, the generated personalized shadow puppet face, V C (D,G) represents the constrained generative adversarial network loss function, and V(D,G) represents the original generative adversarial network loss function. λ is the local geometric feature constraint function, and λ is the weight adjustment parameter. The consistency between the shadow puppet face image and the real human face is adjusted by adjusting λ.
[0012] in:
[0013]
[0014]
[0015] In equation (2), x represents a sample in the real human face image domain X, and y represents a sample in the shadow puppet face image domain Y;
[0016] In equation (3), x′ corresponds to x used in the current iteration, and is the edge image of x, with a loss. From x′ and the corresponding The pixel differences of different corresponding regions are accumulated and calculated as shown in Equation (4), where λ1, λ2, λ3, and λ4 are proportional coefficients.
[0017] in,
[0018]
[0019] In equation (4), (t) represents any part of the eyes, nose, mouth and overall outline, and p represents the pixel position in the region.
[0020] Preferably, by adjusting λ1, λ2, λ3, and λ4 to constrain the geometric deformation of typical facial features, the weights of different facial regions in the converted image are adjusted to generate an image that meets the requirements.
[0021] Preferably, the dataset used to train the generative adversarial network is a standardized dataset that can be directly used for training the network, obtained by preprocessing real human face datasets and shadow puppet face datasets.
[0022] Preferably, the preprocessing process includes: removing the background from each image in the real human face dataset and the shadow puppet face dataset, and setting the bottom of the image to white; adjusting the resolution of each image to a preset value; cropping the real human face region and the shadow puppet face region in each image, and adjusting the position of the real human face region and the shadow puppet face region in the image; adjusting the face in each image to a side view with a uniform orientation; and converting each image to a grayscale image.
[0023] The beneficial effects of this application are as follows: This application achieves personalized shadow puppet face generation by simultaneously transforming image texture and geometric structure through a constrained generative adversarial network-based transformation process. It preserves the main features of a real human face while retaining the style of a shadow puppet face, thus ensuring that the transformed image is a shadow puppet image with recognizable features of a real human face. This overcomes the shortcomings of existing transformation schemes, which, due to texture loss and geometric transformation, fail to reflect the typical facial features of the source image. Furthermore, through flexible parameter adjustment, it is possible to generate synthetic images with an overall style more inclined towards a real human face or more towards a shadow puppet face, based on user needs, and the degree of transformation of each major facial component is controllable. Attached Figure Description
[0024] The accompanying drawings are provided to further illustrate the present application and form part of the specification. They are used together with the embodiments of the present application to explain the application and do not constitute a limitation thereof. In the drawings:
[0025] Figure 1(a) is an example image of a real human face sample after preprocessing;
[0026] Figure 1(b) shows an example of a pre-processed shadow puppet face sample;
[0027] Figure 2 A diagram of a generative adversarial network architecture with local constraints;
[0028] Figure 3(a) shows an example of a generator network structure;
[0029] Figure 3(b) shows an example of the discriminator network structure;
[0030] Figure 4 The flowchart shows the training process of a generative adversarial network with local constraints.
[0031] Figure 5 Network structure diagram used to generate personalized shadow puppets;
[0032] Figure 6 A flowchart for generating personalized shadow puppet faces. Detailed Implementation
[0033] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0034] A personalized shadow puppet face generation method based on constrained generative adversarial networks (GANs) includes: first, establishing a GAN for converting real human face images into shadow puppet face images, wherein the loss function of the GAN has a local geometric feature constraint function; second, during the training process of the GAN, obtaining a GAN model that meets the requirements by constraining the geometric deformation of typical facial features; and third, inputting real human face images into the generator of the GAN model to generate shadow puppet face images with some features consistent with real human face images.
[0035] Specifically, the personalized shadow puppet face generation method includes the following steps:
[0036] S10: Establish a human face dataset and a shadow puppet face dataset;
[0037] S20: Preprocess the human face dataset and shadow puppet face dataset in S10 to obtain a standardized dataset that can be directly used for network training. The preprocessing process includes:
[0038] (1) Remove the background from each image in the real human face dataset and the shadow puppet face dataset, and set the bottom of the image to white to avoid other background content from disturbing the network training and causing abnormal output results;
[0039] (2) Adjust the resolution of each image to the preset value to ensure that the input images are all the same resolution, because if the resolution is different, only a part may be cropped when inputting, and it cannot be guaranteed that the complete face image is input.
[0040] (3) Extract the real face area and the shadow puppet face area from each image, and adjust the position of the real face area and the shadow puppet face area in the image to ensure that the input and output sizes remain the same.
[0041] (4) Adjust the faces in each image to a side view with a uniform orientation; the reason for adjusting the images to a side view with a uniform orientation is to maintain consistency with the characteristics of the shadow puppet images.
[0042] (5) Convert each image to grayscale.
[0043] After preprocessing, examples of real human face samples are shown in Figure 1(a), and examples of shadow puppet face samples are shown in Figure 1(b).
[0044] S30: Establish a generative adversarial network for converting real human face images into shadow puppet face images, such as... Figure 2 As shown, where:
[0045] X represents the real human face image domain;
[0046] Y represents the shadow puppet face image domain;
[0047] This represents the personalized shadow puppet face image domain generated by the generator;
[0048] G stands for generator in a generative adversarial network;
[0049] D stands for the discriminator in a generative adversarial network;
[0050] X′ represents the real human face edge image domain formed by edge extraction of samples in X, and the samples in X have a one-to-one correspondence with the samples in X;
[0051] C represents the feature constraint function, which is composed of four local features: eye, nose, mouth, and overall contour.
[0052] The established generative adversarial network (GAN) includes the generator G and discriminator D from the original GAN, as well as the feature constraint function C introduced in this invention, and the image domain X′ required for calculating the loss through C. In this architecture, the generator G and discriminator D can employ different network structures, and the algorithm for generating X′ from X can be implemented using an edge extraction algorithm or a neural network. This embodiment provides one possible implementation. An example of the network structure of the generator G is shown in Figure 3(a). The ResNetBlock network in Figure 3(a) is a basic block, consisting of two 3×3 convolutional layers. An example of the network structure of the discriminator D is shown in Figure 3(b). Samples in the X′ domain can be obtained using the Canny edge extraction operator from the OpenCV library, or through a pre-trained DeepAnalogy neural network.
[0053] The training process of this generative adversarial network follows the general training process of neural networks, such as... Figure 4 As shown. The loss function is:
[0054]
[0055] In equation (1), G represents the generator and D represents the discriminator. That is, the generated personalized shadow puppet face, V C (D,G) represents the constrained generative adversarial network loss function, and V(D,G) represents the original generative adversarial network loss function. Let λ be the local geometric feature constraint function, and let λ be the weight adjustment parameter. The local geometric feature constraint loss is adjusted by the parameter λ. The weight relationship between the image and the generative adversarial network loss V(D,G) is used to adjust the consistency between the shadow puppet face image and the real human face.
[0056] in:
[0057]
[0058]
[0059] In equation (2), x represents a sample in the real human face image domain X, and y represents a sample in the shadow puppet face image domain Y;
[0060] In equation (3), x′ corresponds to x used in the current iteration, and is the edge image of x, with a loss. From x′ and the corresponding The pixel differences between different corresponding regions are accumulated, and the calculation method is shown in Equation (4).
[0061] in,
[0062]
[0063] In equation (4), (t) represents any one of the four parts: eyes, nose, mouth, and overall contour, and p represents the pixel position in the region. λ1, λ2, λ3, and λ4 are scaling coefficients. By adjusting λ1, λ2, λ3, and λ4, the geometric deformation of typical facial features is constrained, and the weight of different facial regions in the transformed image is adjusted to generate an image that meets the requirements. Among them, the eyes, nose, and mouth can be determined by dividing the image into different blocks and remain fixed during training, while the overall contour uses the entire image as the accumulation region.
[0064] S40: Input a real human face image into the generator of the trained generative adversarial network model to obtain a shadow puppet face image with some features consistent with the real human face, such as... Figure 5 As shown in / 6.
[0065] This application controls whether the overall shape and four key parts retain their geometric shape after conversion by adjusting parameters λ, λ1, λ2, λ3, and λ4. Since shadow puppet images only contain edge curves, typical facial features need to be represented through geometric structures. In this application, typical facial features refer to the eyes, nose, mouth, and overall outline. These parts distinguish individuals, reflecting the individuality of the shadow puppet facial image. This overcomes the shortcoming of existing image conversion techniques, which cannot determine which face the shadow puppet image is based on. It should be noted that individualization means that the key parts of the converted shadow puppet facial image can be distinguished as being derived from a real human face image, thus reflecting the appearance of a real person.
[0066] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A personalized shadow puppet face generation method based on constrained generative adversarial networks, characterized in that, A generative adversarial network (GAN) is established for converting real human face images into shadow puppet face images. This GAN introduces a local geometric feature constraint function, and calculates the loss through the local geometric feature constraint function to generate the required image domain. During the training process of the GAN, the geometric deformation of typical facial features is constrained to obtain a GAN model that meets the requirements. A real human face image is input into the generator of this GAN model to generate a shadow puppet face image with some features consistent with the real human face. The loss function of the generative adversarial network is: In equation (1), G represents the generator and D represents the discriminator. That is, the generated personalized shadow puppet face, V C (D,G) represents the constrained generative adversarial network loss function, and V(D,G) represents the original generative adversarial network loss function. λ is the local geometric feature constraint function, and λ is the weight adjustment parameter. The consistency between the shadow puppet face image and the real human face is adjusted by adjusting λ. in: In equation (2), x represents a sample in the real human face image domain X, and y represents a sample in the shadow puppet face image domain Y; In equation (3), x′ corresponds to x used in the current iteration, and is the edge image of x, with a loss. From x′ and the corresponding The pixel differences of different corresponding regions are accumulated and calculated as shown in Equation (4), where λ1, λ2, λ3, and λ4 are proportional coefficients. in, In equation (4), (t) represents any part of the eyes, nose, mouth and overall outline, and p represents the pixel position in the region.
2. The personalized shadow puppet face generation method based on constrained generative adversarial networks according to claim 1, characterized in that, The typical facial features include the eyes, nose, mouth, and overall facial contours.
3. The personalized shadow puppet face generation method based on constrained generative adversarial networks according to claim 2, characterized in that, The three parts—eyes, nose, and mouth—are determined by dividing the image into different blocks and remain fixed during training. The overall contour is achieved by accumulating the entire image as a region.
4. The personalized shadow puppet face generation method based on constrained generative adversarial networks according to claim 1, characterized in that, By adjusting λ1, λ2, λ3, and λ4 to constrain the geometric deformation of typical facial features, the weights of different facial regions in the transformed image are adjusted to generate an image that meets the requirements.
5. The personalized shadow puppet face generation method based on constrained generative adversarial networks according to claim 1, characterized in that, The dataset used to train the generative adversarial network is a standardized dataset that can be directly used for training the network, obtained by preprocessing real human face datasets and shadow puppet face datasets.
6. The personalized shadow puppet face generation method based on constrained generative adversarial networks according to claim 5, characterized in that, The preprocessing process includes: Remove the background from each image in the real human face dataset and the shadow puppet face dataset, and set the bottom of the image to white; Adjust the resolution of each image to the preset value; Extract the real face region and the shadow puppet face region from each image, and adjust the position of the real face region and the shadow puppet face region in the image; Adjust the faces in each image to a side view with a uniform orientation; Convert each image to grayscale.
Citation Information
Patent Citations
Paper-cut generation system and method based on generative adversarial neural network
CN114170340A
Method employing generative adversarial network for predicting face change
WO2020029356A1