Cartoonized digital human generation method and device, medium and computer program product
By generating cartoonized digital people on electronic devices with weak computing power, using pre-bound skeleton models and face models, combined with a full convolutional neural network for face detection and reconstruction, the problem of insufficient computing power is solved and customized cartoonized digital people is achieved.
Patent Information
- Application Number
- CN202411251928.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2044-09-06
AI Technical Summary
The existing cartoonized digital life generation methods require high computing power and are difficult to apply on electronic devices with weak computing power, such as mobile phones and tablets.
By acquiring user face images to generate cartooned face models and combining them with pre-bound skeleton models, the calculation needs are reduced, and a full convolutional neural network is used to perform face detection and reconstruction parameter prediction, reducing the amount of calculation.
It realizes the generation of customized cartoon digital people on electronic devices with weak computing capabilities, providing a unique sense of experience and healing, and reducing computing needs.
Smart Images

Figure CN120472055A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device, medium, and computer program product for generating a cartoonized digital human. Background Art
[0002] With the development of rendering and artificial intelligence technologies, digital images are constantly developing in fields such as augmented reality (AR), games, and animation. Among them, cartoon-like digital humans are popular among users due to their cute appearance. Currently, in the process of creating cartoon-like digital humans, users can create basic geometric shapes in the drawing software of electronic devices and perform operations such as stretching, squeezing, and rotating these basic geometric shapes to produce the body, head, limbs, and other parts of the cartoon-like digital human. Then, users can add facial features (such as eyes, nose, and mouth), wrinkles in clothing, and textures in hairstyles to complete the creation of the cartoon-like digital human.
[0003] However, the above method has relatively high requirements on the computing power of electronic devices, and is therefore difficult to apply to electronic devices with relatively weak computing power, such as mobile phones and tablets. Summary of the Invention
[0004] To address the problem that existing cartoon-like digital human production methods are difficult to apply to electronic devices with relatively weak computing power, embodiments of the present application provide a cartoon-like digital human production method, device, medium, and computer program product, including:
[0005] In a first aspect, an embodiment of the present application provides a method for generating a cartoonized digital human, which is applied to an electronic device, comprising: acquiring a first image, the first image including at least a user's face; generating a cartoonized face model based on the user's face in the first image; and combining the cartoonized face model with a first cartoonized torso model in a torso material library to obtain a cartoonized digital human.
[0006] It can be understood that the first cartoon torso model in the torso library has completed the bone binding.
[0007] Based on the above solution, by directly selecting a cartoon torso model with skeleton binding and combining it with a cartoon face model from the torso material library, the computing power of the electronic device can be reduced, and it can be applied to electronic devices with relatively weak computing power, such as smartphones and tablets.
[0008] In some optional implementations, the first cartoonized torso model can be a torso model obtained by image rendering of torso parameters selected by the user from a torso material library, or a torso model obtained by image rendering of torso parameters matched by the electronic device based on the user's full-body image, or a torso model obtained by image rendering of torso parameters recommended by the electronic device based on the user's habits, and the embodiments of the present application do not make specific limitations on this.
[0009] In some optional implementations of the first aspect, a cartoonized face model is generated based on the user face in the first image, including: inputting the first image into a face recognition model, identifying an area corresponding to the user face from the first image; determining facial parameters of the user face based on the area corresponding to the user face in the first image; and performing image rendering based on the facial parameters to obtain a cartoonized face model.
[0010] In this embodiment of the present application, by creating a cartoonized digital human corresponding to the user based on the area corresponding to the user's face in the first image, the cartoonized digital human can be customized. In this way, a cartoonized digital human can be customized based on the user's face in any input image, providing a unique experience and a sense of healing.
[0011] In some optional implementations of the first aspect, the face recognition model includes a backbone network, a first convolution set, a second convolution set, a third convolution set, a first upsampling module, a second upsampling module, a first prediction head, a second prediction head and a third prediction head; the backbone network includes a first sub-network, a second sub-network, a third sub-network, a fourth sub-network and a fifth sub-network, the feature map output by the first sub-network is the input of the second sub-network, the feature map output by the second sub-network is the input of the third sub-network, the feature map output by the third sub-network is the feature map output by the fourth sub-network, and the feature map output by the fourth sub-network is the input of the fifth sub-network; the feature map output by the fifth sub-network is the input of the first convolution set, the feature map output by the first convolution set is the input of the first prediction head, and the prediction head is used to determine a first prediction result based on the feature map output by the first convolution set, the first prediction result including a first area in the first image, and the probability that the first area is the area corresponding to the user's face in the first image; the first The feature map output by the convolution set is the input of the first upsampling module, and the feature map output by the first upsampling module is the input of the second convolution set; the feature map output by the fourth subnetwork is the input of the second convolution set, and the feature map output by the second convolution set is the input of the second prediction head. The second prediction head is used to determine the second prediction result based on the feature map output by the second convolution set. The second prediction result includes the second area in the first image and the probability that the second area is the area corresponding to the user's face in the first image; the feature map output by the second convolution set is the input of the second upsampling module, and the feature map output by the second upsampling module is the input of the third convolution set; the feature map output by the third subnetwork is the input of the third convolution set, and the feature map output by the third convolution set is the input of the third prediction head. The third prediction head is used to determine the third prediction result based on the feature map output by the third convolution set. The third prediction result includes the third area in the first image and the probability that the third area is the area corresponding to the user's face in the first image.
[0012] In an embodiment of the present application, a fully convolutional neural network model is used for face detection. Since at least some of the convolutional layers in the fully convolutional neural network model are grouped convolutional layers and depthwise separable convolutional layers, the computing power requirements for electronic devices can be reduced. Therefore, the method can be applied to electronic devices with relatively weak computing power, such as mobile phones and tablets.
[0013] In some optional implementations of the first aspect, the area among the first area, the second area, and the third area having the highest probability of being the area corresponding to the user's face in the first image is used as the area corresponding to the user's face.
[0014] In some optional implementations of the first aspect, determining facial parameters of the user's face based on the area corresponding to the user's face in the first image includes: inputting the first image into a reconstruction parameter prediction model to determine the facial parameters of the user's face.
[0015] In some optional implementations of the first aspect, the reconstruction parameter prediction model is trained in the following manner: obtaining at least one sample image, inputting the sample image into the reconstruction parameter prediction model to be trained, and predicting the reconstruction parameters corresponding to the user's face in each sample image; performing image rendering based on the reconstruction parameters corresponding to the user's face in each sample image, and obtaining a first cartoonized face model corresponding to each sample image; mapping each first cartoonized face model to the coordinate system of the corresponding sample image, and obtaining a mapping image corresponding to each sample image; adjusting the model parameters of the reconstruction parameter prediction model to be trained based on the degree of overlap between each sample image and the user's face in the corresponding mapping image, until the degree of overlap meets the termination condition.
[0016] It can be understood that the termination condition may be based on the fact that the degree of overlap between each sample image and the user's face in the corresponding mapping image is greater than a threshold value of overlap.
[0017] In some optional implementations of the first aspect, the overlap of the user's face in each sample image and the corresponding mapping image includes one or more of the overlap of the positions of facial key points, the overlap of the positions of eye key points, the overlap of facial areas, and the overlap of the positions of facial features.
[0018] In a second aspect, the present application provides an electronic device comprising: a memory for storing instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, for executing the cartoonized digital human generation method mentioned in the first aspect of the present application or any item of the first aspect.
[0019] In a third aspect, the present application provides a readable storage medium having instructions stored thereon. When the instructions are executed on an electronic device, the electronic device executes the method for generating a cartoonized digital human mentioned in the first aspect or any item of the first aspect of the present application.
[0020] In a fourth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions. When executed by an electronic device, the electronic device executes the computer program code of the cartoonized digital human generation method mentioned in the first aspect of the present application or any item of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A schematic diagram of an application scenario is shown;
[0022] Figure 2 According to some embodiments of the present application, a flowchart of a method for generating a cartoonized digital human is shown;
[0023] Figure 3According to some embodiments of the present application, a flowchart of another method for generating a cartoonized digital human is shown;
[0024] Figure 4 According to some embodiments of the present application, a structural schematic diagram of a face recognition model is shown;
[0025] Figure 5 According to some embodiments of the present application, a schematic diagram of the structure of a convolution set is shown;
[0026] Figure 6 According to some embodiments of the present application, a schematic structural diagram of an upsampling module is shown;
[0027] Figure 7 According to some embodiments of the present application, a schematic diagram of the structure of a prediction head is shown;
[0028] Figure 8 According to some embodiments of the present application, a schematic diagram of input and output of a reconstruction parameter prediction model is shown;
[0029] Figure 9 According to some embodiments of the present application, a schematic diagram of the hardware structure of an electronic device is shown. DETAILED DESCRIPTION
[0030] The embodiments of the present application include but are not limited to a method, device, medium, and program product for generating a cartoonized digital human.
[0031] It is understood that the method for generating a cartoonized digital human described in the embodiments of this application can be applied to electronic devices. The electronic device may also be referred to as a terminal, user equipment (UE), or mobile terminal (MT). In some specific implementations, the electronic device may be a smartphone, tablet, or the like.
[0032] It is understood that the method mentioned in the embodiment of the present application can be applied to the user's cartoon digital human customization scene, and can also be applied to the animal's cartoon pet customization scene, and the embodiment of the present application does not make any specific limitations. Figure 1 FIG2 is a schematic diagram of an application scenario. In actual application, the electronic device can generate a cartoon digital human as shown in FIG2 based on the user's face in FIG2 in FIG2 .
[0033] To address the above-mentioned problems, an embodiment of the present application provides a method for generating a cartoonized digital human. In this method, a first image is obtained, the first image including at least a user's face, and a cartoonized face model is generated based on the first image. In response to a user selecting a cartoonized torso model in a torso library, the cartoonized face model and the cartoonized torso model are combined to generate a cartoonized digital human. Each cartoonized torso model in the torso library can be stored in the form of a cartoonized torso model or in the form of torso parameters, and each part of the torso model (e.g., hands, upper body, lower body, etc.) can maintain different postures or perform different actions as the position of the skeleton changes, which is referred to as skeletal binding. In this way, by directly selecting a cartoonized torso model with good skeletal binding from the torso library and combining it with the cartoonized face model, the computing power of the electronic device can be reduced, and the method can be applied to electronic devices with relatively weak computing power, such as smartphones and tablets.
[0034] It is understood that in some optional implementations, a cartoonized face model may be generated in the following manner:
[0035] The electronic device can determine the coordinates of the bone nodes in the image coordinate system corresponding to the first image, and determine the skin weights corresponding to each pixel point in the first image relative to each bone node based on the distance between each pixel point in the first image and the bone node (for example, the pixel points corresponding to the left corner of the eye, the right corner of the eye, etc.). The closer the distance between the pixel point and the bone node, the higher the skin weight corresponding to the pixel point, and the higher the degree to which the pixel point is affected by the bone node. The coordinates of the bone node in the image coordinate system corresponding to the first image are adjusted, thereby adjusting the corresponding vertices in the initial face mesh to generate a cartoonized face model.
[0036] It is understood that in some optional implementations, the electronic device can perform face detection on a full-body image of the user's face and body to obtain a facial region. Furthermore, the electronic device can extract an input image corresponding to the facial region from the full-body image. In this way, by obtaining an input image from the full-body image of the user to create a cartoonized digital human corresponding to the user, the cartoonized digital human can be customized.
[0037] In some cartoon-like digital human generation methods, electronic devices need to perform both art design and modeling for the digital human. During the art design process, electronic devices can create cute, colorful cartoon humans for children, or stylish and trendy ones for young people. During the modeling process, electronic devices can first create basic geometric shapes and then perform operations such as stretching, squeezing, and rotating them to create the head, body, and limbs of the cartoon-like digital human. Then, electronic devices can add features such as facial features, clothing wrinkles, and hairstyle textures to complete the creation of the cartoon-like digital human.
[0038] Currently, users urgently need electronic devices that can customize and generate a corresponding cartoon-like digital human based on the user's face in any input image, providing a unique and therapeutic experience. However, using the above-mentioned cartoon-like digital human generation method to customize the user's face in the input image will result in long production time and high production costs.
[0039] To this end, with the development of artificial intelligence technology, electronic devices can use 3D reconstruction or 3D generation technologies to customize a cartoon-like digital human corresponding to the user's face in any input image. However, 3D reconstruction technologies, such as neural radiance fields (NeRF) and 3D Gaussian splatting (3DGS), require high computing power from electronic devices. Furthermore, 3D reconstruction technology itself tends to focus on realistic reconstruction. Changing the style of the generated cartoon-like digital human, i.e., performing style transfer on the cartoon-like digital human, increases the computational workload of the electronic device.
[0040] Moreover, the cartoonized digital human corresponding to the user's face in any input image customized using three-dimensional generation technology cannot be driven, that is, the generated cartoonized digital human cannot perform limb movements. In order to ensure that the generated cartoonized digital human can perform smooth limb movements, manual bone binding is often required, which results in the electronic device requiring relatively high computing power.
[0041] Furthermore, considering that mobile devices like smartphones and tablets also require customized cartoon-like digital humans corresponding to the user's face in any input image, 3D reconstruction or 3D generation technologies require long reconstruction times, high complexity, and high deployment difficulty, making them difficult to deploy on mobile devices with relatively weak computing power and limited supported operators. Therefore, creating a drivable cartoon-like digital human on mobile devices is currently a difficult problem in the field of digital image production.
[0042] To this end, an embodiment of the present application provides a method for generating a cartoonized digital human, which can be applied to mobile terminals with relatively weak computing power and limited supported operators.
[0043] The following is an introduction to the method for generating a cartoonized digital human mentioned in the embodiments of this application.
[0044] like Figure 2 The figure shows a flowchart of a method for generating a cartoon digital human. The method can be executed by an electronic device, such as the aforementioned smart phone, tablet, etc. The specific process is as follows:
[0045] 201: Face detection.
[0046] In some optional implementations, the electronic device may perform face detection on the input image to obtain an area corresponding to the user's face in the input image.
[0047] It is understood that the input image is the first image mentioned above, which may include at least the user's face. In some optional implementations, the input image may be a full-body image including the user's face and body, a half-body image including the user's face and part of the user's body, or an input image that only includes the user's face.
[0048] In some specific implementations, the electronic device may extract facial key point features associated with facial key points (e.g., eyes, nose, mouth, etc.) from an input image. The electronic device may then predict the coordinates of the facial key points in an image coordinate system corresponding to the input image based on the facial key point features, and determine the area corresponding to the user's face based on the coordinates of the facial key points in the image coordinate system corresponding to the input image.
[0049] For example, based on the facial key point features, the coordinates corresponding to the outer corner of the left eye are predicted to be (50, 80), the coordinates corresponding to the inner corner of the left eye are (80, 80), the coordinates corresponding to the inner corner of the right eye are (120, 80), the coordinates corresponding to the outer corner of the right eye are (150, 80), the coordinates corresponding to the tip of the nose are (100, 120), the coordinates corresponding to the left corner of the mouth are (70, 150), and the coordinates corresponding to the right corner of the mouth are (130, 150). Furthermore, the electronic device can determine the area corresponding to the user's face from the input image based on the distance between the facial key points and the facial contour, for example, the outer corner of the left eye is 20 from the left contour of the face, and the mouth is 30 from the bottom contour of the face, with a rectangular box with an upper left corner vertex of (30, 45) and a lower right corner vertex of (170, 180) to represent the area corresponding to the user's face.
[0050] In some embodiments, the electronic device may also perform face detection on the input image using a pre-trained face recognition model to obtain an area corresponding to the user's face in the input image.
[0051] It is understood that the specific structure and specific detection process of the face recognition model mentioned in the embodiment of the present application will be Figure 3 To avoid repetition, it is not described in detail here.
[0052] 202: Face modeling.
[0053] It is understandable that in some optional implementations, the electronic device may perform face modeling on the input image in which the area corresponding to the user's face is identified, and obtain face parameters corresponding to the user's face in the input image.
[0054] In other optional implementations, since smaller image sizes result in fewer pixels and less detail, the electronic device can crop the area corresponding to the user's face from the input image to obtain the input image. Furthermore, after obtaining the input image, the electronic device can perform facial modeling on the input image to obtain facial parameters corresponding to the user's face in the input image.
[0055] It is understood that the face parameters can be used to construct a cartoon face model. In some optional implementations, the face parameters can include shape parameters.
[0056] In some specific implementations, the electronic device can use a pre-trained reconstruction parameter prediction model to perform face modeling on the input image to obtain facial parameters. It is understood that the training process of the reconstruction parameter prediction model will be described in detail below and will not be repeated here to avoid repetition.
[0057] 203: Select torso parameters from the torso library and merge the face parameters and torso parameters.
[0058] Then, the electronic device can select torso parameters from the torso material library and merge the face parameters and torso parameters to obtain complete human body parameters.
[0059] It is understood that the torso parameters can be used to construct a cartoon torso model. The torso parameters in the torso library are divided into three parts: hands, upper body, and lower body. In some specific implementations, these torso parameters can be shape parameters generated using the Skinned Multi-Person Linear Model-Expressive (SMPL-X). Moreover, the cartoon torso model generated based on these torso parameters has completed skeletal rigging.
[0060] 204: Select a texture map from the texture library.
[0061] It can be understood that in some optional implementations, the electronic device can select a texture map from a texture material library.
[0062] In some specific implementations, the electronic device can respond to a user's selection of a texture map material in a texture library by combining the user's selected texture map material with the aforementioned complete human body parameters to obtain textured human body parameters. The texture map material can be manually designed or generated using an AI model. The texture map material can include skin, eyes, lip color, body hair, blush, eyelashes, eyebrows, beards, wrinkles, nails, etc., which are not specifically limited in the embodiments of this application.
[0063] 205: Select an accessory material from the accessory material library.
[0064] It can be understood that in some optional implementations, the electronic device can select accessory materials from an accessory material library.
[0065] In some specific implementations, the electronic device can respond to the user's selection of an accessory material in the accessory material library by combining the user's selected accessory material with the aforementioned textured human body parameters to obtain textured human body parameters with accessories. Furthermore, the electronic device can perform image rendering based on the textured human body parameters with accessories to obtain a digital human (i.e., the cartoonized digital human mentioned above). The accessory material can be manually designed or generated using an AI model. The accessory material can include hair, glasses, hats, tops, pants, shoes, socks, hand accessories, neck accessories, backpacks, etc., which are not specifically limited in the embodiments of this application.
[0066] like Figure 3 FIG2 shows a flowchart of another method for generating a cartoonized digital human. This method can be performed by an electronic device, such as the aforementioned smartphone, tablet, or other electronic device. The specific process is as follows:
[0067] 301: Acquire a first image, where the first image at least includes a user's face.
[0068] It can be understood that the first image can be any image that includes at least the user's face. For example, the first image can be a full-body image including the user's face and the user's body, or a half-body image including the user's face and part of the user's body, or a facial picture including the user's face.
[0069] 302: Generate a cartoonized face model based on the user's face in the first image.
[0070] In some optional implementations, when the first image is a full-body image including the user's face and body, the electronic device can use a pre-trained face recognition model to perform face detection on the input image to obtain the area corresponding to the user's face in the first image. Furthermore, the electronic device can crop the area corresponding to the user's face from the first image to obtain a face image. In this way, the electronic device can reduce the amount of computation while increasing the details contained in the image, thereby improving the accuracy of face detection. The electronic device can then use the trained reconstruction parameter prediction model to perform face modeling on the face image, obtain face parameters, and perform image rendering based on the face parameters to generate a cartoonized face model.
[0071] In some optional implementations, when the first image is a facial image that includes the user's face, the electronic device may directly use the trained reconstruction parameter prediction model to perform facial modeling on the facial image, obtain facial parameters, and then perform image rendering based on the facial parameters to generate a cartoon-like facial model. The specific structure and training process of the reconstruction parameter prediction model will be described in detail below and will not be repeated here to avoid repetition.
[0072] 303: Combine the cartoonized face model and the first cartoonized torso model in the torso material library to obtain a cartoonized digital human.
[0073] It can be understood that the first cartoonized torso model can be a torso model obtained by image rendering of torso parameters selected by the user from the torso material library, or a torso model obtained by image rendering of torso parameters matched by the electronic device based on the user's full-body image, or a torso model obtained by image rendering of torso parameters recommended by the electronic device based on the user's habits. The embodiments of the present application do not make specific limitations on this.
[0074] Below Figure 2 and Figure 3 The specific structure and detection process of the face recognition model mentioned in the paper are introduced in detail.
[0075] like Figure 4 FIG2 is a block diagram of a face recognition model, which may include a backbone network, a convolution set 1, a convolution set 2, a convolution set 3, an upsampling layer 1, an upsampling layer 2, a prediction head 1, a prediction head 2, and a prediction head 3.
[0076] In some optional implementations, the backbone network may be a residual network with 18 layers (ResNet-18). The backbone network may include five stages (i.e., five sub-networks): stage 1, stage 2, stage 3, stage 4, and stage 5.
[0077] Furthermore, the feature map output by stage 1 can serve as the input to stage 2, the feature map output by stage 2 can serve as the input to stage 3, the feature map output by stage 3 can serve as the input to stage 4, and the feature map output by stage 4 can serve as the input to stage 5. The feature map output by stage 5 can serve as the input to convolution set 1, and the feature map output by convolution set 1 can serve as the input to prediction head 1. Prediction head 1 can obtain prediction result 1 based on the feature map output by convolution set 1. The feature map output by convolution set 1 can also serve as the input to upsampling 1, and the feature map output by upsampling 1 can serve as the input to convolution set 2. The feature map output by stage 4 can serve as the input to convolution set 2, and the feature map output by convolution set 2 can serve as the input to prediction head 2. Prediction head 2 can obtain prediction result 2 based on the feature map output by convolution set 2. The feature map output by convolution set 2 can also serve as the input to upsampling 2, and the feature map output by upsampling 2 can serve as the input to convolution set 3. The feature map output by stage 3 can be used as the input of convolution set 3, and the feature map output by convolution set 3 can be used as the input of prediction head 3. Prediction head 3 can obtain prediction result 3 based on the feature map output by convolution set 3.
[0078] In some specific implementations, stage 1 can perform feature extraction processing on the input image to obtain a first intermediate feature map. The first intermediate feature map can then be decoded to obtain a first intermediate feature map of a higher scale. Stage 2 can perform feature extraction processing on the first intermediate feature map of a higher scale to obtain a second intermediate feature map. The second intermediate feature map can then be decoded to obtain a second intermediate feature map of a higher scale. The scale of the first intermediate feature map is greater than the scale of the second intermediate feature map. Specifically, the height and width of the first intermediate feature map can both be twice that of the second intermediate feature map. Similarly, stages 3, 4, and 5 can also perform the above-mentioned feature map scale reduction processing on the intermediate feature maps. In order to avoid repetition, they will not be repeated here.
[0079] For example, the input image can be represented as Among them, I can represent the input picture, It can represent a set of real numbers, that is, the pixel value of each pixel in the input image is a real number, h can represent the height of the input image, that is, the number of pixels of the input image in the vertical direction, w can represent the width of the input image, that is, the number of pixels of the input image in the horizontal direction, and 3 can represent the number of channels of the portrait image.
[0080] In some optional implementations, the intermediate feature map outputted by the i-th stage (i∈[1,5]) may be in the form shown in the following expression (1):
[0081]
[0082] Among them, f i It can represent the intermediate feature map output by stage i, It can represent a set of real numbers, that is, the pixel value of each pixel in the intermediate feature map is a real number, h can represent the height of the intermediate feature map, that is, the number of pixels in the intermediate feature map in the vertical direction, w can represent the width of the intermediate feature map, that is, the number of pixels in the intermediate feature map in the horizontal direction, c i It can represent the number of channels of the intermediate feature map output by the i-th stage. And, [c1, c2, c3, c4, c5] = [64, 65, 128, 256, 512].
[0083] It is understandable that the backbone network can also be a mobile convolutional neural network (MobileNet), an extremely efficient convolutional neural network for mobile devices (SuffleNet), and other networks, which are not specifically limited in the embodiments of the present application. When the backbone network is MobileNet, [c1, c2, c3, c4, c5] = [64, 128, 256, 1024, 2048]; when the backbone network is SuffleNet, [c1, c2, c3, c4, c5] = [24, 24, 144, 288, 576]. In this way, by continuously increasing the number of channels of the feature map, the richness and diversity of the features can be improved.
[0084] It can be understood that each convolution set in convolution set 1, convolution set 2, and convolution set 3 can be a full convolution structure, and each convolution set can include multiple convolution activation modules (CR for short), each convolution activation module includes a series of convolution layers (convolution, C) and activation units (rectified linear unit, ReLU), and each convolution activation module can perform convolution operations and ReLU nonlinear operations on the input feature map to input the nonlinear transformation feature map to the next convolution activation module.
[0085] In some optional implementations, such as Figure 5 As shown, each convolution set may include a first convolution activation module, a second convolution activation module, a third convolution activation module, a fourth convolution activation module, and a fifth convolution activation module. Exemplarily, the number of channels of the input feature map and the output feature map of each convolution activation module may be the same.
[0086] Among them, the convolution layer in the first convolution activation module can include a 1×1 convolution kernel, so the first convolution activation module can be referred to as 1×1CR; the convolution layer in the second convolution activation module can include a 3×3 convolution kernel, so the second convolution activation module can be referred to as 3×3CR; the convolution layer in the third convolution activation module can include a 1×1 convolution kernel, so the third convolution activation module can be referred to as 1×1CR; the convolution layer in the fourth convolution activation module can include a 1×1 convolution kernel, so the fourth convolution activation module can be referred to as 1×1CR; the convolution layer in the fifth convolution activation module can include a 3×3 convolution kernel, so the fifth convolution activation module can be referred to as 3×3CR.
[0087] Moreover, the nonlinear transformation feature map output by the first convolution activation module can be used as the input of the second convolution activation module, the nonlinear transformation feature map output by the second convolution activation module can be used as the input of the third convolution activation module, the nonlinear transformation feature map output by the third convolution activation module can be used as the input of the fourth convolution activation module, the nonlinear transformation feature map output by the fourth convolution activation module can be used as the input of the fifth convolution activation module, and the nonlinear transformation feature map output by the fifth convolution activation module can be used as the input of the upsampling module or the prediction head.
[0088] In some specific implementations, the feature map of the input convolution set 1 can be in the form shown in the following expression (2):
[0089] i1=f5(2)
[0090] Among them, f5 can represent the intermediate feature map with a smaller scale obtained by further feature extraction processing of the intermediate feature map output by stage 4 in stage 5, and i1 can represent the feature map of the input convolution set 1.
[0091] For other convolution sets except convolution set 1, the input feature map can include the intermediate feature map output by the corresponding stage in the backbone network and the fused feature map output by the previous sampling module. For example, the feature map of the input convolution set j (j≠1) can be i j , the feature map output by convolution set j can be o j , then the feature map of the input convolution set j can be expressed as follows:
[0092] i j =cat(u j-1 ,f 6-j ) (3)
[0093] Among them, i j It can represent the feature map of the input convolution set j, u j-1 It can represent the fusion feature map output by the previous sampling module, f 6-j Represents the intermediate feature map output by stage 6-j in the backbone network. cat can represent a concatenation operation. For example, suppose there are two feature maps C and D, whose shapes are [batch_size, height, width, channels1] and [batch_size, height, width, channels2] respectively. If the concatenation (cat) operation is performed on the channel dimension, the resulting feature map is [batch_size, height, width, channels1+channels2].
[0094] Assuming that the size of the feature map input to each convolution set is a×b, the feature map output by each convolution set can be in the form shown in the following expression (4):
[0095]
[0096] Among them, i j It can represent the feature map of the input convolution set j, o j It can represent the feature map of the convolution set j input prediction head, It can represent a set of real numbers, that is, the pixel value of each pixel in the feature map is a real number, a can represent the width of the feature map, b can represent the height of the feature map, c inj The number of channels that can represent the feature map of input convolution set j, c outj The number of channels that can represent the feature map input to the prediction head.
[0097] In some optional implementations, when the backbone network is ResNet-18, [c out1 , c out2 , c out3 ]=[c3,c4,c5].
[0098] It can be understood that each upsampling module in upsampling 1 and upsampling 2 can increase the scale of the feature map output by the convolution set to fuse it with the large-scale feature map obtained by the previous convolution set to obtain a fused feature map. Figure 6 As shown, each upsampling module in upsampling 1 and upsampling 2 may include a convolutional activation submodule and an upsampling submodule. The convolutional activation submodule may include a convolutional layer and an activation unit connected in series. The convolutional layer may include a 1×1 convolution kernel, and the activation unit may be a ReLU activation function. Therefore, the convolutional activation submodule may be referred to as a 1×1CR. The upsampling submodule may be a 2× bilinear upsampling layer. In some optional implementations, the number of channels of the input feature map and the output feature map of the convolutional layer may be the same.
[0099] It is understandable that in some optional implementations, each upsampling module in upsampling 1 and upsampling 2 may also be an upsampling module that performs a transposed convolution upsampling operation and a nearest neighbor upsampling operation on the feature map output by the convolution set, which is not specifically limited in the embodiments of the present application. Moreover, the convolution kernel in the convolution layer of each upsampling module may also be a convolution kernel of other sizes, or a stack of convolution kernels of various sizes, which is not specifically limited in the embodiments of the present application. In addition, the activation unit may also be a tanh activation function, which is not specifically limited in the embodiments of the present application.
[0100] It is understood that in some optional implementations, such as Figure 7 As shown, each prediction head in prediction head 1, prediction head 2 and prediction head 3 may include a convolutional activation submodule and a convolutional layer (conv) submodule. Among them, the convolutional activation submodule may include a convolutional layer and an activation unit connected in series, and the convolutional layer in the convolutional activation submodule may include a 3×3 convolution kernel. Therefore, the convolutional activation submodule may be referred to as 3×3CR. The convolutional layer submodule may include a 1×1 convolution kernel. Therefore, the convolutional layer submodule may be referred to as 1×1conv. In addition, the nonlinear change feature map output by the convolutional activation submodule may be used as the input of the convolutional layer submodule. In some optional implementations, the channels of the input feature map and the output feature map of 3×3CR may be the same, and the channels of the output feature map of 1×1conv may be 5.
[0101] In addition, in the prediction result output by the prediction head, that is, each pixel point in the output feature map of the convolutional layer submodule in the prediction head has a corresponding pixel point in the input image, and the corresponding pixel point can be called the point corresponding to the index coordinate.
[0102] In some optional implementations, the prediction result output by the prediction head k may be in the form shown in the following expression (5):
[0103]
[0104] Among them, p k It can be said that the prediction head k predicts the index coordinates corresponding to the face area, and the five channels of the index coordinates can be mapped to the input image (index coordinates multiplied by 2 6-k ), for example, the distance between the left boundary of the bounding box centered on the point corresponding to the index coordinate and the point corresponding to the index coordinate, the distance between the right boundary of the bounding box centered on the point corresponding to the index coordinate and the point corresponding to the index coordinate, the distance between the upper boundary of the bounding box centered on the point corresponding to the index coordinate and the point corresponding to the index coordinate, the distance between the lower boundary of the bounding box centered on the point corresponding to the index coordinate and the point corresponding to the index coordinate, and the probability that the object in the bounding box is a face.
[0105] In this way, the electronic device can crop the face area from the input image based on the maximum value of the probability that the object in the bounding box in prediction head 1, prediction head 2 and prediction head 3 is a face, and the information corresponding to the remaining four channels of the bounding box to obtain a face image.
[0106] In some specific implementations, during the training process of the face recognition model, the L2 loss function can be used to predict the up, down, left, and right information of the bounding box, and the binary cross entropy loss function can be used to predict the face probability.
[0107] In this way, a fully convolutional neural network model is used for face detection. Since at least some of the convolutional layers in the fully convolutional neural network model are grouped convolutional layers and depth-separable convolutional layers, this can reduce the computing power requirements of electronic devices. Therefore, it can be applied to electronic devices with relatively weak computing power, such as mobile phones and tablets.
[0108] Below Figure 2 and Figure 3 The training process of the reconstruction parameter prediction model mentioned is introduced in detail.
[0109] In some optional implementations, such as Figure 8 As shown in FIG, a sample input image is input into the reconstruction parameter prediction model to be trained to obtain the reconstruction parameters corresponding to the sample input image. The reconstruction parameters may include shape parameters, camera position, illumination parameters, albedo, expression parameters, and posture parameters.
[0110] An image is rendered based on the reconstruction parameters to obtain a cartoonized face model. The cartoonized face model is mapped to the image coordinate system corresponding to the sample input image to obtain a mapped image. A total loss value of the reconstruction parameter prediction model to be trained is determined based on the positional overlap of facial key points, eye key points, facial region, and facial features in the mapped image and the sample input image. The positional overlap of facial key points may include the positional overlap of 68 facial key points.
[0111] When the total loss value is greater than the preset loss threshold, the parameters of the reconstruction parameter prediction model to be trained are adjusted, and the above steps of inputting the sample input image into the reconstruction parameter prediction model to be trained, rendering the image based on the reconstruction parameters, and mapping the cartoonized face model to the image coordinate system corresponding to the sample input image are repeated until the total loss value of the reconstruction parameter prediction model to be trained is less than or equal to the preset loss threshold, and the trained reconstruction parameter prediction model is obtained.
[0112] It is understood that shape parameters may include facial contour parameters, facial feature position and shape parameters, expression parameters, etc. Facial contour parameters may include the aspect ratio of the face, the sharpness of the chin, and the contour curve of the face. Facial feature position and shape parameters may include eye position, eye size, eye aspect ratio, eye canthus tilt, nose position, nose height, nose width, nose bridge curvature, mouth position, mouth size, lip thickness, etc.
[0113] The camera position is used to determine the shooting angle of the portrait image, so that in the face reconstruction process, the three-dimensional shape of the cartoonized face model can be accurately determined in combination with the shooting angle of the portrait image.
[0114] Different camera positions cause different lighting effects on the user's face, affecting the brightness, shadows, and color of the portrait image. During the face reconstruction process, it is necessary to consider both camera position and lighting parameters to improve the realism of the reconstructed cartoon face model.
[0115] Albedo indicates how well a user's face reflects light. For example, it can represent the color and brightness characteristics of the user's skin, hair, eyes, and other parts of the face. Using albedo for facial reconstruction can improve the fidelity of the cartoon face model to the user's actual face color and gloss.
[0116] Expression parameters may include the degree of upturned corners of the mouth when smiling, the depth of forehead wrinkles when frowning, etc.
[0117] Posture parameters may include head rotation parameters, such as yaw angle, pitch angle, roll angle, etc., where the yaw angle represents the left and right rotation angle of the head, the pitch angle represents the up and down swing angle of the head, and the roll angle represents the left and right tilt angle of the head.
[0118] In some specific implementations, the reconstruction parameter prediction models to be trained may include a skeletal skinning mesh model, a facial landmark animation and editing model (FLAME), a 50-layer residual network (Res-Net50), MobileNet, ShuffleNet, a regularized network (RegNet), etc.
[0119] The skeleton skin mesh model can be expressed in the form shown in the following expression (6):
[0120]
[0121] in, J(β), θ, Ω) represents the skeleton skin mesh model, J(β) represents the bone node, such as the left corner of the eye, the right corner of the eye, etc. Represents the skin vertex, for example, a pixel corresponding to the left cheek, β can represent the shape parameter, θ can represent the posture parameter, Can represent expression parameters, Ω can represent the blend weight, n represents the number of grids.
[0122] For skinned vertices The electronic device can be calculated by the following formula (7):
[0123]
[0124] Among them, T represents the initial posture template, BS(β) represents the shape expression function, BP(θ) represents the posture expression function, represents the expression function, β can represent the shape parameter, θ can represent the posture parameter, Can represent expression parameters,
[0125] The reconstruction parameter prediction model to be trained may also include an illumination model.
[0126] In some optional implementations, the illumination model can be expressed in the form shown in the following expression (8):
[0127] B(α,l,N)=A(α)⊙∑l×H(N) Formula (8)
[0128] Among them, A() can represent the FLAME UV access albedo function, H() can represent the spherical harmonics (SH), and ⊙ can represent the dot product calculation. Can represent expression parameters, l can represent the lighting parameters,
[0129] In some optional implementations, the electronic device may render an image based on the reconstruction parameters and the illumination parameters to obtain a cartoonized face model. Furthermore, the cartoonized face model may be mapped to an image coordinate system corresponding to the input image, for example, using a projection function to map the cartoonized face model to the image coordinate system corresponding to the input image. This allows determining a total loss value of the reconstruction parameter prediction model based on at least one of the positional overlap of facial key points, the positional overlap of eye key points, the positional overlap of facial regions, and the positional overlap of facial features.
[0130] And determine the coordinates of multiple facial landmarks in the image coordinate system corresponding to the input image, and determine the coordinates of these facial landmarks in the image coordinate system corresponding to the input image from the input image.
[0131] When rendering an image based on the reconstruction parameters, a rendering function as shown in formula (9) can be used:
[0132] l r =R(M, B, c) Formula (9)
[0133] Among them, l r It can represent a cartoon face model, R() can represent a rendering function, M represents reconstruction parameters, B represents lighting parameters, and c represents camera position.
[0134] Loss = L landmark +L eye +L photo +L reg Formula (10)
[0135] Among them, Loss can represent the total loss value, L landmark It can represent the position overlap of facial key points, L eye It can express the position overlap of the key points of the eyes, Lphoto It can represent the overlap of face area, L reg It can represent the position coincidence of facial features.
[0136] Specifically, the position coincidence of facial key points can be calculated using the following formula (11):
[0137]
[0138] Among them, L landmark It can represent the position overlap of facial key points, k i It can represent the horizontal coordinates of the facial key points in the image coordinate system corresponding to the input image, project(M i ) can represent the horizontal coordinates of the facial key points when the cartoonized face model is mapped to the image coordinate system corresponding to the input image.
[0139] Specifically, the position overlap of the eye key points can be calculated using the following formula (12):
[0140]
[0141] Among them, L eye It can represent the position coincidence of eye key points.
[0142] Assume that I r If there is a mask label of the user's face (the pixel-level classification 0-1 map of the image, the part where the user's face exists is 1, otherwise it is 0), the face area overlap can be calculated using the following formula (13):
[0143] L pho =||mask(II r )||1 Formula (13)
[0144] Among them, L pho It can indicate the degree of overlap of face areas.
[0145] In some optional implementations, facial features may include shape parameters, expression parameters, and albedo, so the positional overlap of facial features can be determined based on the regularization loss of each facial feature. Specifically, the positional overlap of facial features can be calculated using the following formula (14):
[0146]
[0147] Among them, L reg It can represent the overlap of face area, β can represent the shape parameter, It can represent the expression parameter, and α can represent the albedo.
[0148] After completing the training of the reconstruction parameter prediction model, when performing face reconstruction based on an input image or a face image, only the shape parameters can be output as face parameters, so as to perform image rendering based on the face parameters and obtain a cartoon face model.
[0149] The effects of the cartoonized digital human generation method mentioned in the embodiments of this application are described below in conjunction with some cartoonized digital human generation methods.
[0150] In some cartoon-like digital human generation methods, an electronic device can obtain a target object image and use generative facial prior-generative adversarial network (GFP-GAN) technology to perform three-dimensional facial reconstruction of the target object in the target object image to generate a three-dimensional face model. The electronic device can then use a voice-driven method to perform facial driving on the three-dimensional face model to obtain the final three-dimensional face model. However, this cartoon-like digital human generation method only includes three-dimensional face reconstruction and cannot customize the cartoon-like digital human corresponding to the user's face in any input image.
[0151] In other methods for generating cartoon-like digital humans, electronic devices can obtain a user's textual description of a 3D digital human and then use natural language processing (NLP) technology to parse the description to obtain the target's key attribute information. The electronic device can then determine the values of control parameters in the 3D digital human's parametric model based on the target's key attribute information and input these control parameter values into a stable diffusion model to generate the cartoon-like digital human. However, this method uses a stable diffusion model, making it difficult to deploy on mobile devices with relatively low computing power, such as smartphones and tablets. Furthermore, the cartoon-like digital humans generated using this method are not drivable.
[0152] Therefore, compared with the above-mentioned cartoonized digital human generation method, the cartoonized digital human generation method mentioned in the embodiment of the present application can reduce the computing power requirements of electronic devices, and therefore can be applied to electronic devices with relatively weak computing power, such as mobile phones, tablets, etc.
[0153] It is understood that the method for generating a cartoonized digital human provided in the embodiment of the present application can be applied to electronic devices. The following is an exemplary introduction to the hardware structure of an electronic device to which the method for generating a cartoonized digital human provided in the embodiment of the present application is applicable.
[0154] like Figure 9As shown, the electronic device 900 may include a processor 910, an external memory interface 920, an internal memory 921, a universal serial bus (USB) interface 930, a charging management module 940, a power management module 941, a battery 942, an antenna, a wireless communication module 950, an audio module 960, a speaker 960A, a receiver 960B, a microphone 960C, a headphone jack 960D, a camera 970, a display screen 980, etc.
[0155] It should be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 900. In other embodiments of the present application, the electronic device 900 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0156] The processor 910 may include one or more processing units. For example, the processor 910 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.
[0157] The controller can generate an operation control signal based on the instruction operation code and the timing signal to complete the control of fetching and executing instructions. The processor 910 can control the fetching and executing of instructions through the controller to implement the method for generating a cartoonized digital human provided by the embodiment of the present application. For example, the processor 910 can control the fetching and executing of instructions through the controller to implement the above-mentioned method for generating a cartoonized digital human. Figure 2 or Figure 3 The corresponding steps in the process shown are implemented.
[0158] Processor 910 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 910 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 910. If processor 910 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 910 latency, and thus improves system efficiency.
[0159] The wireless communication function of the electronic device can be implemented through an antenna, a wireless communication module 950, a modem processor, and a baseband processor.
[0160] Antennas are used to transmit and receive electromagnetic wave signals. Each antenna in an electronic device can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antennas can be reused as diversity antennas for wireless local area networks. In some other embodiments, antennas can be used in conjunction with tuning switches.
[0161] The wireless communication module 950 can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 950 can be one or more devices that integrate at least one communication processing module. The wireless communication module 950 receives electromagnetic waves via an antenna, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 910. The wireless communication module 950 can also receive signals to be sent from the processor 910, frequency modulate them, amplify them, and convert them into electromagnetic waves for radiation through the antenna.
[0162] The electronic device implements display functions through a GPU, display screen 980, and an application processor. The GPU is a microprocessor for image processing that connects the display screen 980 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 910 may include one or more GPUs that execute program instructions to generate or modify display information.
[0163] Display screen 980 is used to display images, videos, and the like. Display screen 980 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Mini-LED, a MicroLED, a Micro-OLED, or a quantum dot light-emitting diode (QLED). In some embodiments, the electronic device can include one or N display screens 980, where N is a positive integer greater than one.
[0164] In some cases, the embodiments disclosed herein may be implemented in hardware, firmware, software, or any combination thereof.
[0165] The embodiments disclosed in this application can also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which can be read and executed by one or more processors. For example, instructions can be distributed over a network or by other computer-readable media. Therefore, a machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to a floppy disk, an optical disk, an optical disk, a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic card or an optical card, a flash memory, or a tangible machine-readable memory for transmitting information (e.g., a carrier wave, an infrared signal digital signal, etc.) using the Internet in an electrical, optical, acoustic or other form of propagation signal. Therefore, a machine-readable medium includes any type of machine-readable medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).
[0166] Embodiments of the present application may be implemented as a computer program or program code executed on a programmable system comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.
[0167] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit, or a microprocessor.
[0168] Program code can be implemented with a high-level programming language or an object-oriented programming language to communicate with the processing system. Where necessary, program code can also be implemented in assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any particular programming language. In either case, the language can be a compiled language or an interpreted language.
[0169] The above describes the hardware structure that an electronic device may have. It is understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device may include more or fewer components than shown in the figure, or combine certain components, separate certain components, or arrange the components differently. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.
[0170] In the accompanying drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or order may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. In addition, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and in some embodiments, such features may not be included or may be combined with other features.
[0171] It should be noted that in the examples and description of this patent, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "including a" does not exclude the presence of other identical elements in the process, method, article or device that includes the element.
[0172] While the present application has been shown and described with reference to certain embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the present application.
Claims
1. A method for generating a cartoon digital human, applied to electronic equipment, characterized in that: include: Acquire a first image, where the first image at least includes a user's face; generating a cartoonized face model based on the user's face in the first image; The cartoonized face model is combined with the first cartoonized torso model in the torso material library to obtain a cartoonized digital human.
2. The method according to claim 1, characterized in that The step of generating a cartoonized face model based on the user's face in the first image includes: Inputting the first image into a face recognition model, and identifying an area corresponding to the user's face from the first image; determining facial parameters of the user's face based on an area corresponding to the user's face in the first image; Image rendering is performed based on the facial parameters to obtain the cartoonized facial model.
3. The method according to claim 2, characterized in that The face recognition model includes a backbone network, a first convolution set, a second convolution set, a third convolution set, a first upsampling module, a second upsampling module, a first prediction head, a second prediction head and a third prediction head; The backbone network includes a first sub-network, a second sub-network, a third sub-network, a fourth sub-network and a fifth sub-network, wherein the feature map output by the first sub-network is the input of the second sub-network, the feature map output by the second sub-network is the input of the third sub-network, the feature map output by the third sub-network is the feature map output by the fourth sub-network, and the feature map output by the fourth sub-network is the input of the fifth sub-network; The feature map output by the fifth subnetwork is used as input to the first convolution set, and the feature map output by the first convolution set is used as input to the first prediction head. The prediction head is configured to determine a first prediction result based on the feature map output by the first convolution set, where the first prediction result includes a first region in the first image and a probability that the first region corresponds to the user's face in the first image. The feature map output by the first convolution set is the input of the first upsampling module, and the feature map output by the first upsampling module is the input of the second convolution set; The feature map output by the fourth subnetwork is used as input to the second convolution set, and the feature map output by the second convolution set is used as input to the second prediction head. The second prediction head is used to determine a second prediction result based on the feature map output by the second convolution set, where the second prediction result includes a second region in the first image and a probability that the second region is the region corresponding to the user's face in the first image. The feature map output by the second convolution set is the input of the second upsampling module, and the feature map output by the second upsampling module is the input of the third convolution set; The feature map output by the third subnetwork is the input of the third convolution set, and the feature map output by the third convolution set is the input of the third prediction head. The third prediction head is used to determine a third prediction result based on the feature map output by the third convolution set. The third prediction result includes the third area in the first image and the probability that the third area is the area corresponding to the user's face in the first image.
4. The method according to claim 3, characterized in that Among the first area, the second area, and the third area, the area with the highest probability of being the area corresponding to the user's face in the first image is used as the area corresponding to the user's face.
5. The method according to claim 2, characterized in that The determining of facial parameters of the user's face based on the area corresponding to the user's face in the first image includes: The first image is input into a reconstruction parameter prediction model to determine facial parameters of the user's face.
6. The method according to claim 5, characterized in that The reconstruction parameter prediction model is trained in the following way: Obtain at least one sample image, input the sample image into a reconstruction parameter prediction model to be trained, and predict reconstruction parameters corresponding to the user's face in each sample image; Performing image rendering based on the reconstruction parameters corresponding to the user's face in each of the sample images to obtain a first cartoonized face model corresponding to each of the sample images; Mapping each of the first cartoonized face models to the coordinate system of the corresponding sample image to obtain a mapped image corresponding to each of the sample images; Based on the degree of overlap between each sample image and the user's face in the corresponding mapping image, the model parameters of the reconstruction parameter prediction model to be trained are adjusted until the degree of overlap meets a termination condition.
7. The method according to claim 6, characterized in that The overlap between the user's face in each sample image and the corresponding mapping image includes one or more of the overlap of facial key points, the overlap of eye key points, the overlap of facial regions, and the overlap of facial features.
8. An electronic device, characterized in that: include: The memory is used to store instructions executed by one or more processors of the electronic device, and the processor is one of the one or more processors of the electronic device, and is used to execute the method for generating a cartoonized digital human according to any one of claims 1 to 7.
9. A readable storage medium, characterized in that: The readable storage medium stores instructions, and when the instructions are executed on an electronic device, the electronic device executes the method for generating a cartoonized digital human according to any one of claims 1 to 7.
10. A computer program product, characterized in that The computer program product includes computer instructions, and when executed by an electronic device, the electronic device executes the computer program code of the method for generating a cartoonized digital human according to any one of claims 1 to 7.
Citation Information
Patent Citations
End-to-end face detection and recognition method
CN110399826A
Face cartoonalization method and device and computer storage medium
CN112907708A
Face detection method and device for learning noise region information
CN113128479A
Face image cartoonalization processing method and device, computer equipment and storage medium
CN114820907A
Training data generation method and device, electronic equipment and storage medium
CN116109646A