Model processing method, portrait generation method and related products

By acquiring sample datasets for feature extraction and model parameter adjustment, the problems of harsh edges and low clarity in portrait segmentation were solved, resulting in clear and natural portraits.

CN121982122APending Publication Date: 2026-05-05MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MASHANG CONSUMER FINANCE CO LTD
Filing Date
2024-10-30
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies for human portrait segmentation suffer from harsh edges and low clarity, resulting in portraits that are not natural or clear.

Method used

By acquiring a sample dataset of sample objects, including sample images, sample reference images, and sample portraits, feature extraction is performed to generate accurate hairstyles and facial features. The model parameters are then adjusted and trained to obtain a second model, which generates clear and natural portraits.

Benefits of technology

It achieves clear and natural semantic structures for hairstyles and faces in portraits, improving the accuracy and speed of portrait segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121982122A_ABST
    Figure CN121982122A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model processing method, a portrait generation method and a related product. The model processing method comprises the steps of obtaining a sample data set of a sample object; the sample data set comprises a sample image, a sample reference image and a sample portrait; inputting the sample data set into the model for image processing to obtain a predicted portrait of the sample object; the image processing comprises the following steps: performing feature extraction on a sample image and a sample reference image to obtain a face feature and a hair style feature of a sample object; generating a reference portrait of the sample object according to the face features and the hairstyle features; and determining sample portrait features of the sample object according to the reference portrait, the face features and the hair style features, and generating a predicted portrait of the sample object according to the sample portrait features. And adjusting model parameters according to the sample image, the predicted portrait, the sample portrait and the reference portrait. According to the invention, the model with excellent performance and portrait generation capability can be trained, and the trained model can generate a clear and natural portrait.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and image processing technology, and in particular to a model processing method, a portrait generation method, and related products. Background Technology

[0002] In image processing, portrait segmentation remains a popular topic. Distinguishing between people and backgrounds at the pixel level is a classic and widely applied task in portrait segmentation. Generally, portrait segmentation tasks can be divided into two categories: one is segmentation for full-body and half-body portraits, referred to as general portrait segmentation; the other is segmentation for half-body portraits, referred to as portrait segmentation. Portrait segmentation technology is widely deployed on the internet, mobile phones, and edge devices. Therefore, portrait segmentation needs to achieve both high accuracy and extremely fast inference speed. Balancing accuracy and speed with complex and rapidly changing portrait edges remains a highly challenging task in portrait segmentation. Summary of the Invention

[0003] The purpose of this application is to provide a model processing method, a portrait generation method, and related products to solve the problems of harsh edges and low clarity in human portrait segmentation in the prior art, which result in unnatural and unclear portraits.

[0004] To solve the above-mentioned technical problems, the embodiments of this application are implemented as follows: On the one hand, embodiments of this application provide a model processing method, including: Obtain a sample dataset of the sample objects; the sample dataset includes: sample images, sample reference images, and sample portraits; the sample reference images include background region images, face region images, and hairstyle region images; the sample images are obtained by removing the hairstyle region images from the sample reference images. The sample dataset is input into a first model for image processing to obtain a predicted portrait of the sample object. The image processing includes: extracting features from the sample image and the sample reference image to obtain a first facial feature and a first hairstyle feature of the sample object; generating a reference portrait of the sample object based on the first facial feature and the first hairstyle feature; determining the sample portrait features of the sample object based on the reference portrait, the first facial feature, and the first hairstyle feature, and generating the predicted portrait based on the sample portrait features. Based on the sample image, the predicted portrait, the sample portrait, and the reference portrait, the model parameters of the first model are adjusted to obtain the second model.

[0005] On the other hand, embodiments of this application provide a portrait generation method, including: Obtain an image dataset of a first object; the image dataset includes a first image and a first reference image; the first reference image includes a background region image, a face region image, and a hairstyle region image; the first image is obtained by removing the hairstyle region image from the first reference image. The image dataset is input into a second model for image processing to obtain a portrait of the first object; the image processing includes: extracting features from the first image and the first reference image to obtain facial features and hairstyle features of the first object; determining portrait features of the first object based on the facial features and hairstyle features; and generating a portrait of the first object based on the portrait features.

[0006] Furthermore, embodiments of this application provide a model processing apparatus, including: The first acquisition module is used to acquire a sample dataset of the sample object; the sample dataset includes: a sample image, a sample reference image, and a sample portrait; the sample reference image includes a background region image, a face region image, and a hairstyle region image; the sample image is obtained by removing the hairstyle region image from the sample reference image. A first processing module is used to input the sample dataset into a first model for image processing to obtain a predicted portrait of the sample object; the image processing includes: extracting features from the sample image and the sample reference image to obtain a first facial feature and a first hairstyle feature of the sample object; generating a reference portrait of the sample object based on the first facial feature and the first hairstyle feature; determining the sample portrait features of the sample object based on the reference portrait, the first facial feature, and the first hairstyle feature, and generating the predicted portrait based on the sample portrait features; The model training module is used to adjust the model parameters of the first model based on the sample image, the predicted portrait, the sample portrait, and the reference portrait to obtain the second model.

[0007] Furthermore, embodiments of this application provide a portrait generation apparatus, comprising: The second acquisition module is used to acquire an image dataset of the first object; the image dataset includes a first image and a first reference image; the first reference image includes a background region image, a face region image, and a hairstyle region image; the first image is obtained by removing the hairstyle region image from the first reference image. The second processing module is used to input the image dataset into the second model for image processing to obtain the portrait of the first object; the image processing includes: extracting features from the first image and the first reference image to obtain the facial features and hairstyle features of the first object; determining the portrait features of the first object based on the facial features and the hairstyle features; and generating the portrait of the first object based on the portrait features.

[0008] In another aspect, embodiments of this application provide an electronic device, including a processor and a memory electrically connected to the processor. The memory stores a computer program, and the processor is used to call and execute the computer program from the memory to implement the above-described model processing method, or the processor is used to call and execute the computer program from the memory to implement the above-described portrait generation method.

[0009] In another aspect, embodiments of this application provide a computer-readable storage medium for storing a computer program that can be executed by a processor to implement the above-described model processing method, or the computer program can be executed by a processor to implement the above-described portrait generation method.

[0010] In another aspect, embodiments of this application provide a computer program product, including a computer program, which is executed by a processor to implement the above-described model processing method, or the computer program is executed by a processor to implement the above-described portrait generation method.

[0011] The technical solution of this application involves acquiring a sample dataset of the sample object. This dataset includes sample images, sample reference images, and sample portraits. The sample reference images include background region images, face region images, and hairstyle region images. The sample dataset is input into a first model, and feature extraction is performed on the sample images and sample reference images to obtain the face features and hairstyle features of the sample object. It is evident that the sample data used in training the model includes not only the sample reference images (i.e., the complete portrait) but also the sample images after removing the hairstyle region images (i.e., only the background region images and face region images). This allows the model to extract accurate and rich hairstyle features, i.e., the semantic structure of the hairstyle, rather than simple hairstyle contour recognition, providing strong semantic structural support for model training. Furthermore, a reference portrait of the sample object is generated based on the face features and hairstyle features. The portrait features of the sample object are determined based on the reference portrait, face features, and hairstyle features, and a predicted portrait of the sample object is generated based on the portrait features. The model parameters of the first model are adjusted based on the sample images, predicted portraits, sample portraits, and reference portraits to obtain the second model. Since the sample reference images provide accurate and rich hairstyle features for the model training process, and the sample images provide accurate and rich facial features for the model training process, as the model is trained iteratively, the generated predicted portraits can have both clear and natural hairstyle semantic structure and facial semantic structure. This further enables the trained second model to generate clear and natural portraits. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in one or more embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in one or more embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a schematic diagram illustrating the technical concept of a model processing method and a portrait generation method according to an embodiment of this application; Figure 2 This is a schematic structural diagram of a Unet model according to an embodiment of this application; Figure 3 This is a schematic flowchart of a model processing method according to an embodiment of this application; Figure 4 This is a schematic flowchart of a model processing method according to another embodiment of this application; Figure 5 This is a schematic diagram illustrating a training method for a portrait generation model according to an embodiment of this application. Figure 6 This is a schematic diagram illustrating a training method for a portrait generation model according to another embodiment of this application; Figure 7 This is a schematic flowchart of a portrait generation method according to an embodiment of this application; Figure 8 This is a schematic diagram illustrating a portrait generation method according to an embodiment of this application. Figure 9 This is a schematic diagram illustrating a portrait generation method according to another embodiment of this application; Figure 10 This is a schematic block diagram of a model processing apparatus according to an embodiment of this application; Figure 11 This is a schematic block diagram of a portrait generation device according to an embodiment of this application; Figure 12 This is a schematic block diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0014] This application provides a model processing method, a portrait generation method, and related products to solve the problems of harsh edges and low clarity in human portrait segmentation in the prior art, resulting in unnatural and unclear portraits.

[0015] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0016] In portrait segmentation, image matting techniques are commonly used to extract portraits from human images. The biggest drawback of matting is its blurry and uneven edges, resulting in a lack of naturalness in the segmented portrait. To address this, related technologies introduce background images as references and utilize deep neural networks to train portrait segmentation models. Specifically, during model training, both the segmentation results and the background image are used as reference information. A CS (Context Switching block) module is introduced to effectively select useful information from the image. After decoding, a more accurate matting result is obtained. However, while this approach improves the quality of the segmented portrait to some extent, such as enhancing the smoothness of the portrait outline, it is still essentially a matting method. Due to the inherent limitations of matting, the segmented portrait still retains some background pixel information, resulting in unnatural and uneven edges and low clarity.

[0017] To address the technical issues of unnatural and low-resolution portrait segmentation, this application provides a model processing method. This method involves acquiring a sample dataset of the sample object, including sample images, sample reference images, and sample portraits. The sample reference images include background region images, face region images, and hairstyle region images. The sample dataset is input into a first model, and feature extraction is performed on the sample images and sample reference images to obtain the facial and hairstyle features of the sample object. It is evident that the sample data used in training the model includes not only the sample reference images (i.e., the complete portrait) but also the sample images after removing the hairstyle region images (i.e., only including the background and face region images). This allows the model to extract accurate and rich hairstyle features, i.e., the semantic structure of the hairstyle, rather than simple hairstyle contour recognition, providing strong semantic structural support for model training. Furthermore, a reference portrait of the sample object is generated based on the facial and hairstyle features. The portrait features of the sample object are determined based on the reference portrait, facial features, and hairstyle features, and a predicted portrait of the sample object is generated based on these portrait features. The model parameters of the first model are adjusted based on the sample images, predicted portraits, sample portraits, and reference portraits to obtain a second model. Since the sample reference images provide accurate and rich hairstyle features for the model training process, and the sample images provide accurate and rich facial features for the model training process, as the model is trained iteratively, the generated predicted portraits can have both clear and natural hairstyle semantic structure and facial semantic structure. This further enables the trained second model to generate clear and natural portraits.

[0018] The model processing method and portrait generation method provided in this application can both be executed by an electronic device or by software installed in an electronic device. Specifically, the electronic device can be a terminal device or a server device. The terminal device can include smartphones, laptops, smart wearable devices, vehicle terminals, etc., and the server device can include an independent physical server, a server cluster consisting of multiple servers, or a cloud server capable of cloud computing.

[0019] Figure 1 This is a schematic diagram illustrating the technical concept of the model processing method and portrait generation method provided in this application. Figure 1 As shown, during model training or portrait generation, the human body image and hairstyle mask image are input into the pre-trained model. The pre-trained model generates a portrait based on the human body image and hairstyle mask image. The human body image refers to an image including the hairstyle region, face region, and background region. The hairstyle mask image is the image obtained by masking the hairstyle region image from the human body image. In other words, the pre-trained model can obtain all human body features (including hairstyle features, facial features, and background features) from the human body image, and facial and background features from the hairstyle mask image. By using all human body features as reference features, accurate facial and hairstyle features are determined, and then a portrait is generated based on these features. In the model training method, the pre-trained model can be a Unet (a U-shaped network structure) model. In the portrait generation method, the pre-trained model can be a portrait generation model trained based on the Unet model. As can be seen, the model processing method and portrait generation method provided in this application both generate portraits based on accurate facial features and hairstyle features, that is, the portraits are obtained based on the generation method, rather than using image matting technology, so that the generated portraits have both clear and natural hairstyle semantic structure and facial semantic structure.

[0020] Before detailing the model processing methods, we will first introduce the pre-trained model used in this application, namely the Unet model. Figure 2 This is a schematic structural diagram of a Unet model according to an embodiment of this application, such as... Figure 2As shown, the Unet model is a U-shaped network structure. Its backbone consists of a feature extraction network (encoder) on the left and a feature fusion network (decoder) on the right. The feature extraction network, acting as a downsampling layer, extracts features, encoding the high-resolution image into abstract semantic features. It consists of one feature extraction layer and four convolutional layers. The feature fusion network, acting as an upsampling layer, decodes the abstract semantic features encoded by the feature extraction network to recover the high-resolution image. It consists of four convolutional layers and one convolutional output layer. The four convolutional layers in the feature fusion network correspond to the four convolutional layers in the feature extraction network. Each convolutional layer in the upsampling layer fuses the output features of the corresponding convolutional layer in the downsampling layer, making the image information after upsampling richer and more complete. The convolutional output layer is used to reduce the dimensionality of the image information after upsampling, reducing the number of channels to a specific number to obtain the target image. Of course, the number of convolutional layers in the feature extraction network and the feature fusion network is not limited to 4. As long as the number of convolutional layers in the feature extraction network and the feature fusion network is the same, this embodiment is just an example with 4 convolutional layers.

[0021] Figure 3 This is a schematic flowchart of a model processing method according to an embodiment of this application, such as... Figure 3 As shown, the method includes: S302, Obtain the sample dataset of the sample object. The sample dataset includes: sample image, sample reference image and sample portrait.

[0022] The sample reference image includes a background region image, a face region image, and a hairstyle region image. The sample image is obtained by removing the hairstyle region image from the sample reference image, and includes both a face region image and a background region image. The sample portrait is an image without a background region image, and includes both a face region image and a hairstyle region image.

[0023] For example, consider user A. A photo of user A includes a half-body shot of user A and a background image. The half-body shot includes images of user A's face and hairstyle. This photo of user A serves as a sample reference image. Removing the hairstyle image from user A's photo (e.g., by masking or using an electronic eraser) results in a photo lacking the hairstyle image, which is the sample image. Similarly, removing the background image (or extracting the foreground image) from user A's photo results in a photo lacking the background image, which is the sample portrait. The sample portrait includes images of user A's face and hairstyle.

[0024] When acquiring the sample dataset, multiple sample reference images of the sample object in different poses can be obtained first. Then, a foreground extraction algorithm is used to extract the foreground image from the sample reference images to obtain the sample portrait. The foreground image includes the face region image and the hairstyle region image. The sample portrait is the foreground image in the sample reference images, and any existing foreground extraction algorithm can be used to extract the foreground image, such as using masking techniques to cut out the sample reference images to obtain the foreground image. For each sample reference image, the hairstyle region image in the sample reference image is identified, and a masking process is applied to the hairstyle region image in the sample reference image to obtain the sample image. Alternatively, the hairstyle region image in the sample reference image can be directly erased to obtain the sample face image.

[0025] For example, for the same sample object, obtain 5 sample reference images of the sample object in different poses, then make a copy of each sample reference image, and perform masking on the hairstyle area image in each copied sample reference image to obtain sample images corresponding to the 5 sample reference images respectively.

[0026] In the sample dataset, multiple sample reference images with different poses share the same ID (identification) feature. That is, the same ID corresponds to multiple sample reference images and multiple sample images. During model training, using multiple sample reference images and multiple sample images corresponding to the same ID as input data can increase the diversity of input data, thereby making the information learned by the model richer and more diverse, which is beneficial to speeding up the model training process.

[0027] S304. Input the sample dataset into the first model for image processing to obtain the predicted portrait of the sample object.

[0028] The image processing process includes, for example: Figure 4 Steps S3042-S3046 shown: S3042, Perform feature extraction on the sample image and the sample reference image to obtain the first face region feature and the first hairstyle region feature of the sample object.

[0029] Optionally, the first model includes a first feature extraction module and a second feature extraction module. The first model is the portrait generation model to be trained. After inputting the sample dataset into the portrait generation model to be trained, feature extraction is first performed through these two feature extraction modules. Specifically, the first feature extraction module extracts features from the sample image to obtain the first facial features of the sample image, and the second feature extraction module extracts features from the sample reference image to obtain the reference image features of the sample object. These reference image features are then input into the first feature extraction module. Furthermore, the first feature extraction module processes the first facial features and the reference image features to obtain the first hairstyle features. In other words, the first model (i.e., the portrait generation model to be trained) includes two structurally identical feature extraction modules, which extract features from the sample image and the sample reference image, respectively.

[0030] The reference image features include facial features and hairstyle features. By comparing the reference image features extracted by the second feature extraction module with the facial features extracted by the first feature extraction module, the hairstyle features can be accurately determined. It should be noted that the first feature extraction module can also extract background features from the sample image, and the second feature extraction module can also extract background features from the sample reference image. Since the sample reference image and the sample image include the same background image, and this embodiment uses two feature extraction modules to extract accurate facial and hairstyle features without needing to identify the dividing lines between the facial region, hairstyle region, and background region, the extraction of background features is not specifically described during the feature extraction process.

[0031] S3044, Generate a reference portrait of the sample object based on the first facial features and the first hairstyle features.

[0032] In generating a reference portrait of a sample object, the first facial features and the first hairstyle features can be dimensionality reduced to obtain the second facial features and the second hairstyle features of the sample object, i.e., latent features. Then, a reference portrait of the sample user can be generated based on the second facial features and the second hairstyle features.

[0033] In one embodiment, a generative adversarial network (GAN) can be introduced into the first model to enable the first model to learn the image generation capabilities of the GAN. When generating a reference portrait of a sample object using the GAN, the first facial features and the first hairstyle features can be dimensionality-reduced to obtain the latent features of the sample object. The latent features include the second facial features and the second hairstyle features. Then, the latent features of the sample object are input into the GAN, and the GAN generates a reference portrait of the sample object based on the latent features.

[0034] Latent features can be one-dimensional features. For example, features extracted by a feature extraction network (including facial region features and hairstyle features) can be compressed into 1*512 dimensional features; these 1*512 dimensional features are the latent features. The advantage of generative adversarial networks (GANs) lies in generating detailed and realistic human images, such as portraits, based on low-dimensional features. Of course, the features extracted by the feature extraction network can be directly input into the GAN for portrait generation without dimensionality reduction. However, compared to directly inputting the extracted features, dimensionality reduction significantly reduces the computational cost of the GAN without noticeably affecting the quality of the portrait generated.

[0035] Generative adversarial networks (GANs) can be used to generate StyleGan generators, which can generate new images that simulate real images and have powerful image generation capabilities. After determining facial and hairstyle features, StyleGan generators can generate detailed and realistic portraits.

[0036] In one embodiment, the first model further includes a discriminator for constraining the generative adversarial network (GAN). The discriminator's role is to determine the authenticity of the predicted portrait, which can be understood as the degree of authenticity of the image information contained in the predicted portrait. Since the predicted portrait is generated based on portrait features, and these portrait features are obtained from reference features in the reference portrait generated by the GAN, the authenticity of the predicted portrait reflects the quality of the reference portrait generated by the GAN. The discriminator's assessment of the predicted portrait is used as the training basis for the portrait generation model. For example, if the assessment result is false, the portrait generation model continues to be iteratively trained; if the assessment result is true, the portrait generation model stops for the next iteration. As the model iterates and trains, the GAN's ability to generate portraits improves, thereby enabling the portrait generation model to learn a stronger portrait generation capability.

[0037] S3046. Based on the reference portrait, the first face feature, and the first hairstyle feature, determine the sample portrait features of the sample object, and generate a predicted portrait of the sample object based on the sample portrait features.

[0038] Optionally, when determining the sample portrait features of a sample object based on a reference portrait, first facial features, and first hairstyle features, features can first be extracted from the reference portrait to obtain the first portrait features corresponding to the reference portrait. Then, feature fusion processing can be performed on the second facial features, second hairstyle features, and first facial features to obtain the second portrait features of the sample object. Afterward, feature fusion processing can be performed on the first portrait features and second portrait features to obtain the sample portrait features.

[0039] The latent features (including the second face features and the second hairstyle features) are obtained by dimensionality reduction of the first face features and the first hairstyle features. By fusing the latent features and the first face features extracted by the first feature extraction module, the fused second portrait features can have both low-resolution information during downsampling and high-resolution information during upsampling, thus making the final portrait features richer and more complete.

[0040] S306. Based on the sample image, the predicted portrait, the sample portrait, and the reference portrait, the model parameters of the first model are adjusted to obtain the second model.

[0041] In this step, the model loss value of the first model is determined based on the sample image, the predicted portrait, the sample portrait, and the reference portrait, and the first model is trained based on the model loss value. The second model is the trained portrait generation model, and the calculation method of the model loss value will be described in detail in the following embodiments.

[0042] The technical solution of this application involves acquiring a sample dataset of the sample object. This dataset includes sample images, sample reference images, and sample portraits. The sample reference images include background region images, face region images, and hairstyle region images. The sample dataset is input into a first model, and feature extraction is performed on the sample images and sample reference images to obtain the face features and hairstyle features of the sample object. It is evident that the sample data used in training the model includes not only the sample reference images (i.e., the complete portrait) but also the sample images after removing the hairstyle region images (i.e., only the background region images and face region images). This allows the model to extract accurate and rich hairstyle features, i.e., the semantic structure of the hairstyle, rather than simple hairstyle contour recognition, providing strong semantic structural support for model training. Furthermore, a reference portrait of the sample object is generated based on the face features and hairstyle features. The portrait features of the sample object are determined based on the reference portrait, face features, and hairstyle features, and a predicted portrait of the sample object is generated based on the portrait features. The model parameters of the first model are adjusted based on the sample images, predicted portraits, sample portraits, and reference portraits to obtain the second model. Since the sample reference images provide accurate and rich hairstyle features for the model training process, and the sample images provide accurate and rich facial features for the model training process, as the model is trained iteratively, the generated predicted portraits can have both clear and natural hairstyle semantic structure and facial semantic structure. This further enables the trained second model to generate clear and natural portraits.

[0043] Figure 5 This is a schematic diagram illustrating a training method for a portrait generation model according to an embodiment of this application, such as... Figure 5As shown, the portrait generation model to be trained includes: a first feature extraction module, a second feature extraction module, a third feature extraction module, a first feature fusion module, a second feature fusion module, a generative adversarial network (GAN), and a discriminator for constraining the GAN. During training, a sample dataset of the sample object is obtained, including sample images, sample reference images, and sample portraits. The sample portraits serve as label information for the sample images. The sample images are input into the first feature extraction module, and the sample reference images are simultaneously input into the second feature extraction module. The first feature extraction module extracts features from the sample images to obtain the first facial features of the sample images. The second feature extraction module extracts features from the sample reference images to obtain the reference image features of the sample object. The second feature extraction module inputs the reference image features into the first feature extraction module, and the first feature extraction module determines the first hairstyle feature based on the reference image features and the first facial features.

[0044] Subsequently, the first feature extraction module inputs the first face feature and the first hairstyle feature into the first feature fusion module, and simultaneously inputs the first face feature and the first hairstyle feature into the generative adversarial network (GAN). The first feature fusion module fuses the first face feature and the first hairstyle feature input from the first feature extraction module. The GAN generates a reference portrait based on the first face feature and the first hairstyle feature input from the first feature extraction module. The reference portrait is input into the third feature extraction module, which extracts features from the reference portrait to obtain the corresponding first portrait feature. The first portrait feature is input into the second feature fusion module, which simultaneously inputs the output second portrait feature (i.e., the fused feature of the first face feature and the first hairstyle feature) into the second feature fusion module. The second feature fusion module further fuses the first portrait feature and the second portrait feature to obtain the sample portrait feature of the sample object.

[0045] Next, the sample portrait features are input into the first feature fusion module, which generates a predicted portrait based on these features. The predicted portrait is then input into a discriminator, which assesses the realism of the image information in the predicted portrait against the sample portrait and outputs the result. The result is categorized as either true or false, represented by the numbers "1" and "0," respectively. A "1" output indicates a true prediction with high realism, while a "0" output indicates a false prediction with low realism.

[0046] After the discriminator outputs the discrimination result for the predicted portrait, the model loss value of the first model (i.e. the portrait generation model to be trained) is determined based on the sample image, the predicted portrait, the sample portrait, and the reference portrait, and the model parameters of the first model are adjusted based on the model loss value.

[0047] The following section details how to determine the model loss value for the first model.

[0048] In one embodiment, the model loss value of the first model is determined based on the sample image, the predicted portrait, the sample portrait, and the reference portrait, and then the model parameters of the first model are adjusted based on the model loss value. The method for determining the model loss value includes the following steps A1-A3: Step A1: Based on the predicted portrait and the sample portrait, determine the first loss value, the second loss value, and the third loss value of the first model. The first loss value is used to characterize the first degree of difference between the predicted portrait and the sample portrait, the second loss value is used to characterize the image realism of the predicted portrait relative to the sample portrait, and the third loss value is used to characterize the second degree of difference in facial features between the predicted portrait and the sample portrait.

[0049] The first loss value can be characterized by the mean squared error between the predicted portrait and the sample portrait to represent the degree of difference. That is, the first loss value loss1 can be calculated as follows: Loss1 = mse(net(x), label) Where x represents the sample image, net(x) represents the predicted portrait output by the first model, Label represents the label information, i.e., the sample portrait in the sample dataset, and mse represents the mean squared error between net(x) and label.

[0050] When calculating the second loss value, the realism of the image information in the predicted portrait can be determined based on the realism of the image information in the sample portrait, and then the second loss value can be determined based on the realism of the image information in the predicted portrait. Optionally, the second loss value of the first function can be determined by a discriminator. The predicted portrait and the sample portrait are input into the discriminator, and the realism of the image information in the predicted portrait is discriminated based on the realism of the image information in the sample portrait, thereby determining the second loss value based on the discriminator result of the realism of the image information in the predicted portrait. If the numbers "1" and "0" are used to represent the realism of the image information, that is, if the discriminator outputs "1", it means that the discriminator result of the predicted portrait is true, and the realism of the image information it contains is high. If the discriminator outputs "0", it means that the discriminator result of the predicted portrait is false, and the realism of the image information it contains is low. Since the sample portrait is the label information used to train the first model, the realism of the image information in the sample portrait can be considered as 1, or a value close to 1. If the sample portrait is input into the discriminator, the output of the discriminator is 1 or a value close to 1. The calculation method of the second loss value Loss2 can be expressed as the following formula: Loss2=

[0051]

[0052]

[0053] in, Indicates a sample portrait. This represents the predicted portrait output by the first model. This represents the output of the discriminator after the sample portrait is input into it. This indicates that the output of the discriminator will be predicted after the portrait is input into the discriminator.

[0054] When calculating the third loss value, face recognition can be performed on the predicted portrait and the sample portrait to obtain the third face feature of the predicted portrait and the fourth face feature of the sample portrait. Then, the third loss value is determined based on the difference between the third and fourth face features. Optionally, the predicted portrait and the sample portrait are input into a pre-trained face recognition model. The pre-trained face recognition model performs face recognition on both the predicted portrait and the sample portrait respectively to obtain the third face feature of the predicted portrait and the fourth face feature of the sample portrait. Then, the second loss function is determined based on the third difference between the third and fourth face features. The pre-trained face recognition model can be any existing face recognition model, possessing the ability to identify face features from images containing faces.

[0055] The third loss value, Loss3, can be calculated using the following formula: Loss3 =

[0056] in, Indicates a sample portrait. This represents the predicted portrait output by the first model. This represents the output of the face recognition model after the sample portrait is input, i.e., the second facial feature of the sample portrait. This indicates the output of the face recognition model after the predicted portrait is input, which is the first facial feature of the predicted portrait.

[0057] Step A2: Determine the fourth loss value of the first model based on the sample image and the reference portrait. The fourth loss value is used to characterize the third degree of dissimilarity of facial features in the sample image and the reference portrait.

[0058] When calculating the fourth loss value, the sample image and the reference portrait can be input into a pre-trained identity recognition model. The pre-trained identity recognition model performs face recognition on the sample image and the reference portrait respectively, obtaining the fifth face feature of the sample image and the sixth face feature of the reference portrait. Then, based on the fourth difference between the fifth and sixth face features, the fourth loss value is determined.

[0059] The fourth loss value, Loss4, can be calculated using the following formula: Loss4= 1-cos(arcface(x),arcface(face)) Where x represents the sample image, and face represents the reference portrait generated by the generative adversarial network. arcface(x) represents the output of the identity recognition model after the sample image is input, i.e., the fifth face feature of the sample image. arcface(face) represents the output of the identity recognition model after the reference portrait is input, i.e., the sixth face feature of the reference portrait.

[0060] It should be noted that when calculating the fourth loss value, although facial features are also extracted, the facial features involved in the fourth loss value focus more on ID features that can uniquely represent the user's identity, such as iris features and skin features that can be extracted from the face.

[0061] Since the fourth loss value can characterize the difference in facial features between the sample image and the reference portrait, and further, the difference in ID features between the sample image and the reference portrait, the reference portrait generated by the generative adversarial network (GAN) is constrained based on the fourth loss value. This allows the GAN to eventually generate a reference portrait with the same ID features as the sample image as the number of iterations increases, thus ensuring that the generated predicted portrait also maintains the same ID. Based on this, as the number of iterations increases, the first model can gradually learn the ability of the GAN to generate portraits with the same ID features, ensuring that the portrait generated by the trained second model is more realistic.

[0062] Step A3: Determine the model loss value of the first model based on the first loss value, the second loss value, the third loss value, and the fourth loss value, and adjust the model parameters of the first model based on the model loss value.

[0063] The model loss value (Loss) can be calculated using the following formula: Loss= aLoss1+bLoss2+cLoss3+dLoss4 Where a, b, c, and d are the weights of the first loss value, the second loss value, the third loss value, and the fourth loss value, respectively. The values ​​of these weights can be customized as needed. For example, Loss = 5Loss1 + 2Loss2 + 5Loss3 + 2Loss4.

[0064] As can be seen, the model loss value calculated in this embodiment can simultaneously represent the following information: the difference between the predicted portrait and the sample portrait, the realism of the image information in the predicted portrait, the difference in facial features between the predicted portrait and the sample portrait, and the difference in facial features between the sample image and the reference portrait. This makes the difference between the predicted portrait and the sample portrait smaller and smaller, the realism of the image information in the predicted portrait higher and higher, the difference in facial features between the predicted portrait and the sample portrait smaller and smaller, and the difference in facial features between the sample image and the reference portrait smaller and smaller, as the number of iterations increases. This achieves the generation of detailed and realistic portraits based on generative methods, and the generated portraits have the same ID features as the sample images and the sample reference images.

[0065] Figure 6 This is a schematic diagram illustrating a training method for a portrait generation model according to another embodiment of this application. In this embodiment, the structure of each module in the portrait generation model is divided in more detail. The portrait generation model to be trained includes: a first feature extraction module, a second feature extraction module, a third feature extraction module, a first feature fusion module, a second feature fusion module, a generative adversarial network, and a discriminator for constraining the generative adversarial network.

[0066] like Figure 6 As shown, the first feature extraction module includes a feature extraction layer (such as the pre-trained Mobilenetv3) and four convolutional layers (conv). These four convolutional layers (conv) serve as downsampling layers, each being a 3x3 convolution with a stride of 2 and a kernel size of 64. Furthermore, each convolutional layer is followed by an activation layer (not shown in the figure) to activate the extracted features. Within these four convolutional layers (conv), each convolutional layer (conv) yields the first face features and background features after downsampling the sample image.

[0067] The second feature extraction module also includes a feature extraction layer (such as the pre-trained Mobilenetv3) and four convolutional layers (conv). The structure of the second feature extraction module is similar to that of the first feature extraction module, the only difference being the input data. Each convolutional layer (conv) in the second feature extraction module obtains reference image features after downsampling the sample reference image. These reference image features include the first face feature, the first hairstyle feature, and the background feature. The last convolutional layer (conv) in the second feature extraction module inputs the output reference image features into the last convolutional layer (conv) in the first feature extraction module. The last convolutional layer (conv) in the first feature extraction module fuses the output of the previous convolutional layer (including the first face feature and background feature) with the reference image features; the fusion method can be concatenated. The fused features are reduced to one-dimensional features, i.e., latent features, which include the second face feature and the second hairstyle feature. The latent features are input into the first feature fusion module and the generative adversarial module. It can be seen that the main function of the input sample reference image is to provide hairstyle features, enabling the feature extraction network to extract accurate hairstyle features, rather than using contour recognition to identify the hairstyle region.

[0068] Generative Adversarial Networks (GANs) can be stylegan generators, which can generate new images that simulate real images and have powerful image generation capabilities. The stylegan generator generates a reference portrait based on latent features (including second facial features and second hairstyle features), and then inputs this reference portrait into a third feature extraction module. The third feature extraction module consists of two convolutional layers (conv) used to extract features from the reference portrait, obtaining the corresponding first portrait features. These first portrait features are then input into a second feature fusion module. The two convolutional layers in the third feature extraction module are both 3*3*64 convolutions.

[0069] The first feature fusion module comprises four convolutional layers (conv) and one convolutional output layer (conv1*1). The four convolutional layers (conv) serve as upsampling layers, each a 3*3 convolution with a stride of 1 / 2 and a kernel size of 64. Each convolutional layer is followed by an activation layer (not shown in the diagram) to activate the upsampled features. Each of these four convolutional layers (conv) yields a more refined feature set (including facial and hairstyle features) upsampled by a factor of 2. The third convolutional layer (conv) in the first feature fusion module inputs the upsampled features into the second feature fusion module. The features obtained after upsampling in the third convolutional layer (conv) include the features obtained after upsampling in the previous convolutional layer (conv) and the features obtained after downsampling in downsampling in the second convolutional layer (conv) in the first feature extraction module. The second feature, obtained by fusing this information, is then input into the second feature fusion module.

[0070] The second feature fusion module may include an AdaIN layer. Using AdaIN's fusion function, it fuses the first portrait features input from the third feature extraction module and the second portrait features input from the first feature fusion module to obtain sample portrait features of the sample object. These sample portrait features are then input to the last convolutional layer (conv) in the first feature fusion module. The fourth convolutional layer (conv) fuses the sample portrait features with the features obtained after downsampling from the first convolutional layer (conv) in the first feature extraction module. The fused features are then input to the convolutional output layer (conv1*1), which outputs the predicted portrait. The convolutional output layer (conv1*1) is a 1*1*3 convolution and also inputs the predicted portrait to the discriminator.

[0071] The discriminator consists of four convolutional layers (conv), one fully connected layer, and one output layer. The discriminator has 64,1 channels. The output discriminator evaluates the realism of the image information in the predicted portrait, typically using 1 and 0 to represent whether the predicted portrait is real or fake. The discriminator's evaluation of the realism of the image information in the predicted portrait is used to constrain the generative adversarial network (GAN), thus forming adversarial training with the GAN.

[0072] In this embodiment, the method for calculating the model loss value of the portrait generation model to be trained is the same as the method for calculating the model loss value of the first model, which has been described in detail in the above embodiments and will not be repeated here.

[0073] As can be seen, the portrait generation model training method provided in this application relies on sample data that includes not only sample reference images (i.e., complete portraits) but also sample images (i.e., images after removing the hairstyle area). This allows the model to extract accurate and rich hairstyle features, i.e., hairstyle semantic structure, during training, rather than simple hairstyle contour recognition, thus providing strong semantic structure support for model training. Furthermore, since the sample reference images provide accurate and rich hairstyle features for the model training process, and the sample images provide accurate and rich facial features, the generated predicted portraits can possess both clear and natural hairstyle semantic structure and facial semantic structure as the model iterates through training. This further enables the trained portrait generation model to generate clear and natural portraits. Furthermore, by determining the model loss value of the portrait generation model to be trained based on sample images, predicted portraits, sample portraits, and reference portraits, the model training process is subject to the following constraints: the difference between the predicted portrait generated by the portrait generation model and the sample portrait, the realism of the image information in the predicted portrait, the difference between the facial features in the predicted portrait and the sample portrait, and the difference between the facial features in the sample image and the reference portrait. Therefore, as the number of iterations increases, the difference between the predicted portrait and the sample portrait becomes smaller and smaller, the realism of the image information in the predicted portrait becomes higher and higher, the difference between the facial features in the predicted portrait and the sample portrait becomes smaller and smaller, and the difference between the facial features in the sample image and the reference portrait becomes smaller and smaller. This achieves the generation of detailed and realistic portraits based on generative methods, and the generated portraits have the same ID features as the sample images and the sample reference images.

[0074] Figure 7 This is a schematic flowchart of a portrait generation method according to an embodiment of this application, such as... Figure 7 As shown, the method includes: S702, Obtain the image dataset of the first object, the image dataset including the first image and the first reference image.

[0075] The first reference image includes a background region image, a face region image, and a hairstyle region image. The first image is obtained by removing the hairstyle region image from the first reference image. The first image includes a face region image and a background region image.

[0076] When acquiring the image dataset of the first object, a first reference image of the first object can be acquired first. Then, a masking process can be applied to the hairstyle region image in the first reference image to obtain the first image. Alternatively, the hairstyle region image in the first reference image can be directly erased to obtain the first image.

[0077] S704, input the image dataset into the second model for image processing to obtain the portrait of the first object.

[0078] The second model is trained using the training method of the first model in any of the above embodiments. In the portrait generation scenario, the first model is the portrait generation model to be trained, and the second model is the trained portrait generation model.

[0079] In one embodiment, the second model includes a first feature extraction module and a second feature extraction module. Based on this, when an image dataset is input into the second model for image processing, the image processing procedure can be executed as follows: B1-B3: Step B1: Extract features from the first image and the first reference image to obtain the facial features and hairstyle features of the first object.

[0080] Optionally, the first feature extraction module extracts features from the first image to obtain the facial features of the first object.

[0081] The hairstyle features of the first object are obtained by using the second feature extraction module and by extracting features from the first reference image.

[0082] Step B2: Determine the portrait features of the first subject based on facial features and hairstyle features.

[0083] Step B3: Generate a portrait of the first object based on the portrait features.

[0084] Figure 8 This is a schematic diagram illustrating a portrait generation method according to an embodiment of this application, such as... Figure 8 As shown, the trained portrait generation model includes: a first feature extraction module, a second feature extraction module, and a feature fusion module.

[0085] First, the image dataset of the first object is input into the trained portrait generation model. Specifically, the first person's image is input into the first feature extraction module, and a first reference image is input into the second feature extraction module. The first feature extraction module extracts features from the first image to obtain the facial features. The second feature extraction module extracts features from the first reference image to obtain the reference image features of the first object. The second feature extraction module then inputs the reference image features back into the first feature extraction module, which determines the hairstyle features based on the reference image features and the facial features.

[0086] Next, the first feature extraction module inputs facial features and hairstyle features into the feature fusion module. The feature fusion module fuses the facial features and hairstyle features input from the first feature extraction module to obtain the portrait features of the first object. Then, a portrait of the first object is generated based on the portrait features.

[0087] Figure 9 A schematic diagram illustrating a portrait generation method according to another embodiment of this application, such as... Figure 9 As shown, the first feature extraction module includes a feature extraction layer (such as the pre-trained Mobilenetv3) and four convolutional layers (conv). These four convolutional layers (conv) serve as downsampling layers, each being a 3x3 convolution with a stride of 2 and a kernel size of 64. Furthermore, each convolutional layer is followed by an activation layer (not shown in the figure) to activate the extracted features. Within these four convolutional layers (conv), each convolutional layer (conv) yields downsampled facial and background features from the first image.

[0088] The second feature extraction module also includes a feature extraction layer (such as the pre-trained Mobilenetv3) and four convolutional layers (conv). The structure of the second feature extraction module is similar to that of the first feature extraction module, the only difference being the input data. Each convolutional layer (conv) in the second feature extraction module obtains reference image features after downsampling the first reference image. These reference image features include face features, hairstyle features, and background features. The last convolutional layer (conv) in the second feature extraction module inputs the output reference image features into the last convolutional layer (conv) in the first feature extraction module. The last convolutional layer (conv) in the first feature extraction module fuses the output of the previous convolutional layer (including face and background features) with the reference image features; the fusion can be done concatenated. The fused features are reduced to one-dimensional features, i.e., latent features, which include the reduced face and hairstyle features. These latent features are then input into the feature fusion module. It can be seen that the first reference image provides hairstyle features for the portrait generation model, enabling the feature extraction network to extract accurate hairstyle features, rather than using contour recognition to identify the hairstyle region.

[0089] The feature fusion module consists of four convolutional layers (conv) and one convolutional output layer (conv1*1). The four conv layers are upsampling layers, each a 3x3 convolution with a stride of 1 / 2 and a kernel size of 64. Each convolutional layer is followed by an activation layer (not shown in the diagram) to activate the upsampled features. Within these four conv layers, each layer obtains more refined features (including facial and hairstyle features) upsampled by a factor of 2. The final convolutional layer fuses these features to obtain the portrait features of the first object, which are then input to the convolutional output layer (conv1*1), which outputs the portrait of the first object.

[0090] The technical solution of the above embodiments involves acquiring an image dataset of a first object and inputting the image dataset into a trained second model. The trained second model then generates a portrait of the first object based on the first image and the first reference image. Because accurate and rich facial and hairstyle features can be extracted during the training of the second model, the portrait generated by the trained second model possesses both clear and natural hairstyle and facial semantic structures, achieving the effect of generating a clear and natural portrait.

[0091] In summary, specific embodiments of this subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing can be advantageous.

[0092] The above are the model processing method and portrait generation method provided in the embodiments of this application. Based on the same idea, the embodiments of this application also provide a model processing device and a portrait generation device.

[0093] Figure 10 This is a schematic block diagram of a model processing apparatus according to an embodiment of this application, such as... Figure 10 As shown, the model processing device includes: The first acquisition module 101 is used to acquire a sample dataset of the sample object; the sample dataset includes: a sample image, a sample reference image, and a sample portrait; the sample reference image includes a background region image, a face region image, and a hairstyle region image; the sample image is obtained by removing the hairstyle region image from the sample reference image. The first processing module 102 is used to input the sample dataset into a first model for image processing to obtain a predicted portrait of the sample object; the image processing includes: extracting features from the sample image and the sample reference image to obtain a first face feature and a first hairstyle feature of the sample object; generating a reference portrait of the sample object based on the first face feature and the first hairstyle feature; determining the sample portrait features of the sample object based on the reference portrait, the first face feature and the first hairstyle feature, and generating the predicted portrait based on the sample portrait features; The model training module 103 is used to adjust the model parameters of the first model based on the sample image, the predicted portrait, the sample portrait, and the reference portrait to obtain the second model.

[0094] In one embodiment, when the first processing module 102 generates a reference portrait of the sample object based on the first facial features and the first hairstyle features, it performs the following steps: The first facial features and the first hairstyle features are subjected to dimensionality reduction processing to obtain the second facial features and the second hairstyle features of the sample object; A reference portrait of the sample object is generated based on the second facial features and the second hairstyle features.

[0095] In one embodiment, when the first processing module 102 determines the sample portrait features of the sample object based on the reference portrait, the first facial features, and the first hairstyle features, it performs the following steps: Feature extraction is performed on the reference portrait to obtain the first portrait feature corresponding to the reference portrait; The second facial feature, the second hairstyle feature, and the first facial feature are fused together to obtain the second portrait feature of the sample object. The first portrait feature and the second portrait feature are subjected to feature fusion processing to obtain the sample portrait feature.

[0096] In one embodiment, when the model training module 103 adjusts the model parameters of the first model based on the sample image, the predicted portrait, the sample portrait, and the reference portrait, it performs the following steps: Based on the predicted portrait and the sample portrait, a first loss value, a second loss value, and a third loss value of the first model are determined; the first loss value is used to characterize a first degree of difference between the predicted portrait and the sample portrait; the second loss value is used to characterize the realism of the image information in the predicted portrait; and the third loss value is used to characterize a second degree of difference in facial features between the predicted portrait and the sample portrait. Based on the sample image and the reference portrait, a fourth loss value for the first model is determined; the fourth loss value is used to characterize the third difference in facial features between the sample image and the reference portrait. Based on the first loss value, the second loss value, the third loss value, and the fourth loss value, the model loss value of the first model is determined, and the model parameters are adjusted based on the model loss value.

[0097] In one embodiment, the first model includes: a first feature extraction module and a second feature extraction module; When the first processing module 102 extracts features from the sample image and the sample reference image to obtain the first facial features and the first hairstyle features of the sample object, it performs the following steps: The first feature extraction module extracts features from the sample image to obtain the first facial feature of the sample image. The second feature extraction module extracts features from the sample reference image to obtain the reference image features of the sample object; The first feature extraction module processes the first facial features and the reference image features to obtain the first hairstyle features.

[0098] The apparatus of this application acquires a sample dataset of a sample object, which includes sample images, sample reference images, and sample portraits. The sample reference images include background region images, face region images, and hairstyle region images. The sample dataset is input into a first model to extract features from the sample images and sample reference images, obtaining the face features and hairstyle features of the sample object. It is evident that the sample data relied upon during model training includes not only the sample reference images (i.e., the complete portrait) but also the sample images after removing the hairstyle region images (i.e., only including the background region images and face region images). This enables the model to extract accurate and rich hairstyle features, i.e., the semantic structure of the hairstyle, rather than simple hairstyle contour recognition, providing strong semantic structural support for model training. Furthermore, a reference portrait of the sample object is generated based on the face features and hairstyle features. The portrait features of the sample object are determined based on the reference portrait, face features, and hairstyle features, and a predicted portrait of the sample object is generated based on the portrait features. The model parameters of the first model are adjusted based on the sample images, predicted portraits, sample portraits, and reference portraits to obtain a second model. Since the sample reference images provide accurate and rich hairstyle features for the model training process, and the sample images provide accurate and rich facial features for the model training process, as the model is trained iteratively, the generated predicted portraits can have both clear and natural hairstyle semantic structure and facial semantic structure. This further enables the trained second model to generate clear and natural portraits.

[0099] Those skilled in the art will understand that Figure 10 The model processing device in the document can be used to implement the model processing method described above. The details of the method description should be similar to those in the previous section. To avoid being too complicated, they will not be repeated here.

[0100] Figure 11 This is a schematic block diagram of a portrait generation device according to an embodiment of this application, such as... Figure 11 As shown, the portrait generation device includes: The second acquisition module 111 is used to acquire an image dataset of the first object; the image dataset includes a first image and a first reference image; the first reference image includes a background region image, a face region image and a hairstyle region image; the first image is obtained by removing the hairstyle region image from the first reference image. The second processing module 112 is used to input the image dataset into the second model for image processing to obtain the portrait of the first object; the image processing includes: extracting features from the first image and the first reference image to obtain the facial features and hairstyle features of the first object; determining the portrait features of the first object based on the facial features and the hairstyle features; and generating the portrait of the first object based on the portrait features.

[0101] In one embodiment, the second model includes: a first feature extraction module and a second feature extraction module; When the second processing module 112 extracts features from the first image and the first reference image to obtain the facial features and hairstyle features of the first object, and determines the portrait features of the first object based on the facial features and hairstyle features, it performs the following steps: The first feature extraction module extracts features from the first image to obtain the facial features of the first object. The second feature extraction module extracts features from the first reference image to obtain the hairstyle features of the first object; The facial features and hairstyle features are fused to obtain the portrait features of the first object.

[0102] The apparatus described in the above embodiment acquires an image dataset of a first object and inputs the image dataset into a trained second model. The trained second model then generates a portrait of the first object based on the first image and a first reference image. Because accurate and rich facial and hairstyle features are extracted during the training of the second model, the portrait generated by the trained second model possesses both clear and natural hairstyle and facial semantic structures, achieving the effect of generating a clear and natural portrait.

[0103] Those skilled in the art will understand that Figure 11 The portrait generation device in the document can be used to implement the portrait generation method described above. The details of the method should be similar to those described in the previous section. To avoid being too complicated, they will not be repeated here.

[0104] Following the same line of thought, embodiments of this application also provide an electronic device, such as... Figure 12As shown. Electronic devices can vary considerably due to differences in configuration or performance, and may include one or more processors 1201 and memory 1202. Memory 1202 may store one or more application programs or data. Memory 1202 may be temporary or persistent storage. The application programs stored in memory 1202 may include one or more modules (not shown), each module may include a series of computer-executable instructions for the electronic device. Furthermore, processor 1201 may be configured to communicate with memory 1202 and execute the series of computer-executable instructions in memory 1202 on the electronic device. The electronic device may also include one or more power supplies 1203, one or more wired or wireless network interfaces 1204, one or more input / output interfaces 1205, and one or more keyboards 1206.

[0105] Specifically, in this embodiment, the electronic device includes a memory and one or more programs, wherein one or more programs are stored in the memory, and one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the electronic device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Obtain a sample dataset of the sample objects; the sample dataset includes: sample images, sample reference images, and sample portraits; the sample reference images include background region images, face region images, and hairstyle region images; the sample images are obtained by removing the hairstyle region images from the sample reference images. The sample dataset is input into a first model for image processing to obtain a predicted portrait of the sample object. The image processing includes: extracting features from the sample image and the sample reference image to obtain a first facial feature and a first hairstyle feature of the sample object; generating a reference portrait of the sample object based on the first facial feature and the first hairstyle feature; determining the sample portrait features of the sample object based on the reference portrait, the first facial feature, and the first hairstyle feature, and generating the predicted portrait based on the sample portrait features. Based on the sample image, the predicted portrait, the sample portrait, and the reference portrait, the model parameters of the first model are adjusted to obtain the second model.

[0106] The technical solution of this application involves acquiring a sample dataset of the sample object. This dataset includes sample images, sample reference images, and sample portraits. The sample reference images include background region images, face region images, and hairstyle region images. The sample dataset is input into a first model, and feature extraction is performed on the sample images and sample reference images to obtain the face features and hairstyle features of the sample object. It is evident that the sample data used in training the model includes not only the sample reference images (i.e., the complete portrait) but also the sample images after removing the hairstyle region images (i.e., only the background region images and face region images). This allows the model to extract accurate and rich hairstyle features, i.e., the semantic structure of the hairstyle, rather than simple hairstyle contour recognition, providing strong semantic structural support for model training. Furthermore, a reference portrait of the sample object is generated based on the face features and hairstyle features. The portrait features of the sample object are determined based on the reference portrait, face features, and hairstyle features, and a predicted portrait of the sample object is generated based on the portrait features. The model parameters of the first model are adjusted based on the sample images, predicted portraits, sample portraits, and reference portraits to obtain the second model. Since the sample reference images provide accurate and rich hairstyle features for the model training process, and the sample images provide accurate and rich facial features for the model training process, as the model is trained iteratively, the generated predicted portraits can have both clear and natural hairstyle semantic structure and facial semantic structure. This further enables the trained second model to generate clear and natural portraits.

[0107] In another embodiment, the electronic device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the electronic device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following: Obtain an image dataset of a first object; the image dataset includes a first image and a first reference image; the first reference image includes a background region image, a face region image, and a hairstyle region image; the first image is obtained by removing the hairstyle region image from the first reference image. The image dataset is input into a second model for image processing to obtain a portrait of the first object; the image processing includes: extracting features from the first image and the first reference image to obtain facial features and hairstyle features of the first object; determining portrait features of the first object based on the facial features and hairstyle features; and generating a portrait of the first object based on the portrait features.

[0108] The technical solution of the above embodiments involves acquiring an image dataset of a first object and inputting the image dataset into a trained second model. The trained second model then generates a portrait of the first object based on the first image and the first reference image. Because accurate and rich facial and hairstyle features can be extracted during the training of the second model, the portrait generated by the trained second model possesses both clear and natural hairstyle and facial semantic structures, achieving the effect of generating a clear and natural portrait.

[0109] This application also proposes a computer-readable storage medium that stores one or more computer programs, each including instructions that, when executed by an electronic device including multiple applications, enable the electronic device to perform various processes of the above-described model processing method embodiments, specifically for executing: Obtain a sample dataset of the sample objects; the sample dataset includes: sample images, sample reference images, and sample portraits; the sample reference images include background region images, face region images, and hairstyle region images; the sample images are obtained by removing the hairstyle region images from the sample reference images. The sample dataset is input into a first model for image processing to obtain a predicted portrait of the sample object. The image processing includes: extracting features from the sample image and the sample reference image to obtain a first facial feature and a first hairstyle feature of the sample object; generating a reference portrait of the sample object based on the first facial feature and the first hairstyle feature; determining the sample portrait features of the sample object based on the reference portrait, the first facial feature, and the first hairstyle feature, and generating the predicted portrait based on the sample portrait features. Based on the sample image, the predicted portrait, the sample portrait, and the reference portrait, the model parameters of the first model are adjusted to obtain the second model.

[0110] The technical solution of this application involves acquiring a sample dataset of the sample object. This dataset includes sample images, sample reference images, and sample portraits. The sample reference images include background region images, face region images, and hairstyle region images. The sample dataset is input into a first model, and feature extraction is performed on the sample images and sample reference images to obtain the face features and hairstyle features of the sample object. It is evident that the sample data used in training the model includes not only the sample reference images (i.e., the complete portrait) but also the sample images after removing the hairstyle region images (i.e., only the background region images and face region images). This allows the model to extract accurate and rich hairstyle features, i.e., the semantic structure of the hairstyle, rather than simple hairstyle contour recognition, providing strong semantic structural support for model training. Furthermore, a reference portrait of the sample object is generated based on the face features and hairstyle features. The portrait features of the sample object are determined based on the reference portrait, face features, and hairstyle features, and a predicted portrait of the sample object is generated based on the portrait features. The model parameters of the first model are adjusted based on the sample images, predicted portraits, sample portraits, and reference portraits to obtain the second model. Since the sample reference images provide accurate and rich hairstyle features for the model training process, and the sample images provide accurate and rich facial features for the model training process, as the model is trained iteratively, the generated predicted portraits can have both clear and natural hairstyle semantic structure and facial semantic structure. This further enables the trained second model to generate clear and natural portraits.

[0111] This application also proposes a computer-readable storage medium that stores one or more computer programs, the computer programs including instructions that, when executed by an electronic device including multiple applications, enable the electronic device to perform various processes of the portrait generation method embodiments described above, and specifically for performing: Obtain an image dataset of a first object; the image dataset includes a first image and a first reference image; the first reference image includes a background region image, a face region image, and a hairstyle region image; the first image is obtained by removing the hairstyle region image from the first reference image. The image dataset is input into a second model for image processing to obtain a portrait of the first object; the image processing includes: extracting features from the first image and the first reference image to obtain facial features and hairstyle features of the first object; determining portrait features of the first object based on the facial features and hairstyle features; and generating a portrait of the first object based on the portrait features.

[0112] The technical solution of the above embodiments involves acquiring an image dataset of a first object and inputting the image dataset into a trained second model. The trained second model then generates a portrait of the first object based on the first image and the first reference image. Because accurate and rich facial and hairstyle features can be extracted during the training of the second model, the portrait generated by the trained second model possesses both clear and natural hairstyle and facial semantic structures, achieving the effect of generating a clear and natural portrait.

[0113] This application provides a computer program product, including a computer program that is executed by a processor to implement the various processes of the above-described model processing method embodiment, or the computer program is executed by a processor to implement the various processes of the above-described portrait generation method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0114] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0115] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0116] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0117] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0118] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0119] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0120] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0121] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0122] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0123] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0124] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0125] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0126] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.

Claims

1. A model processing method, characterized in that, include: Obtain a sample dataset of the sample objects; the sample dataset includes: sample images, sample reference images, and sample portraits; the sample reference images include background region images, face region images, and hairstyle region images; the sample images are obtained by removing the hairstyle region images from the sample reference images. The sample dataset is input into a first model for image processing to obtain a predicted portrait of the sample object. The image processing includes: extracting features from the sample image and the sample reference image to obtain a first facial feature and a first hairstyle feature of the sample object; generating a reference portrait of the sample object based on the first facial feature and the first hairstyle feature; determining the sample portrait features of the sample object based on the reference portrait, the first facial feature, and the first hairstyle feature, and generating the predicted portrait based on the sample portrait features. Based on the sample image, the predicted portrait, the sample portrait, and the reference portrait, the model parameters of the first model are adjusted to obtain the second model.

2. The method according to claim 1, characterized in that, The step of generating a reference portrait of the sample object based on the first facial features and the first hairstyle features includes: The first facial features and the first hairstyle features are subjected to dimensionality reduction processing to obtain the second facial features and the second hairstyle features of the sample object; A reference portrait of the sample object is generated based on the second facial features and the second hairstyle features.

3. The method according to claim 2, characterized in that, The step of determining the sample portrait features of the sample object based on the reference portrait, the first facial features, and the first hairstyle features includes: Feature extraction is performed on the reference portrait to obtain the first portrait feature corresponding to the reference portrait; The second facial feature, the second hairstyle feature, and the first facial feature are fused together to obtain the second portrait feature of the sample object. The first portrait feature and the second portrait feature are subjected to feature fusion processing to obtain the sample portrait feature.

4. The method according to claim 2, characterized in that, The step of adjusting the model parameters of the first model based on the sample image, the predicted portrait, the sample portrait, and the reference portrait includes: Based on the predicted portrait and the sample portrait, a first loss value, a second loss value, and a third loss value of the first model are determined; the first loss value is used to characterize a first degree of difference between the predicted portrait and the sample portrait; the second loss value is used to characterize the realism of the image information in the predicted portrait; and the third loss value is used to characterize a second degree of difference in facial features between the predicted portrait and the sample portrait. Based on the sample image and the reference portrait, a fourth loss value for the first model is determined; the fourth loss value is used to characterize the third difference in facial features between the sample image and the reference portrait. Based on the first loss value, the second loss value, the third loss value, and the fourth loss value, the model loss value of the first model is determined, and the model parameters are adjusted based on the model loss value.

5. The method according to claim 1, characterized in that, The first model includes: a first feature extraction module and a second feature extraction module; The step of extracting features from the sample image and the sample reference image to obtain the first facial features and the first hairstyle features of the sample object includes: The first feature extraction module extracts features from the sample image to obtain the first facial feature of the sample image. The second feature extraction module extracts features from the sample reference image to obtain the reference image features of the sample object; The first feature extraction module processes the first facial features and the reference image features to obtain the first hairstyle features.

6. A portrait generation method, characterized in that, include: Obtain an image dataset of a first object; the image dataset includes a first image and a first reference image; The first reference image includes a background region image, a face region image, and a hairstyle region image; The first image is obtained by removing the hairstyle area from the first reference image; The image dataset is input into a second model for image processing to obtain a portrait of the first object; the image processing includes: extracting features from the first image and the first reference image to obtain the facial features and hairstyle features of the first object; Based on the facial features and hairstyle features, determine the portrait features of the first object; generate a portrait of the first object based on the portrait features.

7. The method according to claim 6, characterized in that, The second model includes: a first feature extraction module and a second feature extraction module; The step of extracting features from the first image and the first reference image to obtain the facial features and hairstyle features of the first object; and determining the portrait features of the first object based on the facial features and hairstyle features, includes: The first feature extraction module extracts features from the first image to obtain the facial features of the first object. The second feature extraction module extracts features from the first reference image to obtain the hairstyle features of the first object; The facial features and hairstyle features are fused to obtain the portrait features of the first object.

8. A model processing device, characterized in that, include: The first acquisition module is used to acquire a sample dataset of the sample object; the sample dataset includes: a sample image, a sample reference image, and a sample portrait; the sample reference image includes a background region image, a face region image, and a hairstyle region image; the sample image is obtained by removing the hairstyle region image from the sample reference image. A first processing module is used to input the sample dataset into a first model for image processing to obtain a predicted portrait of the sample object; the image processing includes: extracting features from the sample image and the sample reference image to obtain a first facial feature and a first hairstyle feature of the sample object; generating a reference portrait of the sample object based on the first facial feature and the first hairstyle feature; determining the sample portrait features of the sample object based on the reference portrait, the first facial feature, and the first hairstyle feature, and generating the predicted portrait based on the sample portrait features; The model training module is used to adjust the model parameters of the first model based on the sample image, the predicted portrait, the sample portrait, and the reference portrait to obtain the second model.

9. A portrait generation device, characterized in that, include: The second acquisition module is used to acquire an image dataset of the first object; the image dataset includes a first image and a first reference image; The first reference image includes a background region image, a face region image, and a hairstyle region image; The first image is obtained by removing the hairstyle area from the first reference image; The second processing module is used to input the image dataset into the second model for image processing to obtain the portrait of the first object; the image processing includes: extracting features from the first image and the first reference image to obtain the facial features and hairstyle features of the first object; Based on the facial features and hairstyle features, determine the portrait features of the first object; generate a portrait of the first object based on the portrait features.

10. An electronic device, characterized in that, The device includes a processor and a memory electrically connected to the processor, the memory storing a computer program, the processor being configured to call and execute the computer program from the memory to implement the model processing method as described in any one of claims 1-5, or the processor being configured to call and execute the computer program from the memory to implement the portrait generation method as described in any one of claims 6-7.

11. A computer-readable storage medium, characterized in that, The storage medium is used to store a computer program that can be executed by a processor to implement the model processing method as described in any one of claims 1-5, or the computer program can be executed by a processor to implement the portrait generation method as described in any one of claims 6-7.

12. A computer program product, characterized in that, The method includes a computer program that is executed by a processor to implement the model processing method as described in any one of claims 1-5, or the computer program that is executed by a processor to implement the portrait generation method as described in any one of claims 6-7.