Image Processing Method, Apparatus, Device, and Medium
By performing face segmentation and fusion processing on the target image, using multiple control vectors of the hairstyle generation model to control the hairstyle transformation image characteristics, the problem of unclear texture details in the prior art is solved, and high-quality hairstyle transformation images are achieved.
Patent Information
- Application Number
- CN202210672542.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-06-14
AI Technical Summary
In the prior art, the image hairstyle transformation method is difficult to ensure the texture details of each part, resulting in poor texture details of the hairstyle transformation image.
The mask image is obtained by performing face segmentation processing on the target image, and after the fusion process, the encoder is input to obtain the image encoding vector, and the hairstyle generation model is used to control multiple image features on the hairstyle transformed image according to the multiple control vectors.
Improves the texture detail clarity of the hairstyle-changing image, reduces artifacts, and improves the quality of the hairstyle-changing image.
Smart Images

Figure CN115115560B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of computer vision technology, specifically to the field of image generation technology, and particularly to an image processing method, apparatus, device, and medium. Background Art
[0002] In related technologies, for the hair style transformation method of an image, a commonly used method is to use the mask images of different regions segmented from the image as a guide. For example, different values are directly assigned to different regions to change the style of the region, such as realizing the transformation of the hair style. However, the texture details of the hair style transformation image obtained in this way are poor. Another commonly used method is to use a generative adversarial network, taking a face image as the input to obtain a hair style transformation image. Since the generative adversarial network generates the hair style transformation image according to the overall feature vector of the input face image, and the texture details of the hair are relatively complex, it is difficult to ensure that each local part of the image meets the requirements for the hair style transformation image generated by using the overall feature vector. Usually, the contours such as the face change and the texture details of the hair are also poor. Summary of the Invention
[0003] In view of the above-mentioned defects or deficiencies in the prior art, it is desired to provide an image processing method, apparatus, device, and medium, which combine mask images to respectively adjust multiple image features on the hair style transformation image, and improve the clarity of the texture details of the hair style transformation image.
[0004] In a first aspect, the present application provides an image processing method, which includes: performing face segmentation processing on a target image to obtain a mask image of the target image, where the mask image is used to represent the hair region of the target image and the face region of the target image; performing fusion processing on the mask image and the target image to obtain a fusion image; inputting the fusion image into an encoder for encoding processing to obtain an image encoding vector; and inputting the image encoding vector into a hair style generation model to obtain a hair style transformation image of the target image, where the hair style generation model is based on multiple control vectors obtained from the image encoding vector and controls multiple image features on the hair style transformation image according to the multiple control vectors.
[0005] Second aspect, the present application provides an image processing apparatus, the apparatus comprising: a segmentation unit configured to perform face segmentation processing on a target image to obtain a mask image of the target image, wherein the mask image is used to characterize a hair region of the target image and a face region of the target image; a fusion unit configured to perform fusion processing on the mask image and the target image to obtain a fused image; an encoding unit configured to input the fused image into an encoder for encoding processing to obtain an image encoding vector; a generation unit configured to input the image encoding vector into a hairstyle generation model to obtain a hairstyle transformation image of the target image, wherein the hairstyle generation model is based on a plurality of control vectors obtained from the image encoding vector and controls a plurality of image features on the hairstyle transformation image.
[0006] In a possible implementation manner of the second aspect, the generation unit is specifically configured to: obtain the plurality of control vectors according to the image encoding vector; respectively input the plurality of control vectors into a plurality of convolutional networks of the hairstyle generation model to obtain a plurality of feature maps; and perform fusion processing on the plurality of feature maps to obtain the hairstyle transformation image of the target image.
[0007] In a possible implementation manner of the second aspect, it further comprises: a decoder configured to perform decoding and reconstruction processing on the image encoding vector to obtain a plurality of reconstructed images, wherein hairstyles included in the reconstructed images are different from those included in the target image; the generation unit is specifically configured to: input the image encoding vector into the hairstyle generation model to obtain a plurality of output images, wherein resolutions of the plurality of output images are different; and perform fusion processing on the plurality of output images and the plurality of reconstructed images to obtain the hairstyle transformation image of the target image.
[0008] In a possible implementation manner of the second aspect, the generation unit is specifically configured to: input output image 1 to output image n into fusion module 1 to fusion module n one by one, and input reconstructed image 1 to reconstructed image n into the fusion module 1 to the fusion module n one by one; the fusion module 1 to the fusion module n are connected in sequence, where n is the number of the plurality of output images; perform fusion processing on the output image 1 and the reconstructed image 1 through the fusion module 1 to obtain a fused image 1; perform fusion processing on the output image x, the reconstructed image x, and an upsampled image y through the fusion module x to obtain a fused image x; where x is a positive integer greater than 1 and less than or equal to n, the upsampled image y is obtained by upsampling the fused image y output by the fusion module y, and y = x - 1; when x is n, obtain the fused image n output by the fusion module n, and use the fused image n as the hairstyle transformation image of the target image.
[0009] In a possible implementation of the second aspect, the encoding unit is specifically configured to:
[0010] Input the image encoding vector into a plurality of deconvolution networks of the decoder to obtain the plurality of reconstructed images.
[0011] In a possible implementation of the second aspect, the training process of the decoder includes: inputting a sample image into the encoder to obtain an encoding vector of the sample image; inputting the encoding vector of the sample image into an initial decoder, and training the initial decoder according to the loss between the output of the initial decoder and the label image to obtain the decoder, where the hairstyle included in the label image is different from the hairstyle included in the sample image.
[0012] In a possible implementation of the second aspect, the training process of the encoder includes: inputting a sample image into an initial encoder, and training the initial encoder according to the loss between the output of the initial encoder and the latent vector of the sample image to obtain the encoder.
[0013] In a possible implementation of the second aspect, the segmentation unit is specifically configured to: input the target image into an autoencoder to obtain a mask image of the target image, where the training samples of the autoencoder are face images, and the labels of the training samples are the mask images of the face images, and the mask images of the face images are used to represent the face region and the hairstyle region of the face images.
[0014] In a possible implementation of the second aspect, the training process of the autoencoder includes: inputting the training samples into an initial autoencoder, and training the initial autoencoder according to the loss between the output of the initial autoencoder and the labels of the training samples to obtain the autoencoder.
[0015] In a possible implementation of the second aspect, the fusion unit is specifically configured to: add and fuse the pixels of the mask image and the corresponding pixels of the target image to obtain the fused image; perform attention mechanism feature fusion on the region of the mask image and the region of the target image to obtain the fused image; add and fuse the pixel color channels of the mask image and the corresponding pixel color channels of the target image to obtain the fused image.
[0016] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the embodiment of the present application is implemented.
[0017] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in the embodiment of the present application is implemented.
[0018] For the image processing method, device, equipment and medium proposed in the present application, the hairstyle generation model obtains a plurality of control vectors according to the image coding vector of the input fused image, and can respectively control a plurality of image features on the hairstyle transformation image according to the plurality of control vectors. This way of separately controlling different image features on the hairstyle transformation image by different control vectors avoids affecting the change of other image features when adjusting a certain image feature, and can control the hairstyle transformation image with more detailed features, improving the texture effect of the hairstyle transformation image. In addition, the image coding vector input to the hairstyle generation model is the feature of the image after fusing the target image and the mask image, and the image coding vector includes the features of the hair region and the face region of the mask image representing the mask image. Therefore, according to the features of the hair region and the face region of the mask image, the face contour and hair range of the hairstyle transformation image can be effectively controlled, reducing the artifacts in the hairstyle transformation image and improving the display effect of the hairstyle change image.
[0019] It can be seen that the present application can combine the mask image to separately adjust a plurality of image features on the hairstyle transformation image. Compared with the prior art, such as the way of simply relying on the face image to generate the hairstyle transformation image, the clarity of the texture details of the hairstyle transformation image is improved, and the quality of the hairstyle transformation image can be effectively improved.
[0020] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0021] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives and advantages of the present application will become more apparent:
[0022] Figure 1 It is a schematic flowchart of the image processing method provided by the embodiment of the present application;
[0023] Figure 2 It is a schematic diagram of the autoencoder provided by the embodiment of the present application;
[0024] Figure 3 It is a schematic diagram of the processing model of the image processing method provided by the embodiment of the present application;
[0025] Figure 4 It is a schematic diagram of the fusion model for fusing the output image and the reconstructed image provided by the embodiment of the present application;
[0026] Figure 5 Schematic diagram of the fusion module provided by an embodiment of the present application;
[0027] Figure 6 Structural schematic diagram of the image processing device provided by an embodiment of the present application;
[0028] Figure 7 Structural schematic diagram of the computer device provided by an embodiment of the present application. Detailed implementation manners
[0029] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention, rather than limiting the invention. In addition, it should be noted that for the sake of convenience of description, only the parts related to the invention are shown in the drawings.
[0030] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.
[0031] Currently, when performing a hairstyle transformation on a person in an image, such as changing the hairstyle, obtaining a transformed hairstyle by adding bangs, or obtaining a hairstyle transformation image with bangs added, the texture of the hairstyle, bangs, etc. on the obtained hairstyle transformation image is not clear.
[0032] Based on this, the present application proposes an image processing method, device, equipment and storage medium, which can separately control multiple image features on the hairstyle transformation image by using the fusion image obtained by fusing the target image and the mask image of the target image, thereby greatly improving the clarity of the texture of the hairstyle and the like on the hairstyle transformation image.
[0033] In the implementation environment of the embodiment of the present application, a personal computer device, a mobile terminal, etc. can perform face segmentation processing on the target image to obtain the mask image of the target image; perform fusion processing on the mask image and the target image to obtain a fusion image; input the fusion image into an encoder for encoding processing to obtain an image encoding vector; input the image encoding vector into a hairstyle generation model to obtain the hairstyle transformation image of the target image.
[0034] Alternatively, it can also be implemented by a server. For example, a personal computer device and a mobile terminal send requests to the server. The server performs face segmentation processing on the target image to obtain the mask image of the target image; performs fusion processing on the mask image and the target image to obtain a fusion image; inputs the fusion image into an encoder for encoding processing to obtain an image encoding vector; inputs the image encoding vector into a hairstyle generation model to obtain the hairstyle transformation image of the target image, and finally returns the hairstyle transformation image to the personal computer device and the mobile terminal.
[0035] Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.
[0036] An embodiment of the present application provides an image processing method, which can be applied to Figure 7 the computer device shown in Figure 1 As shown, the method includes the following steps:
[0037] 101. Perform face segmentation processing on the target image to obtain a mask image of the target image, where the mask image is used to represent the hair region and the face region of the target image.
[0038] Among them, the target image can be a target image including any face, and the mask image of the target image can be a black-and-white mask image. For example: the face region and the hair region on the mask image are white, and other regions on the mask image are black.
[0039] For face segmentation processing of the target image, first identify the face and hair on the target image, then perform segmentation to obtain the face region, hair region and other regions, and finally, assign different regions to black and white to obtain the mask image.
[0040] In a possible implementation, performing face segmentation processing on the target image to obtain a mask image of the target image includes: inputting the target image into an autoencoder to obtain a mask image of the target image, where the training samples of the autoencoder are face images, and the labels of the training samples are mask images of the face images, and the mask images of the face images are used to represent the face region and the hairstyle region of the face images.
[0041] As Figure 2 shown, the autoencoder includes an encoder 1 and a decoder 1. The target image passes through the encoder 1 to extract the features of the target image, and then generates a mask image corresponding to the target image through the decoder 1.
[0042] In this example, the encoder 1 includes multiple convolutional layers. As Figure 2 shown, taking the encoder 1 including three convolutional layers as an example, the encoder 1 is stacked by 3 convolutional layers. For example, the first convolutional layer, the second convolutional layer and the third convolutional layer from left to right. The first convolutional layer, the second convolutional layer and the third convolutional layer can extract image features of the target image at different resolutions; the decoder 1 includes multiple deconvolutional layers. AsFigure 2 As shown, taking the decoder 1 including three transposed convolution layers as an example, the decoder 1 is stacked by 3 transposed convolution layers, such as the first transposed convolution layer, the second transposed convolution layer, and the third transposed convolution layer from right to left. The first transposed convolution layer, the second transposed convolution layer, and the third transposed convolution layer can generate a mask image of the target image through transposed convolution operations according to the image features of the target image at different resolutions transmitted by the encoder 1. In order to improve the accuracy of the obtained mask image, in one example, as Figure 2 shown, the image features of the encoder 1 and the decoder 1 at the corresponding resolutions are connected in a skip connection manner, that is: the first convolution layer is connected to the first transposed convolution layer, the second convolution layer is connected to the second transposed convolution layer, and the third convolution layer is connected to the third transposed convolution layer. By connecting the image features of the encoder 1 and the decoder 1 at the corresponding resolutions in this skip connection manner, it can be ensured that each transposed convolution layer of the decoder 1 in the autoencoder can obtain richer features and reduce information loss. Thus, a more accurate mask image can be decoded.
[0043] A downsampling layer (i.e., pooling layer) can also be provided between the convolution layers of the autoencoder. For example, a downsampling layer is provided between the first convolution layer and the second convolution layer, and a downsampling layer is provided between the second convolution layer and the third convolution layer. The input of the pooling layer comes from the previous convolution layer, and its main function is to provide strong robustness. For example: max-pooling takes the maximum value in a small area. At this time, if other values in this area change slightly, or the image is slightly translated, the result after pooling remains unchanged. The pooling layer is sandwiched between consecutive convolution layers, and it is used to compress the amount of data and parameters and reduce overfitting.
[0044] It can be understood that the autoencoder used in this example is a neural network model, that is: the target image is input into this neural network model to obtain the mask image of the target image. Before using this neural network model, it can be pre-trained, that is: the autoencoder is trained.
[0045] In one embodiment, there is also provided the training process of the above-mentioned autoencoder, including: inputting training samples into the initial autoencoder, and training the initial autoencoder according to the loss between the output of the initial autoencoder and the labels of the training samples to obtain the autoencoder. That is to say, first, obtain the training samples of face images and obtain the mask images of the training samples, and use the mask images as labels. Then, train the initial autoencoder according to the loss between the output of the initial autoencoder and the labels of the training samples to obtain the autoencoder. That is to say, pre-obtain the training samples of face images and obtain the mask images of the face images, use the mask images of the face images as labels, train the initial autoencoder according to the training samples of the face images and the mask images of the face images, adjust the parameters of the initial autoencoder according to the loss between the output of the initial autoencoder and the labels of the training samples. When the loss meets the requirements, such as being less than a certain preset value, the training is completed and the autoencoder is obtained. In this way, when any target image including a face is input into the autoencoder, the mask image of the target image can be obtained.
[0046] 102. Perform a fusion process on the mask image and the target image to obtain a fused image.
[0047] Such as Figure 3 shown, Figure 3 the "+" in
[0048] (1) Pixel-by-pixel addition method: Add and fuse the pixels of the mask image and the corresponding pixels of the target image to obtain a fused image.
[0049] (2) Fusion method using an attention mechanism: Perform feature fusion on the regions of the mask image and the target image using the attention mechanism to obtain a fused image.
[0050] (3) Concatenate fusion method: Add and fuse the pixel color channels of the mask image and the corresponding pixel color channels of the target image to obtain a fused image. For example: The color of the pixels of the target image is represented by a three-channel color vector, that is, [R, G, B], while the color of the pixels of the mask image is represented by a one-channel color vector. For example: black is 0 and white is 255. Therefore, after adding and fusing the corresponding pixel color channels, the pixel color is represented by a four-channel color, that is: the color representations of the first channel, the second channel, the third channel, and the fourth channel are: [R, G, B, 0 or 255]. In this embodiment, the concatenate fusion method is used to fuse the target image and the mask image to obtain a fused image. That is: through Figure 2The mask image of the target image extracted by the shown autoencoder (i.e., the face segmentation mask) and the target image input to the autoencoder are fused by concatenation, such that the pixel colors in the fused image adopt a four-channel color representation [R, G, B, 0 or 255]. Since the fourth channel represents the features of the hair region and the face region of the target image, therefore, the fusion of the mask image and the target image can, when generating the hairstyle transformation image subsequently, control the contour range of the face and the range of the hair in the generated hairstyle transformation image according to the features of the hair region and the face region represented by the fourth channel, effectively reducing the artifacts in the generated hairstyle transformation image and improving the quality of the hairstyle transformation image.
[0051] 103. Input the fused image into the encoder for encoding processing to obtain an image encoding vector.
[0052] Among them, the image encoding vector can be the latent code W, that is, a one-dimensional feature vector obtained by mapping from a high-dimensional space to a one-dimensional space.
[0053] In a possible implementation manner, as Figure 3 shown, extract the image features of the fused image through the encoder 2 to obtain the latent code W, that is, the image encoding vector.
[0054] In this example, the encoder 2 is a feature extraction network including multiple convolutional layers. Therefore, before obtaining the image encoding vector from the fused image using the encoder 2, a process of training the encoder 2 is also provided, including: inputting the sample image into the initial encoder, and training the initial encoder according to the loss between the output of the input initial encoder and the latent vector of the sample image to obtain the encoder 2. For example: for the encoder 2 using the ResNet-50 feature extraction network, in the training stage, the distance metric function L1 loss can be used as the loss function between the latent code W (i.e., the image encoding vector) output by the encoder 2 and the latent vector of the sample image. Through the distance metric function L1 loss, the loss between the output of the input initial encoder and the latent vector of the sample image is obtained. When this loss meets the set conditions, the training of the initial encoder is completed to obtain the encoder 2.
[0055] L1 loss is a loss function based on distance metric. Usually, the input data is mapped to a feature space based on distance metric, such as the Euclidean space. The mapped samples are regarded as points in the space, and a suitable loss function is used to measure the distance between the true value and the predicted value of the samples in the feature space. Generally speaking, the smaller the distance between two points in the feature space, the better the prediction performance of the encoder 2.
[0056] 104. Input the image encoding vector into the hairstyle generation model to obtain the hairstyle transformation image of the target image, where the hairstyle generation model is based on multiple control vectors obtained from the image encoding vector and controls multiple image features on the hairstyle transformation image according to the multiple control vectors.
[0057] In a possible implementation, inputting the image encoding vector into the hairstyle generation model to obtain the hairstyle transformation image of the target image includes: obtaining multiple control vectors according to the image encoding vector; respectively inputting the multiple control vectors into multiple convolutional networks of the hairstyle generation model to obtain multiple feature maps; and performing a fusion process on the multiple feature maps to obtain the hairstyle transformation image of the target image. That is to say, the hairstyle generation model does not directly generate multiple feature maps based on the image encoding vector latent code. Namely, it is different from the traditional generator directly feeding the latent code W to the input layer of the generator. That is, different from directly feeding the latent code W to the input layer, first, the latent code W can be non-linearly mapped through a mapping network such as a fully connected layer including 8 layers, f: Z→W, where Z and W have the same dimension, such as 512×1. This mapping network encodes the latent code W into an intermediate vector, and the intermediate vector is then passed to a generation network such as a 18-layer generation network. Therefore, each layer in the generation network can generate a control vector, so 18 control vectors can be obtained, enabling different control vectors to control different image features, where the image features include but are not limited to color features, texture features, etc.
[0058] In the above description, the multiple image features refer to multiple features extracted from the image, such as color features, texture features, etc. The multiple feature maps refer to images with different resolutions, for example, images with different resolutions output by the transposed convolutional layers of multiple convolutional networks of the hairstyle generation model.
[0059] Since the hairstyle generation model is based on multiple control vectors obtained from the image encoding vectors, and correspondingly controls multiple image features on the hairstyle transformation image, the clarity of the texture details of the hairstyle transformation image can be improved. For example, a part of the multiple control vectors controls a part of the image features on the hairstyle transformation image, and another part of the multiple control vectors controls another part of the image features on the hairstyle transformation image. When a part of the multiple control vectors controls a part of the image features on the hairstyle transformation image, it will not affect another part of the image features on the hairstyle transformation image. Similarly, when another part of the multiple control vectors controls another part of the image features on the hairstyle transformation image, it will not affect a part of the image features on the hairstyle transformation image. That is to say, when adjusting some image features on the image to meet the requirements, other image features will not change. Therefore, this way of separately controlling different image features can make the texture details, quality, etc. of the finally obtained image higher.
[0060] As Figure 3 shown, the image encoding vector is input into the hairstyle generation model 3 to obtain the hairstyle transformation image of the target image. In this example, the hairstyle generation model 3 uses a neural network including StyleGAN. Since StyleGAN contains rich portrait texture information, it can be used as prior knowledge for generating the texture and details of the hairstyle transformation image. Specifically, inputting the latent vector latent code W into StyleGAN and modulating the weights of each convolutional layer in StyleGAN helps to generate rich texture details.
[0061] Among them, the main working principle of StyleGAN is as follows: Instead of directly feeding the latent code W into the input layer, that is, different from the traditional generator directly feeding the latent code W into the input layer of the generator, that is, different from directly feeding the latent code W into the input layer, first, the latent code W is non-linearly mapped through a mapping network of an 8-layer fully connected layer f: Z→W, where Z and W have the same dimension, such as 512×1. This mapping network encodes the latent code W into an intermediate vector, and then the intermediate vector is passed to a generation network such as an 18-layer generation network. Therefore, 18 control vectors are obtained, enabling different control vectors to control different image features, where the image features include but are not limited to color features, texture features, etc. Among them, the role of the mapping network is to untangle the entanglement between image features. For example, it controls the color of the face generation on the image with a resolution of 32×32, but this will also change the image features such as texture controlled on the image with a resolution of 32×32. Therefore, the mapping network can untangle the features of the latent code W. Furthermore, specific image features can be adjusted based on different control vectors without affecting other image features. Thus, different image features can be adjusted separately, and the texture details of each image feature can meet the requirement of clarity, resulting in a high-quality hairstyle transformation image with high-definition texture details.
[0062] In a specific application, the target image is a portrait including a certain hairstyle, and a portrait with a different hairstyle is required. As Figure 3 shown, after processing the target image ( Figure 3 the leftmost image in Figure 3 ) through an autoencoder, a mask image of the target image is obtained. After cascade fusion processing of the two, a fusion image is obtained. The fusion image is input into encoder 2, and an image encoding vector W is obtained through encoder 2. The image encoding vector W is respectively input into decoder 2 and hairstyle generation model 3. Decoder 2 can obtain multiple reconstructed images with different hairstyles from the target image according to the image encoding vector W. The hairstyle generation model 3 generates multiple output images according to the image encoding vector W. Then, after the multiple output images and multiple reconstructed images are fused through fusion model 4, finally, the hairstyle transformation image of the portrait after hairstyle transformation is obtained ( Figure 3 the rightmost image in
[0063] ) is obtained. As Figure 3 shown, a hairstyle transformation image with a different hairstyle from the target image and with bangs added is obtained.
[0063] In the image processing method provided by the embodiment of the present application, the hairstyle generation model obtains multiple control vectors based on the image encoding vector of the input fused image. For example, the image encoding vector is non-linearly mapped through a mapping network of a multi-layer fully connected layer. The mapping network encodes the image encoding vector into an intermediate vector, and then the intermediate vector is passed to a multi-layer generation network. Therefore, each layer in the multi-layer generation network can generate a control vector, and thus multiple control vectors can be obtained. Then, multiple image features on the hairstyle transformation image can be respectively controlled according to the multiple control vectors. This way of separately controlling different image features on the hairstyle transformation image by different control vectors avoids affecting the change of other image features when adjusting a certain image feature, and can control the hairstyle transformation image with more detailed features, improving the texture effect of the hairstyle transformation image. In addition, the image encoding vector input to the hairstyle generation model includes the features of the image after fusing the target image and the mask image. Therefore, according to the features of the hair region and the face region of the mask image, the face contour and hair range of the hairstyle transformation image can be effectively controlled, reducing the artifacts in the hairstyle transformation image and improving the display effect of the hairstyle change image.
[0064] In a possible implementation manner, the method further includes: inputting the image encoding vector into a decoder for decoding and reconstruction processing to obtain multiple reconstructed images, where the hairstyles included in the reconstructed images are different from the hairstyles included in the target image. On this basis, inputting the image encoding vector into the hairstyle generation model to obtain the hairstyle transformation image of the target image includes: inputting the image encoding vector into the hairstyle generation model to obtain multiple output images, where the resolutions of the multiple output images are different; performing fusion processing on the multiple output images and the multiple reconstructed images to obtain the hairstyle transformation image of the target image.
[0065] Among them, the reconstructed image is an image obtained by performing a hairstyle transformation on the head portrait in the target image. That is to say, compared with the target image, the image quality of the reconstructed image is different and the hairstyles of the head portraits in the two are different, that is: the head portraits in the two represent the same person with different hairstyles. And the multiple output head portraits are the same as the multiple feature maps in the above embodiment, that is: the hairstyle generation model obtains multiple control vectors according to the image encoding vector; the multiple control vectors are respectively input into multiple convolutional networks of the hairstyle generation model, and then multiple output images can be obtained. The output image is generated based on the target image, and its different image features on the output image can be separately controlled by different control vectors, which can improve the quality of the output image. Since the output image is generated based on the target image by the hairstyle transformation image, compared with the target image, the hairstyle generation model uses a neural network including StyleGAN, for example. Therefore, the quality of the output image generated by the hairstyle generation model is higher, but the portrait in the image is different from the portrait in the target image.
[0066] Such as Figure 3As shown, the image encoding vector W is input into the hairstyle generation model 3 to obtain multiple output images; the image encoding vector W is input into the decoder 2 for decoding and reconstruction processing to obtain multiple reconstructed images; the multiple output images and the multiple reconstructed images are fused to obtain the hairstyle transformation image of the target image.
[0067] In this example, the image encoding vector is input into the decoder 2 for decoding and reconstruction processing to obtain multiple reconstructed images, including: inputting the image encoding vector into multiple deconvolution networks of the decoder 2 to obtain multiple reconstructed images. That is, multiple reconstructed images are generated through multiple stacked deconvolution layers.
[0068] In a possible implementation manner, the present application provides a process for training the decoder 2, including: inputting a sample image into the encoder to obtain the encoding vector of the sample image; inputting the encoding vector of the sample image into the initial decoder, and training the initial decoder according to the loss between the output of the input initial decoder and the label image to obtain the decoder 2, where the hairstyle included in the label image is different from the hairstyle included in the sample image. For example: the label image and the sample image are supervised through the reconstruction loss function L1 loss and perceptual loss to implement the training of the decoder 2. The encoder 2 and the decoder 2 can form a generative adversarial network. Since the texture of hair is generally complex, the hair strands in the image generated by this generative adversarial network lack a realistic texture and it is difficult to ensure the clarity of the hairstyle texture.
[0069] As can be seen from the above embodiment, the hairstyle generation model 3 can obtain the hairstyle transformation image of the target image with clear texture details. There may be some differences in some forms on the human face, such as: the facial features on the five sense organs are different from those on the target image. That is, the hairstyle generation model uses the StyleGAN neural network, which can separately control different image features on the output image through different control vectors, improving the quality of the image. However, based on the working principle of the StyleGAN neural network, there are differences between the portrait in the output image and the portrait in the target image. Therefore, in the present application, in order to make the face on the hairstyle transformation image more consistent with the face on the target image, multiple reconstructed images that are more consistent with the face on the target image are obtained by using the decoder 2. Then, the multiple reconstructed images are fused with the multiple output images output by the hairstyle generation model 3. Thus, the face on the obtained hairstyle transformation image is consistent with the face on the target image, and the clarity of the texture details of the hairstyle transformation image is improved.
[0070] In a possible implementation manner, such as Figure 4As shown in the figure, the fusion model 4 can perform a fusion process on multiple output images and multiple reconstructed images to obtain a hairstyle transformation image of the target image, including: respectively inputting the multiple output images and the multiple reconstructed images into a plurality of cascaded fusion modules one by one, and performing a fusion process on the multiple output images and the multiple reconstructed images to obtain the hairstyle transformation image of the target image.
[0071] As Figure 4 shown in the figure, assuming that the number of the multiple reconstructed images, the multiple output images, and the multiple fusion modules is denoted as n, where n is a positive integer, the fusion model 4 performs a fusion process on the multiple output images and the multiple reconstructed images to obtain the hairstyle transformation image of the target image, specifically including: inputting output image 1 to output image n into fusion module 1 to fusion module n one by one, and inputting reconstructed image 1 to reconstructed image n into fusion module 1 to fusion module n one by one; fusion module 1 to fusion module n are connected in sequence, and n is the number of the multiple output images; performing a fusion process on output image 1 and reconstructed image 1 through fusion module 1 to obtain fusion image 1; performing a fusion process on output image x, reconstructed image x, and the upsampled image y through fusion module x to obtain fusion image x; x is a positive integer greater than 1 and less than or equal to n, and the upsampled image y is obtained by upsampling the fusion image y output by fusion module y, where y = x - 1; when x is n, obtaining the fusion image n output by fusion module n, and taking the fusion image n as the hairstyle transformation image of the target image.
[0072] Combined with Figure 4 shown in the figure, Feat_D1, Feat_D2 to Feat_Dn are multiple reconstructed images reconstructed by the decoder 2, and Feat_G1, Feat_G2 to Feat_Gn are multiple output images output by the hairstyle generation model 3. The fusion modules 1 to n will respectively perform the image fusion process, that is: the fusion modules 1 to n gradually fuse Feat_D1, Feat_D2 to Feat_Dn, Feat_G1, Feat_G2 to Feat_Gn to obtain the hairstyle change image. It is named the progressive fusion method.
[0073] The progressive fusion method, that is: the specific fusion process of the fusion modules 1 to n is as follows: fusion module 1 fuses Feat_G1 and Feat_D1, and upsamples the fused features, and then inputs them into fusion module 2; fusion module 2 fuses Feat_G2 and Feat_D2 and the output features of fusion module 1, and upsamples the fused features, and then inputs them into the next fusion module, and so on, until the fusion image output by the last fusion module n is used as the hairstyle transformation image of the target image.
[0074] Among them, the fusion modules 1 to n have the same structure. As Figure 5 shown, taking the fusion module 1 as an example, the input will first be cascaded through a concatenation layer, and then the features will be extracted through a convolutional structure layer (conv + bn + relu). Subsequently, the features will pass through a global pooling layer, a 1x1 convolutional layer, and a sigmoid layer in sequence to extract the global information of each channel. Then, the global information and the features will be multiplied channel by channel through a multiplication layer. Next, the result of the multiplication and the features will be added pixel by pixel through an addition layer to achieve fusion. This way of channel-wise weighting and pixel-wise weighting can better integrate the images, effectively preserve the features on the target image, and improve the texture details of the hairstyle transformation image. Furthermore, it can enhance the quality of the hairstyle transformation image and obtain a high-quality hairstyle transformation image.
[0075] In an embodiment of the present application, the outputs of the global pooling layer, the 1x1 convolutional layer, and the sigmoid layer in each of the fusion modules 1 to n can also be supervised to further enhance the clarity of the hairstyle transformation image.
[0076] Figure 6 It is a block diagram of an image processing device according to an embodiment of the present application.
[0077] As Figure 6 shown, the image processing device includes: a segmentation unit 601, a fusion unit 602, an encoding unit 603, and a generation unit 604, where:
[0078] The segmentation unit 601 is configured to perform face segmentation processing on the target image to obtain a mask image of the target image, where the mask image is used to represent the hair region and the face region of the target image; the fusion unit 602 is configured to perform fusion processing on the mask image and the target image to obtain a fusion image; the encoding unit 603 is configured to input the fusion image into an encoder for encoding processing to obtain an image encoding vector; the generation unit 604 is configured to input the image encoding vector into a hairstyle generation model to obtain a hairstyle transformation image of the target image, where the hairstyle generation model is based on multiple control vectors obtained from the image encoding vector and controls multiple image features on the hairstyle transformation image according to the multiple control vectors.
[0079] In a possible implementation manner, the generation unit 604 is specifically configured to: obtain multiple control vectors according to the image encoding vector; input the multiple control vectors into multiple convolutional networks of the hairstyle generation model respectively to obtain multiple feature maps; and perform fusion processing on the multiple feature maps to obtain a hairstyle transformation image of the target image.
[0080] In a possible implementation, it further includes: a decoder, configured to perform decoding and reconstruction processing on the image encoding vector to obtain a plurality of reconstructed images, where the hairstyles included in the reconstructed images are different from the hairstyles included in the target image; the generating unit 604 is specifically configured to: input the image encoding vector into a hairstyle generation model to obtain a plurality of output images, where the resolutions of the plurality of output images are different; perform fusion processing on the plurality of output images and the plurality of reconstructed images to obtain a hairstyle transformation image of the target image.
[0081] In a possible implementation, the generating unit 604 is specifically configured to: input output image 1 to output image n into fusion module 1 to fusion module n one by one, and input reconstructed image 1 to reconstructed image n into the fusion module 1 to the fusion module n one by one; the fusion module 1 to the fusion module n are connected in sequence, where n is the number of the plurality of output images; perform fusion processing on the output image 1 and the reconstructed image 1 through the fusion module 1 to obtain fusion image 1; perform fusion processing on output image x, reconstructed image x, and upsampled image y through the fusion module x to obtain fusion image x; where x is a positive integer greater than 1 and less than or equal to n, the upsampled image y is obtained by upsampling the fusion image y output by the fusion module y, and y = x - 1; when x is n, obtain the fusion image n output by the fusion module n, and use the fusion image n as the hairstyle transformation image of the target image.
[0082] In a possible implementation, the encoding unit 603 is specifically configured to:
[0083] Input the image encoding vector into a plurality of transposed convolutional networks of the decoder to obtain a plurality of reconstructed images.
[0084] In a possible implementation, the training process of the decoder includes: inputting a sample image into an encoder to obtain an encoding vector of the sample image; inputting the encoding vector of the sample image into an initial decoder, and training the initial decoder according to the loss between the output input to the initial decoder and a label image to obtain a decoder, where the hairstyles included in the label image are different from the hairstyles included in the sample image.
[0085] In a possible implementation, the training process of the encoder includes: inputting a sample image into an initial encoder, and training the initial encoder according to the loss between the output input to the initial encoder and a latent vector of the sample image to obtain an encoder.
[0086] In a possible implementation, the splitting unit 601 is specifically configured to: input the target image into an autoencoder to obtain a mask image of the target image, where the training samples of the autoencoder are face images, and the labels of the training samples are the mask images of the face images. The mask images of the face images are used to represent the face regions and hairstyle regions of the face images.
[0087] In a possible implementation, the training process of the autoencoder includes: inputting the training samples into an initial autoencoder, and training the initial autoencoder according to the loss between the output of the initial autoencoder and the labels of the training samples to obtain the autoencoder.
[0088] In a possible implementation, the fusion unit 602 is specifically configured to: add and fuse the pixels of the mask image and the corresponding pixels of the target image to obtain a fused image; perform feature fusion on the regions of the mask image and the target image using an attention mechanism to obtain a fused image; add and fuse the pixel color channels of the mask image and the corresponding pixel color channels of the target image to obtain a fused image.
[0089] In the image processing apparatus provided in the embodiments of the present application, the hairstyle generation model obtains a plurality of control vectors according to the image encoding vector of the input fused image. For example, the image encoding vector is non-linearly mapped through a mapping network of a multi-layer fully connected layer. The mapping network encodes the image encoding vector into an intermediate vector, and the intermediate vector is then passed to a multi-layer generation network. Therefore, each layer in the multi-layer generation network can generate a control vector. Therefore, a plurality of control vectors can be obtained. Then, according to the plurality of control vectors, a plurality of image features on the hairstyle transformation image can be controlled respectively. This way of separately controlling different image features on the hairstyle transformation image by different control vectors avoids affecting other image features when adjusting a certain image feature, and can control the hairstyle transformation image with more detailed features, improving the texture effect of the hairstyle transformation image. In addition, the image encoding vector input to the hairstyle generation model includes the features of the image after fusing the target image and the mask image. Therefore, according to the features of the hair region and the face region of the mask image, the face contour and hair range of the hairstyle transformation image can be effectively controlled, reducing the artifacts in the hairstyle transformation image and improving the display effect of the hairstyle transformation image.
[0090] It should be understood that the various units described in the image processing apparatus and the reference Figure 1It corresponds to each step in the described image processing method. Thus, the operations and features described above for the method also apply to the image processing apparatus and the units included therein, and will not be elaborated here. The image processing apparatus can be pre-implemented in a browser or other secure applications of a computer device, or can be loaded into the browser or its secure applications of the computer device by means of downloading or the like. The corresponding units in the image processing apparatus can cooperate with the units in the computer device to implement the solutions of the embodiments of the present application.
[0091] Among the several modules or units mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0092] It should be noted that for the details not disclosed in the image processing apparatus of the embodiments of the present application, please refer to the details disclosed in the above embodiments of the present application, and will not be elaborated here.
[0093] Next, refer to Figure 7 , Figure 7 which shows a schematic structural diagram of a computer device suitable for implementing the embodiments of the present application. As Figure 7 shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the programs stored in the read-only memory (ROM) 702 or the programs loaded from the storage section 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation instructions of the system are also stored. The CPU 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0094] The following components are connected to the I / O interface 705; an input section 706 including a keyboard, a mouse, etc.; an output section 707 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed, so that the computer program read from it can be installed into the storage section 708 as needed.
[0095] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart Figure 1 can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 709 and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above functions defined in the system of the present application are executed.
[0096] It should be noted that the computer-readable medium shown in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0097] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operation instructions of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two connected blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that executes the specified functions or operation instructions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0098] The units or modules involved in the embodiments described in the present application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. For example, it can be described as: a processor includes a first collection module, a second collection module, and a transmission module. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0099] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments, or may exist separately without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above programs are used by one or more processors to execute the image processing method described in the present application.
[0100] The above description is only the preferred embodiments of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present application.
Claims
1. An image processing method, characterized in that, Including: Performing face segmentation on a target image to obtain a mask image of the target image, where the mask image is used to represent the hair region of the target image and the face region of the target image; Performing fusion processing on the mask image and the target image to obtain a fused image; Inputting the fused image into an encoder for encoding processing to obtain an image encoding vector; Inputting the image encoding vector into a decoder for decoding and reconstruction processing to obtain a plurality of reconstructed images, where the hairstyles included in the reconstructed images are different from the hairstyles included in the target image; Inputting the image encoding vector into the hairstyle generation model to obtain a plurality of output images, where the resolutions of the plurality of output images are different; Performing fusion processing on the plurality of output images and the plurality of reconstructed images to obtain a hairstyle transformation image of the target image, where the hairstyle generation model is based on a plurality of control vectors obtained from the image encoding vector and controls a plurality of image features on the hairstyle transformation image according to the plurality of control vectors, and different control vectors can control different image features on the hairstyle transformation image.
2. The image processing method according to claim 1, wherein The step of inputting the image encoding vector into the hairstyle generation model to obtain a hairstyle transformation image of the target image includes: Obtaining the plurality of control vectors according to the image encoding vector; Inputting the plurality of control vectors into a plurality of convolutional networks of the hairstyle generation model respectively to obtain a plurality of feature maps; Performing fusion processing on the plurality of feature maps to obtain a hairstyle transformation image of the target image.
3. The image processing method according to claim 1, wherein Performing fusion processing on the plurality of output images and the plurality of reconstructed images to obtain a hairstyle transformation image of the target image includes: Inputting output image 1 to output image n into fusion module 1 to fusion module n one by one, and inputting reconstructed image 1 to reconstructed image n into the fusion module 1 to the fusion module n one by one; the fusion module 1 to the fusion module n are connected in sequence, and n is the number of the plurality of output images; Performing fusion processing on the output image 1 and the reconstructed image 1 through the fusion module 1 to obtain a fused image 1; Performing fusion processing on output image x, reconstructed image x, and an upsampled image y through the fusion module x to obtain a fused image x; x is a positive integer greater than 1 and less than or equal to n, and the upsampled image y is obtained by upsampling the fused image y output by the fusion module y, and y = x - 1; When x is n, obtaining the fused image n output by the fusion module n and using the fused image n as the hairstyle transformation image of the target image.
4. The image processing method according to claim 1, wherein Inputting the image encoding vector into a decoder for decoding and reconstruction processing to obtain a plurality of reconstructed images includes: Inputting the image encoding vector into a plurality of transposed convolutional networks of the decoder to obtain the plurality of reconstructed images.
5. The image processing method according to claim 4, characterized in that, The training process of the decoder includes: Inputting a sample image into the encoder to obtain an encoding vector of the sample image; Input the encoded vector of the sample image into the initial decoder, and train the initial decoder according to the loss between the output of the initial decoder and the label image to obtain the decoder, where the hairstyle included in the label image is different from the hairstyle included in the sample image.
6. The image processing method according to any one of claims 1-4, characterized in that, The training process of the encoder includes: Input the sample image into the initial encoder, and train the initial encoder according to the loss between the output of the initial encoder and the latent vector of the sample image to obtain the encoder.
7. The image processing method according to any one of claims 1-5, characterized in that Perform face segmentation on the target image to obtain the mask image of the target image, including: Input the target image into the autoencoder to obtain the mask image of the target image, where the training samples of the autoencoder are face images, and the labels of the training samples are the mask images of the face images, and the mask images of the face images are used to represent the face area and hairstyle area of the face images.
8. The image processing method according to claim 7, wherein The training process of the autoencoder includes: Input the training samples into the initial autoencoder, and train the initial autoencoder according to the loss between the output of the initial autoencoder and the labels of the training samples to obtain the autoencoder.
9. The image processing method according to any one of claims 1-5, characterized in that, Perform fusion processing on the mask image and the target image to obtain a fused image, including at least one of the following: Additively fuse the pixels of the mask image and the corresponding pixels of the target image to obtain the fused image; Perform attention mechanism feature fusion on the regions of the mask image and the target image to obtain the fused image; Additively fuse the pixel color channels of the mask image and the corresponding pixel color channels of the target image to obtain the fused image.
10. An image processing apparatus, characterized in that, Includes: A segmentation unit for performing face segmentation on the target image to obtain the mask image of the target image, where the mask image is used to represent the hair area and the face area of the target image; A fusion unit for performing fusion processing on the mask image and the target image to obtain a fused image; An encoding unit for inputting the fused image into an encoder for encoding processing to obtain an image encoding vector; A decoder for performing decoding and reconstruction processing on the image encoding vector to obtain multiple reconstructed images, where the hairstyle included in the reconstructed image is different from the hairstyle included in the target image; A generation unit for inputting the image encoding vector into the hairstyle generation model to obtain multiple output images, where the resolutions of the multiple output images are different; perform fusion processing on the multiple output images and the multiple reconstructed images to obtain the hairstyle transformation image of the target image; where the hairstyle generation model is based on multiple control vectors obtained from the image encoding vector and controls multiple image features on the hairstyle transformation image according to the multiple control vectors, and different control vectors can control different image features on the hairstyle transformation image.
11. The image processing apparatus according to claim 10, wherein The generation unit is specifically used for: Obtain the multiple control vectors according to the image encoding vector; Input the multiple control vectors into multiple convolutional networks of the hairstyle generation model respectively to obtain multiple feature maps; Perform a fusion process on the multiple feature maps to obtain a hairstyle transformation image of the target image.
12. The image processing apparatus according to claim 10, wherein The generating unit is specifically configured to: Input output image 1 to output image n into fusion module 1 to fusion module n one by one, and input reconstruction image 1 to reconstruction image n into the fusion module 1 to the fusion module n one by one; the fusion module 1 to the fusion module n are connected in sequence, and n is the number of the multiple output images; Perform a fusion process on the output image 1 and the reconstruction image 1 through the fusion module 1 to obtain a fusion image 1; Perform a fusion process on the output image x, the reconstruction image x, and the upsampled image y through the fusion module x to obtain a fusion image x; x is a positive integer greater than 1 and less than or equal to n, and the upsampled image y is obtained by upsampling the fusion image y output by the fusion module y, and y = x - 1; When x is n, obtain the fusion image n output by the fusion module n, and use the fusion image n as the hairstyle transformation image of the target image.
13. The image processing apparatus according to claim 10, wherein The encoding unit is specifically configured to: Input the image encoding vector into multiple transposed convolutional networks of the decoder to obtain the multiple reconstruction images.
14. The image processing apparatus according to claim 13, wherein The training process of the decoder includes: Input a sample image into the encoder to obtain an encoding vector of the sample image; Input the encoding vector of the sample image into an initial decoder, and train the initial decoder according to the loss between the output of the initial decoder and the label image to obtain the decoder, where the hairstyle included in the label image is different from the hairstyle included in the sample image.
15. The image processing apparatus according to any one of claims 10 to 13, characterized in that, The training process of the encoder includes: Input a sample image into an initial encoder, and train the initial encoder according to the loss between the output of the initial encoder and the latent vector of the sample image to obtain the encoder.
16. The image processing apparatus according to any one of claims 10 to 14, characterized in that, The segmentation unit is specifically configured to: Input the target image into an autoencoder to obtain a mask image of the target image, where the training samples of the autoencoder are face images, and the labels of the training samples are the mask images of the face images, and the mask images of the face images are used to characterize the face region and the hairstyle region of the face images.
17. The image processing apparatus according to claim 16, wherein The training process of the autoencoder includes: Input the training samples into an initial autoencoder, and train the initial autoencoder according to the loss between the output of the initial autoencoder and the labels of the training samples to obtain the autoencoder.
18. The image processing apparatus according to any one of claims 10 to 14, characterized in that, Perform a fusion process on the mask image and the target image to obtain a fusion image, including at least one of the following: Add and fuse the pixels of the mask image and the corresponding pixels of the target image to obtain the fusion image; Perform attention mechanism feature fusion on the region of the mask image and the region of the target image to obtain the fusion image; Add and fuse the pixel color channels of the mask image and the corresponding pixel color channels of the target image to obtain the fusion image.
19. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the image processing method according to any one of claims 1-9.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the image processing method according to any one of claims 1-9.
21. A computer program product, comprising a computer program carried on a computer-readable medium, characterized in that, The computer program includes program code for executing the image processing method according to any one of claims 1-9.
Citation Information
Patent Citations
Image processing method and device, processor, electronic equipment and storage medium
CN110399849A
Face image generation method and device, equipment and medium
CN111652828A