Image conversion method, device, electronic device and storage medium
The image conversion network jointly trained by the latent code encoding network and the generative adversarial network solves the problem of insufficient realism of image conversion in the existing technology and achieves an effect that the target image is more similar to the image to be converted.
Patent Information
- Application Number
- CN202111406904.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2041-11-24
AI Technical Summary
Existing image conversion technologies are difficult to effectively control the details of the converted images, resulting in a weaker sense of reality and affecting the user experience.
An image conversion network jointly trained by a latent code encoding network and a generative adversarial network is used to obtain multiple latent codes of different scales corresponding to the image to be converted, and then perform image conversion to enhance the realism of the target image.
The similarity between the target image and the image to be converted is improved, the realism of the image is enhanced, and the amount of calculation is reduced.
Smart Images

Figure CN114092320B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and more specifically, to an image conversion method, device, electronic device, and storage medium. Background Art
[0002] With the development of science and technology, the application scope of image conversion technology is becoming increasingly wider. Currently, image conversion can be achieved through generative adversarial networks, but the details of the converted image cannot be better controlled, resulting in a weak sense of reality in the converted target image, which affects the user experience. Summary of the Invention
[0003] In view of the above problems, the present application proposes an image conversion method, device, electronic device and storage medium to solve the above problems.
[0004] In a first aspect, an embodiment of the present application provides an image conversion method, comprising: obtaining an image to be converted; inputting the image to be converted into a trained image conversion network, and obtaining a target image output by the trained image conversion network, wherein the trained image conversion network is obtained by jointly training a latent code encoding network and a trained generative adversarial network, and the latent code encoding network is used to obtain multiple latent codes of different scales corresponding to the image to be converted.
[0005] In a second aspect, an embodiment of the present application provides an image conversion device, comprising: an image to be converted acquisition module, for acquiring an image to be converted; a target image acquisition module, for inputting the image to be converted into a trained image conversion network, and acquiring a target image output by the trained image conversion network, wherein the trained image conversion network is obtained by jointly training a latent code encoding network and a trained generative adversarial network, and the latent code encoding network is used to obtain multiple latent codes of different scales corresponding to the image to be converted.
[0006] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory is coupled to the processor, the memory stores instructions, and when the instructions are executed by the processor, the processor executes the above method.
[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, and the program code can be called by a processor to execute the above method.
[0008] The image conversion method, device, electronic device and storage medium provided in the embodiments of the present application obtain an image to be converted, input the image to be converted into a trained image conversion network, and obtain a target image output by the trained image conversion network, wherein the trained image conversion network is obtained by jointly training a latent code encoding network and a trained generative adversarial network, and the latent code encoding network is used to obtain multiple latent codes of different scales corresponding to the image to be converted, thereby achieving image conversion by obtaining the latent code corresponding to the image to be processed, which can reduce the sketch feeling of the target image and make the target image more similar to the image to be converted. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0010] Figure 1 A schematic diagram of the process of an image conversion method provided in an embodiment of the present application is shown;
[0011] Figure 2 A first image schematic diagram of the image conversion method provided in an embodiment of the present application is shown;
[0012] Figure 3 A second image schematic diagram of the image conversion method provided in an embodiment of the present application is shown;
[0013] Figure 4 A schematic diagram of the process of an image conversion method provided in an embodiment of the present application is shown;
[0014] Figure 5 A schematic diagram of a first network architecture of the image conversion method provided in an embodiment of the present application is shown;
[0015] Figure 6 Shows the application Figure 4 Schematic diagram of the flow of step S220 of the image conversion method shown;
[0016] Figure 7 A second network architecture diagram of the image conversion method provided in an embodiment of the present application is shown;
[0017] Figure 8 A schematic diagram of the process of an image conversion method provided in an embodiment of the present application is shown;
[0018] Figure 9 Shows the application Figure 8 Schematic diagram of the flow of step S320 of the image conversion method shown;
[0019] Figure 10 Shows the application Figure 8 FIG. 1 is a flow chart of step S340 of the image conversion method shown;
[0020] Figure 11 The following is a block diagram of an image conversion device according to an embodiment of the present invention;
[0021] Figure 12 A block diagram of an electronic device for executing an image conversion method according to an embodiment of the present application is shown;
[0022] Figure 13 A storage unit for storing or carrying program codes for implementing the method according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0024] Image conversion technology has a wide range of applications. It can be used in entertainment to enrich users' leisure activities, and it can also be applied to crime solving, assisting public security departments in obtaining target images to help solve cases. Currently, image conversion techniques often use a Pixel2Pixel generative adversarial network to convert the target image. Alternatively, a cyclic generative adversarial network designed using a U-Net model architecture can be used to convert the target image. However, both of these conversion methods can cause the target image to overly closely match the sketch texture, making the target image look less like a photograph, less similar to the image being processed, and more discordant.
[0025] To address the above issues, the inventors, after extensive research, have discovered and proposed the image conversion method, device, server, and storage medium provided in the embodiments of this application. By acquiring the latent code corresponding to the image to be processed and performing image conversion, the image conversion can reduce the realism of the target image, making the target image more similar to the image to be converted. The specific image conversion method is described in detail in the subsequent embodiments.
[0026] See also Figure 1 , Figure 1 FIG. 1 is a flow chart showing an image conversion method according to an embodiment of the present invention. In a specific embodiment, the image conversion method is applied to Figure 11 The image conversion device 200 and the electronic device 100 equipped with the image conversion device 200 are shown. Figure 12 ). The following will take electronic equipment as an example to illustrate the specific process of this embodiment. Figure 1The process shown in FIG. 1 is described in detail. Specifically, the image conversion method may include the following steps:
[0027] Step S110: Acquire the image to be converted.
[0028] In this embodiment, an image to be converted is obtained. The image to be converted may include an image used for sketching to generate a photo, an image used for multi-angle generation of a face, an image used for segmenting a picture to generate a corresponding picture, a face image, a vehicle image, a house image, a comic image, and the like, which are not limited here.
[0029] In some implementations, a portrait of the person to be found may be determined as the image to be converted, and the image to be converted may be acquired.
[0030] In some embodiments, the image to be converted can be directly acquired. As one approach, the electronic device can store an image collection, and the user can select an image from the collection and determine the selected image as the image to be converted. Alternatively, the electronic device can store an image collection, and the user can select an image from the collection and zoom in on the image. The user can select an area within the image, determine the area to be converted, retain the image in that area, and determine the image in that area as the image to be converted. For example, if the image collection includes an image containing multiple faces, the user can select an area of one of the faces, determine that area as the image to be converted, and acquire the image to be converted.
[0031] Step S120: inputting the image to be converted into a trained image conversion network to obtain a target image output by the trained image conversion network, wherein the trained image conversion network is obtained by jointly training a latent code encoding network and a trained generative adversarial network, and the latent code encoding network is used to obtain multiple latent codes of different scales corresponding to the image to be converted.
[0032] In this embodiment, after obtaining the image to be converted, the obtained image to be converted can be input into the image conversion network to obtain the target image output by the image conversion network. Figure 2 as well as Figure 3 The target image 20 is obtained by inputting the image to be converted 10 into the image conversion network. The image conversion network is obtained by jointly training the latent code encoding network and the trained generative adversarial network. The latent code encoding network is used to obtain multiple latent codes of different scales corresponding to the image to be converted.
[0033] It should be noted that the latent code is a feature vector, which can also be called a feature map. Specifically, it can be a multi-dimensional vector, and each value in the vector is in the range of [-1,1]. For example, it can be an 18*512 vector, and each value in the vector is in the range of [-1,1]. It can be understood that by inputting the latent code into the generative adversarial network, an image corresponding to the latent code can be generated. It can be understood that the latent code can also be understood as the features of the image extracted from the image based on the neural network. The latent code can represent the image. When the latent code is determined, the image generated based on the feature latent code is also determined. From another perspective, the latent code can also be understood as the vector output after the image passes through the convolution layer in the neural network.
[0034] In some embodiments, the generative adversarial network may be StyleGAN2-ada, which is an adaptive discriminator data enhancement method applied to StyleGAN2. A generative adversarial network is a deep learning model with many different types, including but not limited to a style-based generative adversarial network (stylegan2). The generative adversarial network mainly includes two independent neural networks, a generator and a discriminator. The generator's task is to sample a noise z from a random uniform distribution and then output synthetic data G(z). The discriminator obtains a real data x or synthetic data G(z) as input and outputs the probability that the sample is "true". During the training process, the generator strives to deceive the discriminator, while the discriminator strives to learn how to correctly distinguish between true and false samples. In this way, the two form an adversarial relationship, and the ultimate goal is for the generator to be able to generate fake samples that are sufficiently real.
[0035] In some embodiments, separate input latent codes are provided at different levels of the image to be detected, allowing the generator to control visual features at different levels. The first few layers can control higher-level details, such as head shape, pose, and hairstyle, but these are not limited here. The last few layers can control finer details, such as hair and eye color, but these are not limited here, thereby enhancing the realism of the image to be converted.
[0036] In some embodiments, the generative adversarial network can be StyleGAN2-ada, and the image conversion network can be expressed as Among them, G represents the generation of adversarial network StyleGAN2-ada, E represents the latent code encoding network, Represented as the average latent code of the generative adversarial network StyleGAN2-ada itself.
[0037] The image conversion method provided in the embodiment of the present application obtains an image to be converted, inputs the image to be converted into a trained image conversion network, and obtains a target image output by the trained image conversion network, wherein the trained image conversion network is obtained by jointly training a latent code encoding network and a trained generative adversarial network, and the latent code encoding network is used to obtain multiple latent codes of different scales corresponding to the image to be converted, thereby achieving image conversion by obtaining the latent code corresponding to the image to be processed, which can enhance the realism of the target image and make the target image more similar to the image to be converted.
[0038] See also Figure 4 , Figure 4 The flow chart of the image conversion method provided by the embodiment of the present application is shown in FIG. Figure 4 The process shown in FIG. 1 is described in detail. Specifically, the image conversion method may include the following steps:
[0039] Step S210: Acquire the image to be converted.
[0040] The detailed description of step S210 can be found in step S11 and will not be repeated here.
[0041] Step S220: inputting the image to be converted into the latent code encoding network, and obtaining a plurality of latent codes of different scales corresponding to the image to be converted output by the latent code encoding network.
[0042] In some embodiments, see Figure 5 The image to be converted 10 is input into a latent code encoding network A, and multiple latent codes at different scales corresponding to the image to be converted 10 are obtained from the output of the latent code encoding network A. The multiple latent codes are responsible for generating multiple styles. The latent code encoding network A includes an FPN module and a map2style module. The FPN module is used to fuse feature images of different scales, and the map2style module is used to perform latent code conversion on feature images of different scales. Figure 5 The FPN module is used to perform feature fusion on the feature image 11A, the feature image 12A, and the feature image 13A.
[0043] In some embodiments, the latent code encoding network uses ResNet101 (residual network) as the backbone network, and downsamples the network to be converted to obtain an 8-fold downsampled feature image 11, a 16-fold downsampled feature image 12, and a 32-fold downsampled feature image 13, and then connects the FPN module to obtain feature images 11A, 12A, and 13A of corresponding sizes. The map2style module is used to convert the feature images 11A, 12A, and 13A into 512-dimensional styleGAN2-ada latent codes w1, w2, w3, w4, w5, w6, w7, w8, w9, w10, w11, w12, w13, w14, w15, w16, w17, and w18.
[0044] See also Figure 6 , Figure 6 Shows the application Figure 4 The following is a flow chart of step S220 of the image conversion method. Figure 6 The process shown in FIG. 1 is described in detail. Specifically, the image conversion method may include the following steps:
[0045] Step S221: inputting the image to be converted into the latent code encoding network, downsampling the image to be converted, and obtaining a plurality of first feature images of different scales.
[0046] In this embodiment, the image to be converted is input into a latent code encoding network, and the image to be converted is downsampled to obtain a plurality of first feature images of different scales.
[0047] In some embodiments, see Figure 5 , the image to be converted 10 is input into the latent code encoding network, and the image to be converted is downsampled to obtain a feature image 11 downsampled 8 times, a feature image 12 downsampled 16 times, and a feature image 13 downsampled 32 times, wherein the feature image 11, the feature image 12, and the feature image 13 are all first feature images. Downsampling the image to be converted can reduce the computational complexity of the electronic device.
[0048] Step S222: performing feature fusion on the multiple first feature images of different scales to obtain multiple second feature images of different scales.
[0049] In some embodiments, multiple first feature images of different scales may be fused to obtain multiple second feature images of different scales. Figure 5 , the plurality of second feature images may include a feature image 11A, a feature image 12A, and a feature image 13A. Figure 13 A and feature image 13 are the same feature image, and feature image 12A is composed of feature Figure 12 The feature image 11A is obtained by fusion of the feature image 13A after upsampling by 2 times. Figure 12 A is obtained by performing feature fusion on the feature image after upsampling 11 by a factor of 2. Figure 5 The 1*1conv module is used to fuse the feature images, and the upSample module is used to upsample the feature images by a factor of 2.
[0050] Step S223: performing latent code conversion on the plurality of second feature images of different scales to obtain the plurality of latent codes of the different scales.
[0051] In this embodiment, latent code conversion may be performed on a plurality of second feature images of different scales to obtain a plurality of latent codes of different scales.
[0052] In some implementations, multiple second feature images can be converted into multiple latent codes using a map2style module, which is internally composed of several fully connected convolutional layers. The multiple latent codes can output w1-w3 for the second feature image corresponding to 32x downsampling, responsible for style generation for posture, hairstyle, and face shape; w4-w7 for the second feature image corresponding to 16x downsampling, responsible for style generation for facial features, eye closure, and hairstyle details; and w8-w18 for the second feature image corresponding to 8x downsampling, responsible for style generation for subtle changes and color variations.
[0053] Step S230: performing affine transformation on the multiple latent codes to obtain affine transformation information corresponding to each of the multiple latent codes.
[0054] In this embodiment, an affine transformation is performed on multiple latent codes to obtain affine transformation information corresponding to each of the multiple latent codes. It should be noted that, in geometry, an affine transformation refers to a linear transformation followed by a translation in a vector space to transform it into another vector space. Affine transformations include scaling, translation, rotation, reflection, and shear mapping, which are not limited here.
[0055] Step S240: Inputting the affine transformation information corresponding to each of the multiple latent codes into the trained generative adversarial network to obtain the target image output by the trained generative adversarial network.
[0056] In some embodiments, the affine transformation information corresponding to each of the multiple latent codes is input into a trained generative adversarial network to obtain a target image output by the trained generative adversarial network. Figure 5 as well as Figure 7 , Figure 7 correspond Figure 5 The part of generating adversarial network B. Figure 5 It is the basic module of the generator of StyleGAN2-ada. StyleGAN2-ada is composed of multiple such basic modules stacked together, where c1 represents the input of a 4*4*512 constant, A represents the affine transformation of multiple latent codes to control the style of the generated image, and b i Represents random noise, B represents the converted random noise, which is used to enrich the details of the generated image, the Mod module is used to scale the weights of the convolutional layer, the Demod module is used to demodulate the weights of the convolutional layer to obtain new convolution weights to ensure numerical stability, and the UpSample layer is used to upsample the feature image to obtain a high-definition target image. For example, the input of a 4*4 feature image is upsampled layer by layer to obtain a 1024*1024 high-definition output target image.
[0057] An embodiment of the present application provides an image conversion method, compared to Figure 1 The image conversion method shown in the figure processes the image to be processed through a latent code encoding network to obtain multiple latent codes, and then inputs the multiple latent codes into a generative adversarial network to obtain a target image. Thus, image conversion is achieved by obtaining the latent code corresponding to the image to be processed. This can enhance the realism of the target image, make the target image more similar to the image to be converted, and reduce the amount of computation.
[0058] See also Figure 8 , Figure 8 The flowchart of the image conversion method provided by the embodiment of the present application is shown. In a specific embodiment, the preset neural network includes a latent code encoding network and a trained generative adversarial network. Figure 8 The process shown in FIG. 1 is described in detail. Specifically, the image conversion method may include the following steps:
[0059] Step S310: Acquire a training image set, wherein the training image set includes a training image to be converted and a target training image.
[0060] In this embodiment, a training image set is obtained. The training image set may include training images to be converted and target training images. The training images to be converted may include, but are not limited to, facial sketches, vehicle sketches, house sketches, cartoon images, and the like. The target training image is a photograph corresponding to the image to be converted. For example, a portrait of the person to be found may be determined as the training image to be converted, and a real photograph of the person to be found may be determined as the target training image.
[0061] Step S320: using the training image to be converted as an input parameter and the target training image as an output parameter, training a preset neural network to obtain the trained image conversion network.
[0062] In some embodiments, a pre-set neural network can be trained using a training image to be converted as an input parameter and a target training image as an output parameter to obtain a trained image conversion network. Furthermore, after obtaining the trained image conversion network, the accuracy of the trained image conversion network can be verified to determine whether the target training image output by the trained image conversion network based on the input training image to be converted meets preset requirements. If the target training image output by the trained image conversion network based on the input training image to be converted does not meet the preset requirements, the pre-set neural network can be trained using a new training image set, or multiple training image sets can be obtained to calibrate the trained image conversion network. This is not limited herein.
[0063] See also Figure 9 , Figure 9 Shows the application Figure 8 The following is a flow chart of step S320 of the image conversion method. Figure 9 The process shown in FIG. 1 is described in detail. Specifically, the image conversion method may include the following steps:
[0064] Step S321: Processing the training image to be converted based on the latent code encoding network to obtain multiple latent codes of different scales corresponding to the training image to be converted.
[0065] In some embodiments, see Figure 5 The training image 10 to be converted is processed according to the latent code encoding network A, and multiple latent codes at different scales corresponding to the training image 10 to be converted are obtained from the latent code encoding network A. The multiple latent codes are responsible for generating multiple styles. The latent code encoding network A includes an FPN module and a map2style module. The FPN module is used to perform feature fusion on feature images of different scales, and the map2style module is used to perform latent code conversion on feature images of different scales. Figure 5 The FPN module is used to perform feature fusion on the feature image 11A, the feature image 12A, and the feature image 13A.
[0066] In some embodiments, the latent code encoding network uses ResNet101 (residual network) as the backbone network, and the feature maps (feature image 11, feature image 12, feature image 13) of the network to be converted are downsampled 8 times, 16 times, and 32 times respectively, and the FPN module is connected to obtain feature maps of corresponding sizes (feature image 11A, feature image 12A, feature image 13A), and the map2style module obtains the 512-dimensional styleGAN2-ada latent codes w1, w2, w3, w4, w5, w6, w7, w8, w9, w10, w11, w12, w13, w14, w15, w16, w17, and w18.
[0067] Step S322: training the trained generative adversarial network based on the latent codes of different scales corresponding to the image to be converted and the target training image to obtain the trained image conversion network.
[0068] In this embodiment, an affine transformation is performed on the latent codes of different scales corresponding to the image to be converted to obtain affine transformation information corresponding to each of the multiple latent codes. The trained generative adversarial network is trained based on the affine transformation information corresponding to each of the multiple latent codes and the target training image to obtain a trained image conversion network. It should be noted that in geometry, an affine transformation refers to a linear transformation followed by a translation in a vector space to transform it into another vector space. Affine transformation changes include scaling, translation, rotation, reflection, and shear mapping, which are not limited here.
[0069] In some embodiments, the generative adversarial network can be StyleGAN2-ada, and the trained image conversion network can be expressed as Among them, G represents the generation of adversarial network StyleGAN2-ada, E represents the latent code encoding network, Represented as the average latent code of the generative adversarial network StyleGAN2-ada itself.
[0070] Step S330: Obtain influencing factors of the preset neural network.
[0071] In this embodiment, the influencing factors of the preset neural network can be obtained, and the influencing factors may include a first target test image, a second target test image, and multiple latent codes. The first target test image can be a real image corresponding to the test image to be converted, and the second target test image can be an image output by the trained image conversion network after the test image to be converted.
[0072] Step S340: Calculate the influencing factors based on a preset loss function to determine the loss values corresponding to the influencing factors.
[0073] In some embodiments, the electronic device may be pre-set and store a preset loss function, calculate the influencing factors, and determine the loss values of the influencing factors.
[0074] In some embodiments, in order to better determine the similarity between the first target test image and the second target test image to determine whether the image conversion network converges, a multidimensional similarity loss function is constructed. The multidimensional similarity loss function includes an image similarity loss function, a feature similarity loss function, a latent code regression loss function, and a facial embedding similarity loss function.
[0075] See also Figure 10 , Figure 10 Shows the application Figure 8 The following is a flow chart of step S340 of the image conversion method. Figure 10 The process shown in FIG. 1 is described in detail. Specifically, the image conversion method may include the following steps:
[0076] Step S341: Based on the image similarity loss function, the similarity between the first target test image and the second target test image is calculated to determine a first loss value, wherein the first target test image is a real image corresponding to the test image to be converted, and the second target test image is an image output by the trained image conversion network after the test image to be converted.
[0077] In this embodiment, the similarity between the first target test image and the second target test image can be calculated based on the image similarity loss function to determine the first loss value. The image similarity loss function can be L2(x) = ||y-sketch2pic(x)||2, where L2(x) represents the first loss value, sketch2pic(x) represents the second target image test image, and y represents the first target test image. The first target test image is the real image corresponding to the test image to be converted, and the second target test image is the image output by the trained image conversion network after the test image to be converted is converted. The image similarity loss function can calculate the similarity between the first target test image and the second target test image.
[0078] Step S342: Based on the feature similarity loss function, the similarity between the feature information of the first target test image and the feature information of the second target test image is calculated to determine a second loss value.
[0079] In this embodiment, the similarity between the feature information of the first target test image and the feature information of the second target test image can be calculated based on the feature similarity loss function to determine the second loss value. feature (x)=||F(y)-F(sketch2pic(x))||2, L feature (x) represents the second loss value, and F represents the AlexNet network, which is used to extract feature information of the first target test image and feature information of the second target test image. The feature similarity loss function can calculate the similarity between the feature information of the first target test image and the feature information of the second target test image.
[0080] Step S343: Based on the latent code regression loss function, calculate the similarity between the first latent code and the second latent code to determine a third loss value, wherein the first latent code is obtained by the latent code encoding network processing the test image to be converted, and the second latent code is the average latent code corresponding to the trained generative adversarial network.
[0081] In this embodiment, the similarity between the first latent code and the second latent code can be calculated based on the latent code regression loss function to determine the third loss value. L reg (x) is represented as the third loss value, E(x) is represented as the first latent code, and the first latent code is obtained by processing the test image to be converted by the latent code encoding network. The second latent code is represented by the randomly generated average latent code corresponding to the trained generative adversarial network. The latent code regression loss function can calculate the similarity between the multiple latent codes output by the latent code encoding network and the average latent code randomly generated by StyleGAN2-ada itself.
[0082] Step S344: Based on the facial embedding similarity loss function, the similarity between the facial embedding information of the first target test image and the facial embedding information of the second target test image is calculated to determine a fourth loss value.
[0083] In this embodiment, the similarity between the facial embedding information of the first target test image and the facial embedding information of the second target test image can be calculated based on the face embedding similarity loss function to determine the fourth loss value. embedding (x) = 1-<R(y),R(sketch2pic(x))> , L embedding(x) is represented as the fourth loss value, represents the face recognition network (Arcfacenetwork network), the Arcface network is used to extract the facial embedding information of the first target test image and the facial embedding information of the second target test image, L embedding (x) = 1-<R(y),R(sketch2pic(x))> This is expressed as extracting the facial embedding information of the first target test image and the facial embedding information of the second target test image, and then comparing the cosine similarity of the facial embedding information of the first target test image and the facial embedding information of the second target test image. The Faceembedding similarity loss function can calculate the similarity between the facial embedding information of the first target test image and the facial embedding information of the second target test image.
[0084] In some embodiments, the electronic device may pre-set and store a first weight corresponding to the first loss value, a second weight corresponding to the second loss value, a third weight corresponding to the third loss value, and a fourth weight corresponding to the fourth loss value, and may calculate the weight of the first loss value according to L(x)=λ1L2(x)+λ2L feature (x)+λ3L reg (x)+λ4L embedding (x) calculates the influencing factors and determines the loss value corresponding to the influencing factors, where L2(x) represents the first loss value, λ1 represents the first weight corresponding to the first loss value, and L feature (x) represents the second loss value, λ2 represents the second weight corresponding to the second loss value, L reg (x) represents the third loss value, λ3 represents the third weight corresponding to the third loss value, L embedding (x) represents the fourth loss value, and λ4 represents the fourth weight corresponding to the fourth loss value.
[0085] In some embodiments, the electronic device may pre-set and store a first weight corresponding to the first loss value, a second weight corresponding to the second loss value, a third weight corresponding to the third loss value, and a fourth weight corresponding to the fourth loss value, and may add at least two of the loss values according to the weights to obtain the loss value. For example, the loss value may be obtained according to L(x)=λ1L2(x)+λ2L feature (x) Calculate the influencing factors, determine the loss value corresponding to the influencing factors, add the first loss value and the second loss value according to the first weight and the second weight to obtain the loss value; L(x) = λ2L feature (x)+λ3L reg (x) Calculate the influencing factors, determine the loss value corresponding to the influencing factors, add the second loss value and the third loss value according to the second weight and the third weight to obtain the loss value; L(x) = λ3L reg(x)+λ4L embedding (x) Calculate the influencing factors, determine the loss values corresponding to the influencing factors, add the third loss value and the fourth loss value according to the third weight and the fourth weight to obtain the loss value; L(x) = λ1L2(x) + λ2L featur ( e x)+λ3L reg (x) Calculate the influencing factors, determine the loss values corresponding to the influencing factors, add the first loss value, the second loss value, and the third loss value according to the first weight, the second weight, and the third weight to obtain the loss value.
[0086] Step S350: Based on the loss value, the preset neural network is trained to obtain the trained image conversion network.
[0087] In this embodiment, a loss value calculated according to a preset loss function can be obtained, and based on this loss value, whether the image conversion network has converged can be determined. The preset neural network can then be trained to obtain a trained image conversion network. If the image conversion network does not converge, the image conversion network parameters need to be optimized, also using the aforementioned loss value, to obtain a more optimized image conversion network.
[0088] Step S360: Acquire the image to be converted.
[0089] Step S370: Inputting the image to be converted into a trained image conversion network to obtain a target image output by the trained image conversion network, wherein the trained image conversion network is obtained by jointly training a latent code encoding network and a trained generative adversarial network, and the latent code encoding network is used to obtain multiple latent codes of different scales corresponding to the image to be converted.
[0090] For the detailed description of steps S360 to S370 , please refer to steps S110 to S120 , which will not be repeated here.
[0091] An embodiment of the present application provides an image conversion method, compared to Figure 1 The image conversion method shown calculates the influencing factors through a preset loss function and obtains the loss value, which can optimize the image conversion network, so that the target image converted by the image conversion network is more similar to the image to be converted, thereby enhancing the realism of the target image.
[0092] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0093] See also Figure 11 , Figure 11 The module block diagram of the image conversion device provided by the embodiment of the present application is shown. The image conversion device 200 is applied to the above electronic equipment. Figure 11 The image conversion device 200 includes: a to-be-converted image acquisition module 210 and a target image acquisition module 220, wherein:
[0094] The image to be converted acquiring module 210 is configured to acquire the image to be converted.
[0095] The target image acquisition module 220 is used to input the image to be converted into a trained image conversion network and obtain the target image output by the trained image conversion network, wherein the trained image conversion network is obtained by jointly training a latent code encoding network and a trained generative adversarial network, and the latent code encoding network is used to obtain multiple latent codes of different scales corresponding to the image to be converted.
[0096] Furthermore, the target image acquisition module 220 includes: a latent code acquisition submodule, an affine change submodule, and an affine transformation information input submodule, wherein:
[0097] The latent code acquisition submodule is used to input the image to be converted into the latent code encoding network, and obtain multiple latent codes of different scales corresponding to the image to be converted output by the latent code encoding network.
[0098] The affine transformation submodule is configured to perform affine transformation on the plurality of latent codes to obtain affine transformation information corresponding to each of the plurality of latent codes.
[0099] The affine transformation information input submodule is used to input the affine transformation information corresponding to each of the multiple latent codes into the trained generative adversarial network to obtain the target image output by the trained generative adversarial network.
[0100] Furthermore, the latent code acquisition submodule includes: a first feature map acquisition unit, a second feature map acquisition unit and a latent code conversion unit, wherein:
[0101] The first feature map acquisition unit is used to input the image to be converted into the latent code encoding network, downsample the image to be converted, and obtain multiple first feature images of different scales.
[0102] The second feature image acquisition unit is used to perform feature fusion on the multiple first feature images of different scales to obtain multiple second feature images of different scales.
[0103] The latent code conversion unit is configured to perform latent code conversion on the plurality of second feature images of different scales to obtain the plurality of latent codes of the different scales.
[0104] Furthermore, the image conversion device 200 includes: a training image set acquisition module and an image conversion network training module, wherein:
[0105] The training image set acquisition module is used to acquire a training image set, wherein the training image set includes a training image to be converted and a target training image.
[0106] The image conversion network training module is used to train a preset neural network using the training image to be converted as an input parameter and the target training image as an output parameter to obtain the trained image conversion network.
[0107] Furthermore, the image conversion network training module includes: a training image processing submodule to be converted and a generative adversarial network training submodule, wherein:
[0108] The training image to be converted processing submodule is used to process the training image to be converted based on the latent code encoding network to obtain multiple latent codes of different scales corresponding to the training image to be converted.
[0109] The generative adversarial network training submodule is used to train the trained generative adversarial network based on the latent codes of different scales corresponding to the image to be converted and the target training image to obtain the trained image conversion network.
[0110] Furthermore, the image conversion network training module further includes: an influencing factor acquisition submodule, a loss value determination submodule, and a trained image conversion network acquisition submodule, wherein:
[0111] The influencing factor acquisition submodule is used to obtain the influencing factors of the preset neural network.
[0112] The loss value determination submodule is used to calculate the influencing factors based on a preset loss function to determine the loss values corresponding to the influencing factors.
[0113] The trained image conversion network acquisition submodule is used to train the preset neural network based on the loss value to obtain the trained image conversion network.
[0114] Furthermore, the loss value determination submodule includes: a first loss value determination unit, a second loss value determination unit, a third loss value determination unit, and a fourth loss value determination unit, wherein:
[0115] A first loss value determination unit is configured to calculate the similarity between a first target test image and a second target test image based on an image similarity loss function to determine a first loss value, wherein the first target test image is a real image corresponding to the test image to be converted, and the second target test image is an image output by converting the test image to be converted through the trained image conversion network.
[0116] The second loss value determining unit is configured to calculate the similarity between the feature information of the first target test image and the feature information of the second target test image based on a feature similarity loss function, and determine a second loss value.
[0117] a third loss value determining unit, configured to calculate a similarity between a first latent code and a second latent code based on a latent code regression loss function to determine a third loss value, wherein the first latent code is obtained by processing the test image to be converted by the latent code encoding network, and the second latent code is an average latent code corresponding to the trained generative adversarial network.
[0118] The fourth loss value determining unit is used to calculate the similarity between the facial embedding information of the first target test image and the facial embedding information of the second target test image based on the facial embedding similarity loss function to determine a fourth loss value.
[0119] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0120] In several embodiments provided in this application, the coupling between modules may be electrical, mechanical or other forms of coupling.
[0121] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.
[0122] See also Figure 12 , which shows a structural block diagram of an electronic device 100 provided in an embodiment of the present application. The electronic device 100 can be an electronic device capable of running applications, such as a smartphone, a tablet computer, an e-book, etc. The electronic device 100 in the present application may include one or more of the following components: a processor 110, a memory 120, and one or more applications, wherein the one or more applications may be stored in the memory 120 and configured to be executed by one or more processors 110, and the one or more programs are configured to execute the method described in the aforementioned method embodiment.
[0123] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the electronic device 100. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, and accesses data stored in the memory 120 to perform various functions and process data within the electronic device 100. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing displayed content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may be implemented separately via a communication chip.
[0124] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created by the terminal 100 during use (such as a phone book, audio and video data, chat record data), etc.
[0125] See also Figure 13 , which shows a block diagram of a computer-readable storage medium provided in an embodiment of the present application. The computer-readable medium 300 stores program code, which can be called by a processor to execute the method described in the above method embodiment.
[0126] The computer-readable storage medium 300 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium 300 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 300 has storage space for program code 310 for executing any of the method steps described above. These program codes can be read from or written to one or more computer program products. The program code 310 can be compressed, for example, in a suitable form.
[0127] In summary, the image conversion method, device, electronic device, and storage medium provided by the embodiments of the present application obtain an image to be converted, input the image to be converted into a trained image conversion network, and obtain a target image output by the trained image conversion network, wherein the trained image conversion network is obtained by jointly training a latent code encoding network and a trained generative adversarial network, and the latent code encoding network is used to obtain multiple latent codes of different scales corresponding to the image to be converted, thereby achieving image conversion by obtaining the latent code corresponding to the image to be processed, which can enhance the realism of the target image and make the target image more similar to the image to be converted.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An image conversion method, characterized in that: The method comprises: Acquire an image to be converted, wherein the image to be converted is a sketch image; Inputting the image to be converted into a latent code encoding network of a trained image conversion network, downsampling the image to be converted, and obtaining a plurality of first feature images of different scales; Using the FPN module of the latent code coding network to perform feature fusion on the multiple first feature images of different scales to obtain multiple second feature images of different scales; Using the map2style module of the latent code encoding network to perform latent code conversion on the multiple second feature images of different scales to obtain the multiple latent codes of the different scales, and the feature maps of different downsampling scales correspondingly generate latent codes of different styles; Performing affine transformation on the multiple latent codes to obtain affine transformation information corresponding to each of the multiple latent codes; The affine transformation information corresponding to each of the multiple latent codes is input into a trained generative adversarial network of the trained image conversion network to obtain a target image output by the trained generative adversarial network, wherein the trained image conversion network is obtained by jointly training a latent code encoding network and a trained generative adversarial network, and the latent code encoding network is used to obtain multiple latent codes of different scales corresponding to the image to be converted, and the target image is a photographic image.
2. The method according to claim 1, characterized in that Before obtaining the image to be converted, the method further includes: Acquire a training image set, wherein the training image set includes a training image to be converted and a target training image; The training image to be converted is used as an input parameter and the target training image is used as an output parameter to train a preset neural network to obtain the trained image conversion network.
3. The method according to claim 2, characterized in that The preset neural network includes a latent code encoding network and a trained generative adversarial network, and the training image to be converted is used as an input parameter and the target training image is used as an output parameter to train the preset neural network to obtain the trained image conversion network, including: Processing the training image to be converted based on the latent code encoding network to obtain multiple latent codes of different scales corresponding to the training image to be converted; The trained generative adversarial network is trained based on latent codes of different scales corresponding to the image to be converted and the target training image to obtain the trained image conversion network.
4. The method according to claim 2, characterized in that The method further comprises: Obtaining influencing factors of the preset neural network; Calculate the influencing factors based on a preset loss function to determine the loss values corresponding to the influencing factors; Based on the loss value, the preset neural network is trained to obtain the trained image conversion network.
5. The method according to claim 4, characterized in that The calculating the influencing factors based on a preset loss function to determine the loss values corresponding to the influencing factors includes: Based on an image similarity loss function, calculating the similarity between a first target test image and a second target test image to determine a first loss value, wherein the first target test image is a real image corresponding to the test image to be converted, and the second target test image is an image output by converting the test image to be converted through the trained image conversion network; Calculating the similarity between the feature information of the first target test image and the feature information of the second target test image based on a feature similarity loss function to determine a second loss value; calculating a similarity between a first latent code and a second latent code based on a latent code regression loss function to determine a third loss value, wherein the first latent code is obtained by processing the test image to be converted by the latent code encoding network, and the second latent code is an average latent code corresponding to the trained generative adversarial network; and / or Based on the facial embedding similarity loss function, the similarity between the facial embedding information of the first target test image and the facial embedding information of the second target test image is calculated to determine a fourth loss value.
6. The method according to claim 5, characterized in that The calculating the influencing factors based on a preset loss function to determine the loss values corresponding to the influencing factors includes: Based on L(x)=λ1L2(x)+λ2L feature (x)+λ3L reg (x)+λ4L embedding (x) Calculate the influencing factors to determine the loss value corresponding to the influencing factors, wherein L2(x) represents the first loss value, λ1 represents the first weight corresponding to the first loss value, and L feature (x) represents the second loss value, λ2 represents the second weight corresponding to the second loss value, L reg (x) represents the third loss value, λ3 represents the third weight corresponding to the third loss value, L embedding (x) represents the fourth loss value, and λ4 represents the fourth weight corresponding to the fourth loss value.
7. An image conversion device, characterized in that: The device comprises: The image to be converted acquisition module is used to acquire the image to be converted, wherein the image to be converted is a sketch image; The target image acquisition module is configured to input the image to be converted into a latent code encoding network of a trained image conversion network, downsample the image to be converted, and obtain multiple first feature images of different scales; use the FPN module of the latent code encoding network to perform feature fusion on the multiple first feature images of different scales to obtain multiple second feature images of different scales; use the map2style module of the latent code encoding network to perform latent code conversion on the multiple second feature images of different scales to obtain multiple latent codes of different scales, and feature maps of different downsampling scales generate latent codes of different styles; perform affine transformation on the multiple latent codes to obtain affine transformation information corresponding to each of the multiple latent codes; input the affine transformation information corresponding to each of the multiple latent codes into a trained generative adversarial network of the trained image conversion network to obtain a target image output by the trained generative adversarial network, wherein the trained image conversion network is obtained by jointly training the latent code encoding network and the trained generative adversarial network, and the latent code encoding network is configured to obtain multiple latent codes of different scales corresponding to the image to be converted, and the target image is a photographic image.
8. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory is coupled to the processor and stores instructions. When the instructions are executed by the processor, the processor performs the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program code, which can be called by a processor to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image multi-style conversion method based on latent variable feature generation
CN110992252A