Processing Method of Image Generator, Image Generation Method and Device
The image generator iteratively trains with an attribute discriminator to enhance attribute editing accuracy, addressing imprecise attribute conversions in existing technologies by accurately transforming image attributes while preserving subject identity and image quality.
Patent Information
- Application Number
- CN202110706137.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-06-24
AI Technical Summary
When modifying image attributes, existing image attribute conversion technology can easily lead to changes in non-target attributes such as facial age, resulting in inaccurate conversion.
The original data is mapped into an implicitly encoded vector through the image generator, and converted to the target attribute direction based on the attribute editing parameters. The target attribute loss is constructed using the image attribute discriminator, and the image generator and discriminator iteratively trained to optimize the attribute editing parameters to ensure the accuracy of image attribute conversion.
Improves the accuracy of image attribute conversion, ensuring accurate editing of target attributes without affecting other non-target attributes, such as the consistent age of facials.
Smart Images

Figure CN113822953B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and particularly to a processing method for an image generator, an image generation method, and an apparatus. Background Art
[0002] With the development of artificial intelligence, image attribute conversion technology has emerged. The image attribute conversion technology can convert the attributes of an input image, such as modifying the style of the input image, modifying the expression of a person in the input image to a smile, modifying the hair color of a person in the input image to pink, and so on. The image attribute conversion technology is widely applied in fields such as social networking and image editing, and is also applied to constructing an image sample library, and so on.
[0003] However, the current image attribute conversion technology is not mature enough, resulting in poor performance of the converted image attributes. For example, when modifying the expression of a person in the input image, the facial age of the person in the input image will also be changed at the same time, resulting in inaccurate image attribute conversion. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide a processing method for an image generator, an image generation method, and an apparatus that can improve the accuracy of image attribute conversion.
[0005] A processing method for an image generator, the method comprising:
[0006] Obtaining a sample image of a target attribute and a trained image generator;
[0007] Mapping, by the image generator, the original data for generating an image into a latent code vector;
[0008] Based on current attribute editing parameters, converting the latent code vector in the direction of the target attribute, and after obtaining a target latent code vector carrying the target attribute, generating, by the image generator, a target image corresponding to the target latent code vector;
[0009] Constructing a target attribute loss based on the correlation degrees of the target attributes corresponding to the sample image and the target image determined by a to-be-trained image attribute discriminator;
[0010] After updating the network parameters of the image attribute discriminator and the attribute editing parameters according to the target attribute loss, returning to the step of obtaining the sample image of the target attribute to continue training, and until the training ends, obtaining an image attribute converter corresponding to the target attribute according to the image generator and the attribute editing parameters obtained at the end of the training.
[0011] An apparatus for processing an image generator, the apparatus comprising:
[0012] An acquisition module, configured to acquire a sample image of a target attribute and a trained image generator;
[0013] A feature mapping module, configured to map, by using the image generator, raw data for generating an image into a latent encoding vector;
[0014] An attribute conversion module, configured to convert the latent encoding vector in a direction of the target attribute based on current attribute editing parameters, and after obtaining a target latent encoding vector carrying the target attribute, generate, by using the image generator, a target image corresponding to the target latent encoding vector;
[0015] A loss construction module, configured to construct a target attribute loss based on a correlation degree of the target attribute corresponding to each of the sample image and the target image determined by an image attribute discriminator to be trained;
[0016] A training module, configured to update network parameters of the image attribute discriminator and the attribute editing parameters according to the target attribute loss, and then return to the step of acquiring the sample image of the target attribute to continue training. When the training ends, an image attribute converter corresponding to the target attribute is obtained according to the image generator and the attribute editing parameters obtained at the end of the training corresponding to the target attribute.
[0017] In one embodiment, the feature mapping module is further configured to: initialize a latent vector space; randomly sample a latent vector from the latent vectors in the latent vector space to obtain a raw latent vector for generating an image; input the raw latent vector into a feature mapping network in the image generator; and map the raw latent vector into the latent encoding vector by using the feature mapping network.
[0018] In one embodiment, the attribute conversion module is further configured to: read the current attribute editing parameters; randomly sample an attribute conversion amplitude from an attribute conversion amplitude set to obtain an attribute conversion amplitude; and convert the latent encoding vector in a direction of the target attribute according to the current attribute editing parameters and the attribute conversion amplitude to obtain a target latent encoding vector carrying the target attribute.
[0019] In one embodiment, the attribute conversion module is further configured to: input the target latent encoding vector into a feature synthesis network in the image generator; and output, by using the feature synthesis network, a target image corresponding to the target latent encoding vector.
[0020] In one embodiment, the loss construction module is further configured to: determine, by using the image attribute discriminator, a first deviation degree of the target attribute relevance degree of the sample image relative to the target attribute relevance degree of the target image, and a second deviation degree of the target attribute relevance degree of the target image relative to the target attribute relevance degree of the sample image; and construct the target attribute loss based on the first deviation degree and the second deviation degree.
[0021] In one embodiment, the training module is further configured to: update the network parameters of the image attribute discriminator according to the target attribute loss; determine, by using the image authenticity discriminator of the image generator, the image authenticity degrees corresponding to the sample image and the target image respectively, and construct an image authenticity loss based on the image authenticity degrees corresponding to the sample image and the target image respectively; and update the attribute editing parameters according to the target loss determined by the image authenticity loss and the target attribute loss.
[0022] In one embodiment, the training module is further configured to: update the network parameters of the image attribute discriminator according to the target attribute loss; determine, by using the image identity discriminator to be trained, the identity categories corresponding to the target image and the original image corresponding to the original data respectively, and construct an identity classification loss based on the identity categories corresponding to the target image and the original image respectively; update the network parameters of the image identity discriminator according to the identity classification loss; and update the attribute editing parameters according to the target loss determined by the identity classification loss and the target attribute loss.
[0023] In one embodiment, the identity classification loss includes a first identity classification loss and a second identity classification loss; the training module is further configured to: update the network parameters of the image identity discriminator according to the first identity classification loss; and update the attribute editing parameters according to the target loss determined by the second identity classification loss and the target attribute loss.
[0024] In one embodiment, the loss construction module is further configured to: determine, by using the image authenticity discriminator of the image generator, the image authenticity degrees corresponding to the sample image and the target image respectively, and construct an image authenticity loss based on the image authenticity degrees corresponding to the sample image and the target image respectively; the training module is further configured to: update the attribute editing parameters according to the target loss determined by the identity classification loss, the image authenticity loss and the target attribute loss.
[0025] In one embodiment, the loss construction module is further configured to: determine, by the image authenticity discriminator, a third deviation degree of the image authenticity degree of the sample image relative to the image authenticity degree of the target image, and a fourth deviation degree of the image authenticity degree of the target image relative to the image authenticity degree of the sample image; and construct the image authenticity loss based on the third deviation degree and the fourth deviation degree.
[0026] In one embodiment, the feature mapping module is further configured to: input the original data into a feature mapping network in the image generator; map, by the feature mapping network, the original data into a latent coding vector; a processing device of the image generator further includes a feature synthesis module, and the feature synthesis module is configured to: output, by a feature synthesis network in the image generator, an original image corresponding to the original data according to the latent coding vector.
[0027] In one embodiment, the target attribute is a first target attribute, the sample image is a first sample image of the first target attribute, and the attribute editing parameter is a first attribute editing parameter obtained by training the model of the image generator using the first sample image; the obtaining module is further configured to: obtain a second sample image of a second target attribute, where the second target attribute and the first target attribute are non-binary attributes; the training module is further configured to: train, by the second sample image of the second target attribute and the image generator, the attribute editing parameter to determine a second attribute editing parameter corresponding to the second target attribute; and obtain an image attribute converter corresponding to the first target attribute and the second target attribute according to the image generator and the first attribute editing parameter and the second attribute editing parameter.
[0028] In one embodiment, the obtaining module is further configured to: obtain a to-be-processed image to be converted to a target attribute; the feature mapping module is further configured to: determine a latent coding vector corresponding to the to-be-processed image; the attribute conversion module is further configured to: convert, by an attribute editing parameter corresponding to the target attribute in the image attribute converter, the latent coding vector corresponding to the to-be-processed image in the direction of the target attribute to obtain a target latent coding vector corresponding to the to-be-processed image and carrying the target attribute; a processing device of the image generator further includes an image generation module, and the image generation module is further configured to: generate, by the image generator in the image attribute converter, a target image corresponding to the to-be-processed image and carrying the target attribute according to the target latent coding vector corresponding to the to-be-processed image.
[0029] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the processing method of the above image generator are implemented.
[0030] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the steps of the processing method of the above image generator are implemented.
[0031] A computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the processing method of the above image generator.
[0032] For the processing method, device, computer device and storage medium of the above image generator, the original data for generating an image is mapped into a latent encoding vector by the image generator, the latent encoding vector is transformed in the direction of the target attribute based on the current attribute editing parameter, after obtaining the target latent encoding vector carrying the target attribute, the target image corresponding to the target latent encoding vector is generated by the image generator, a target attribute loss is constructed based on the correlation degree of the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained, after updating the network parameters of the image attribute discriminator and the attribute editing parameter according to the target attribute loss, return to the step of obtaining the sample image with the target attribute to continue training. In this way, through the iterative adversarial training of the image generator and the image attribute discriminator, the attribute of the target image generated by the image generator based on the attribute editing parameter is constrained by the image attribute discriminator, so that the finally trained attribute editing parameter has an accurate target attribute editing ability, and the accuracy of image attribute conversion can be improved when performing image attribute conversion through the image generator and the trained attribute editing parameter.
[0033] An image generation method, the method includes:
[0034] Obtain a to-be-processed image to be converted to a target attribute;
[0035] Determine a latent encoding vector corresponding to the to-be-processed image;
[0036] Through the attribute editing parameter corresponding to the target attribute in the trained image attribute converter, transform the latent encoding vector in the direction of the target attribute to obtain a target latent encoding vector carrying the target attribute and corresponding to the to-be-processed image;
[0037] Among them, the attribute editing parameter corresponding to the target attribute in the image attribute converter is determined according to the target attribute loss constructed during the model training of the attribute editing parameter using the sample image of the target attribute and the trained image generator; the target attribute loss is constructed based on the degree of correlation of the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained; the target image is obtained by mapping the original data for generating an image into a latent coding vector by the image generator, then converting the latent coding vector in the direction of the target attribute based on the current attribute editing parameter to obtain a target latent coding vector carrying the target attribute, and finally generating the target image by the image generator according to the target latent coding vector.
[0038] Through the image generator in the image attribute converter, a target image corresponding to the to-be-processed image and carrying the target attribute is generated according to the target latent coding vector corresponding to the to-be-processed image.
[0039] An image generation device, the device includes:
[0040] An acquisition module, configured to acquire a to-be-processed image to be converted to a target attribute;
[0041] A feature mapping module, configured to determine a latent coding vector corresponding to the to-be-processed image;
[0042] An attribute conversion module, configured to convert the latent coding vector in the direction of the target attribute through the attribute editing parameter corresponding to the target attribute in the trained image attribute converter, to obtain a target latent coding vector carrying the target attribute and corresponding to the to-be-processed image;
[0043] Among them, the attribute editing parameter corresponding to the target attribute in the image attribute converter is determined according to the target attribute loss constructed during the model training of the attribute editing parameter using the sample image of the target attribute and the trained image generator; the target attribute loss is constructed based on the degree of correlation of the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained; the target image is obtained by mapping the original data for generating an image into a latent coding vector by the image generator, then converting the latent coding vector in the direction of the target attribute based on the current attribute editing parameter to obtain a target latent coding vector carrying the target attribute, and finally generating the target image by the image generator according to the target latent coding vector.
[0044] An image generation module, configured to generate a target image corresponding to the to-be-processed image and carrying the target attribute through the image generator in the image attribute converter according to the target latent coding vector corresponding to the to-be-processed image.
[0045] A computer device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the above image generation method are implemented.
[0046] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the above image generation method are implemented.
[0047] A computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of the above image generation method.
[0048] For the above image generation method, device, computer device and storage medium, a to-be-processed image to be converted to a target attribute is obtained, a hidden coding vector corresponding to the to-be-processed image is determined, and the hidden coding vector is converted in the direction of the target attribute through the attribute editing parameters corresponding to the target attribute in the trained image attribute converter, so as to obtain a target hidden coding vector corresponding to the to-be-processed image and carrying the target attribute. Through the image generator in the image attribute converter, a target image corresponding to the to-be-processed image and carrying the target attribute is generated according to the target hidden coding vector corresponding to the to-be-processed image. Since the attribute editing parameters trained in the embodiments of the present application have accurate target attribute editing capabilities, the accuracy of image attribute conversion can be improved. Description of the Drawings
[0049] Figure 1 It is an application environment diagram of the processing method of the image generator in an embodiment;
[0050] Figure 2 It is a flowchart of the processing method of the image generator in an embodiment;
[0051] Figure 3 It is a structural block diagram of the processing network of the image generator in an embodiment;
[0052] Figure 4 It is a schematic diagram of the attribute conversion effect in an embodiment;
[0053] Figure 5 It is a structural block diagram of the processing network of the image generator in another embodiment;
[0054] Figure 6 It is a structural block diagram of the processing network of the image generator in yet another embodiment;
[0055] Figure 7 It is a structural block diagram of the processing network of the image generator in still another embodiment;
[0056] Figure 8 Flow chart of the processing method of the image generator in another embodiment;
[0057] Figure 9 Flow chart of the image generation method in one embodiment;
[0058] Figure 10 Schematic diagram of images in different styles in one embodiment;
[0059] Figure 11 Comparison schematic diagram of the attribute conversion effect in one embodiment;
[0060] Figure 12 Comparison schematic diagram of the attribute conversion effect in another embodiment;
[0061] Figure 13 Comparison schematic diagram of the attribute conversion effect in yet another embodiment;
[0062] Figure 14 Structural block diagram of the processing device of the image generator in one embodiment;
[0063] Figure 15 Structural block diagram of the image generation device in one embodiment;
[0064] Figure 16 Structural block diagram of the image generation device in one embodiment;
[0065] Figure 17 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners
[0066] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0067] The processing method of the image generator and the image generation method provided by the embodiments of the present application relate to the technology of Artificial Intelligence (AI). Artificial Intelligence is a technology that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, sense the environment, acquire knowledge and use knowledge to obtain the best results. In other words, Artificial Intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial Intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning and decision-making.
[0068] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0069] The processing method of the image generator provided in the embodiments of this application mainly relates to the machine learning technology (Machine Learning, ML) of artificial intelligence. Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0070] For example, in the embodiments of this application, through the sample images of target attributes and the image generator, model training is performed on the attribute editing parameters, and according to the image generator and the trained attribute editing parameters, an image attribute converter corresponding to the target attribute is obtained.
[0071] The image generation method provided in the embodiments of this application mainly relates to the computer vision technology (Computer Vision, CV) of artificial intelligence. Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as target recognition and measurement, and further perform graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image generation, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0072] The image generation method provided in the embodiments of this application mainly relates to the image generation technology in the field of computer vision technology. For example, in the embodiments of this application, through the trained image attribute converter corresponding to the target attribute, the image to be processed is converted into a target image with the target attribute.
[0073] The processing method of the image generator provided by this application can be applied to the application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network. The terminal 102 can be, but is not limited to, various smart phones, tablet computers, laptop computers, desktop computers, portable wearable devices, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0074] In one embodiment, the terminal 102 obtains a sample image of a target attribute and a trained image generator, and sends the sample image and the image generator to the server 104. The server 104 maps the original data for generating an image into a latent encoding vector through the image generator, and based on the current attribute editing parameter, converts the latent encoding vector in the direction of the target attribute to obtain a target latent encoding vector carrying the target attribute. After that, through the image generator, a target image corresponding to the target latent encoding vector is generated. Based on the degree of correlation between the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained, a target attribute loss is constructed. After updating the network parameters and attribute editing parameters of the image attribute discriminator according to the target attribute loss, the step of obtaining the sample image of the target attribute is returned to continue training. Until the training ends, according to the image generator and the attribute editing parameter corresponding to the target attribute obtained at the end of training, an image attribute converter corresponding to the target attribute is obtained.
[0075] The processing method of the image generator provided by the embodiments of this application may be the processing device of the image generator provided by the embodiments of this application, or a computer device integrated with the processing device of the image generator, where the processing device of the image generator may be implemented in a hardware or software manner. The computer device may be Figure 1 the terminal 102 or the server 104 shown in
[0076] The processing method of the image generator provided by the embodiments of the present application can be applied to the training scenario of attribute editing parameters. Different attribute editing parameters correspond to different attributes. Through specific attribute editing parameters, an image can be converted into a specific attribute. An attribute is a characteristic of an image. According to the nature of the attribute, the attributes can be divided into binary attributes and non-binary attributes. Binary attributes are, for example, the binary characteristics of a person in an image, such as hair length, frowning / not frowning, wearing / not wearing glasses, opening / closing eyes, mouth opening and closing, gender, and so on; non-binary attributes are, for example, the non-binary characteristics of a person in an image, such as pupil color, facial expression, bangs style, person's posture, and so on. Another example is the image style, such as comic style, anime style, supermodel style, star style, and so on.
[0077] In one embodiment, the computer device obtains a sample image of a target attribute and a trained image generator; through the image generator, maps the original data for generating an image into a latent coding vector; based on the current attribute editing parameters, converts the latent coding vector in the direction of the target attribute to obtain a target latent coding vector carrying the target attribute, and then through the image generator, generates a target image corresponding to the target latent coding vector; based on the degree of correlation between the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained, constructs a target attribute loss; after updating the network parameters and attribute editing parameters of the image attribute discriminator according to the target attribute loss, returns to the step of obtaining the sample image of the target attribute to continue training. Until the training ends, according to the image generator and the attribute editing parameters corresponding to the target attribute obtained at the end of the training, an image attribute converter corresponding to the target attribute is obtained.
[0078] It can be understood that by using the image generator and sample images of different attributes to train the model of the attribute editing parameters, according to the image generator and the trained attribute editing parameters, an image attribute converter corresponding to more than one attribute can be obtained.
[0079] The image generation method provided by the present application can also be applied to the application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through the network.
[0080] In one embodiment, the terminal 102 obtains a to-be-processed image to be converted to a target attribute, and sends the to-be-processed image to be converted to the target attribute to the server 104. The server 104 determines a latent encoding vector corresponding to the to-be-processed image, and uses the attribute editing parameters corresponding to the target attribute in the trained image attribute converter to convert the latent encoding vector in the direction of the target attribute, obtaining a target latent encoding vector corresponding to the to-be-processed image and carrying the target attribute. Then, through the image generator in the image attribute converter, a target image corresponding to the to-be-processed image and carrying the target attribute is generated according to the target latent encoding vector corresponding to the to-be-processed image; wherein, the attribute editing parameters corresponding to the target attribute in the image attribute converter are determined according to the target attribute loss constructed when the model is trained for the attribute editing parameters using the sample image of the target attribute and the trained image generator; the target attribute loss is constructed based on the correlation degree of the target attributes corresponding to the sample image and the target image determined by the to-be-trained image attribute discriminator; the target image is generated by the image generator after mapping the original data for generating the image to a latent encoding vector, then converting the latent encoding vector in the direction of the target attribute based on the current attribute editing parameters to obtain a target latent encoding vector carrying the target attribute, and finally generating the target image according to the target latent encoding vector by the image generator.
[0081] The image generation method provided by the embodiments of the present application may be executed by the image generation device provided by the embodiments of the present application, or a computer device integrated with the image generation device, where the image generation device may be implemented in a hardware or software manner. The computer device may be Figure 1 the terminal 102 or the server 104 shown in
[0082] The image generation method provided by the embodiments of the present application can be applied to the image attribute conversion scenario. The embodiments of the present application can implement the conversion of binary attributes. For example, the gender of a person in an image can be converted from male to female; the embodiments of the present application can also implement the conversion of non-binary attributes. For example, the style of an image can be converted from an amateur style to a star style, or the hair color of a person in an image can be converted from black to brown, and at the same time, the style of the image can be converted from an amateur style to a supermodel style; the embodiments of the present application can also implement the conversion of multiple attributes. For example, the mouth of a person in an image can be converted from open to closed, the gender of the person can be converted from male to female, and at the same time, the style of the image can be converted from an amateur style to a supermodel style. The embodiments of the present application can perform attribute conversion on various types of images such as anime images, comic images, and real-person images.
[0083] In one embodiment, a computer device obtains a to-be-processed image to be converted to a target attribute; determines a latent encoding vector corresponding to the to-be-processed image; converts the latent encoding vector in the direction of the target attribute through an attribute editing parameter corresponding to the target attribute in a trained image attribute converter to obtain a target latent encoding vector corresponding to the to-be-processed image and carrying the target attribute; and generates, through an image generator in the image attribute converter, a target image corresponding to the to-be-processed image and carrying the target attribute according to the target latent encoding vector corresponding to the to-be-processed image.
[0084] In one embodiment, as Figure 2 shown, a processing method of an image generator is provided. In this embodiment, this method is mainly illustrated by applying it to the Figure 1 computer device (terminal 102 or server 104) in the above, including the following steps:
[0085] Step S202, obtain a sample image of the target attribute and a trained image generator.
[0086] In this application, the inventor designed an active learning network. Referring to Figure 3 , the active learning network may include an image generator and an image attribute discriminator. Among them, the image generator is used to generate images. During the process of the image generator generating images, the computer device uses attribute editing parameters to perform attribute editing operations, so that the target images generated by the image generator carry the attributes corresponding to the attribute editing parameters. The image attribute discriminator is used to form an adversarial framework with the image generator to constrain the attributes of the target images generated by the image generator, so that the attributes of the target images are consistent with the attributes of the sample images. In this way, through the adversarial training of the image attribute discriminator and the image generator, the attribute editing parameters learn the editing ability of the attributes carried by the sample images.
[0087] Among them, the image generator is a network structure with image generation capabilities. The image attribute discriminator is a network structure with the ability to recognize images with different attributes. The target attribute is the attribute to be learned by the method of the embodiments of the present application. An attribute is a characteristic of an image. According to the nature of the attribute, attributes can be classified into binary attributes and non-binary attributes. Binary attributes are, for example, binary characteristics of a person in an image, such as frowning / not frowning, wearing / not wearing glasses, opening / closing eyes, mouth opening and closing, hair length, gender, and so on; non-binary attributes are, for example, non-binary characteristics of a person in an image, such as pupil color, facial expression, bangs style, person's pose, etc. Another example is the image style, such as comic style, anime style, supermodel style, star style, and so on. It can be understood that binary attributes can be further divided into two attributes. For example, gender can be further divided into male and female, and non-binary attributes can be further divided into more than two attributes. For example, hair color can be further divided into black / yellow / brown, bangs style can be further divided into straight bangs / slanting bangs / no bangs, pupil color can be further divided into black / brown / blue, and so on.
[0088] In one embodiment, the image generator can adopt a general image generation model, such as a GAN (Generative Adversarial Network) model. Specifically, it can be a pre-trained model of StyleGAN, a pre-trained model of StyleGAN2, a pre-trained model of ProgressGAN, and so on.
[0089] In one embodiment, the sample image can be an image collected by an image acquisition device, an image generated by an image generation model, an image from a publicly available training set in the current machine learning field, a video frame extracted from a video, an image downloaded from a website, an image output by a terminal with painting functions, and so on. The sample image can be a real image or a virtual image such as an anime.
[0090] In one embodiment, the sample image can be a face image. The sample image can be a real face image or a virtual face image such as an anime.
[0091] In one embodiment, the computer device uses the sample images of the target attribute to iteratively and adversarially train the image generator and the image attribute discriminator to optimize the attribute editing parameters, so that the trained attribute editing parameters have the editing ability of the target attribute.
[0092] Step S204, through the image generator, map the original data for generating the image into a latent code vector.
[0093] Among them, the original data is a vector for an image generator to generate an image, such as a vector conforming to a uniform distribution, a normal distribution, or a standard normal distribution. A vector represents data in a numerical manner. The latent encoding vector is a vector obtained by feature extraction from the original data and used to describe the features of the original data.
[0094] In one embodiment, through an image generator, mapping the original data for generating an image into a latent encoding vector includes: initializing a latent vector space; randomly sampling a latent vector from the latent vector space to obtain an original latent vector for generating an image; inputting the original latent vector into a feature mapping network in the image generator; and mapping the original latent vector into a latent encoding vector through the feature mapping network.
[0095] Among them, the latent vector space is the vector space where the original data is located.
[0096] In one embodiment, a computer device randomly samples from the latent vector space to obtain an original latent vector for generating an image, and takes the original latent vector as the original data. The computer device inputs the original latent vector into a feature mapping network in the image generator, and maps the original latent vector into a latent encoding vector through the feature mapping network.
[0097] In one embodiment, the image generator may include a feature mapping network and a feature synthesis network. The steps for the computer device to generate an image through the image generator include: the computer device maps the original data into a latent encoding vector through the feature mapping network, outputs the original image corresponding to the latent encoding vector through the feature synthesis network, and the original image is the image generated by the image generator based on the original data. This process can be represented by the following formula:
[0098] w = G map (z)
[0099] x r = G syn (w)
[0100] Among them, G map represents the feature mapping network; G syn represents the feature synthesis network; z represents the original data for generating an image; w represents the encoding vector; x r represents the original image generated according to the original data.
[0101] The goal of this application is to find the target attribute direction in the latent encoding vector space of the image generator. Therefore, it is necessary to perform attribute editing operations on the latent encoding vector in the latent encoding vector space. The latent encoding vector space is the vector space where the latent encoding vector is located.
[0102] Step S206: Based on the current attribute editing parameters, transform the latent encoding vector in the direction of the target attribute. After obtaining the target latent encoding vector carrying the target attribute, use an image generator to generate a target image corresponding to the target latent encoding vector.
[0103] Among them, the attribute editing parameters are used to perform attribute editing operations on the latent encoding vector.
[0104] In one embodiment, the computer device maps the original data into a latent encoding vector through a feature mapping network, performs an attribute editing operation on the latent encoding vector in the latent encoding vector space through the attribute editing parameters, and after obtaining the target latent encoding vector carrying the target attribute, inputs the target latent encoding vector into the feature synthesis network in the image generator, and outputs a target image corresponding to the target latent encoding vector through the feature synthesis network. This process can be represented by the following formula:
[0105] x f =G syn (w + θ)
[0106] Among them, θ represents the attribute editing parameters; w + θ represents the target latent encoding vector carrying the target attribute; x f represents the target image.
[0107] Step S208: Based on the degree of correlation between the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained, construct a target attribute loss.
[0108] Among them, the degree of correlation between the target attributes is used to describe the degree of conformity to the target attributes.
[0109] In one embodiment, the computer device predicts the degree of correlation between the target attributes corresponding to the sample image and the target image respectively through the image attribute discriminator, and constructs a target attribute loss based on the difference between the degrees of correlation between the target attributes corresponding to the sample image and the target image.
[0110] Step S210: After updating the network parameters of the image attribute discriminator and the attribute editing parameters according to the target attribute loss, return to the step of obtaining the sample image of the target attribute and continue training until the training ends. According to the image generator and the attribute editing parameters corresponding to the target attribute obtained at the end of training, obtain an image attribute converter corresponding to the target attribute.
[0111] In one embodiment, the computer device updates the network parameters and the attribute editing parameters of the image attribute discriminator according to the target attribute loss, specifically by reducing the target attribute loss to update the network parameters and the attribute editing parameters of the image attribute discriminator. As the network parameters and the attribute editing parameters of the image attribute discriminator are continuously optimized, the attribute editing parameters increasingly possess the target attribute editing ability, the target images generated by the image generator based on the attribute editing parameters increasingly conform to the target attributes, and the recognition accuracy of the image attribute discriminator for images with different attributes is also increasingly high. In this way, through the iterative adversarial training between the image generator and the image attribute discriminator, the finally trained attribute editing parameters possess the accurate target attribute editing ability.
[0112] In one embodiment, the computer device constructs a first target attribute loss and a second target attribute loss based on the correlation degrees of the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained, updates the network parameters of the image attribute discriminator according to the first target attribute loss, and simultaneously updates the attribute editing parameters according to the second target attribute loss. It can be understood that a general loss function can meet the requirements for the first target attribute loss and the second target attribute loss in the embodiments of the present application, and the embodiments of the present application do not limit the type of the loss function adopted for the first target attribute loss and the second target attribute loss.
[0113] In one embodiment, when the number of training times reaches a specified number, or the change amount of the target attribute loss is less than a specified threshold, etc., the training ends.
[0114] In one embodiment, the computer device obtains an image attribute converter corresponding to the target attribute according to the image generator and the attribute editing parameters corresponding to the target attribute obtained at the end of the training. The image attribute converter can convert the attribute of the image to be processed into the target attribute. It can be understood that through the image generator, the attribute editing parameters are respectively trained using more than one sample image with different attributes. According to the image generator and the attribute editing parameters trained using the sample images with different attributes respectively, image attribute converters corresponding to more than one attribute can be obtained.
[0115] Refer to Figure 4 , Figure 4 shows the attribute conversion effect of the image attribute converter trained through the embodiments of the present application. It can be seen that for the image attribute converter trained through this embodiment, whether it is performing non-binary attribute conversion, binary attribute conversion or multi-attribute conversion on the image, it has excellent attribute conversion effects.
[0116] The training method provided in this embodiment reduces the dependence on training data. Only positive samples of the target attribute are required, and negative samples are not needed. The training method provided in this embodiment can easily achieve multi-attribute learning, such as learning of more than one binary attribute, learning of more than one non-binary attribute, and learning of more than one binary attribute and non-binary attribute, improving the applicability of the attribute editing task. The training method provided in this embodiment can reduce the entanglement between multiple attributes and improve the accuracy of image attribute conversion.
[0117] In the processing method of the above image generator, the original data for generating an image is mapped into a latent encoding vector by the image generator, the latent encoding vector is transformed in the direction of the target attribute based on the current attribute editing parameter, and after obtaining the target latent encoding vector carrying the target attribute, the image generator generates a target image corresponding to the target latent encoding vector. Based on the degree of relevance of the target attribute between the sample image determined by the image attribute discriminator to be trained and the target image, a target attribute loss is constructed. After updating the network parameters and attribute editing parameters of the image attribute discriminator according to the target attribute loss, the step of obtaining the sample image of the target attribute is returned to continue training. In this way, through the iterative adversarial training of the image generator and the image attribute discriminator, the image attribute discriminator constrains the attributes of the target image generated by the image generator based on the attribute editing parameter, so that the finally trained attribute editing parameter has an accurate target attribute editing ability. When performing image attribute conversion through the image generator and the trained attribute editing parameter, the accuracy of image attribute conversion can be improved.
[0118] In one embodiment, transforming the latent encoding vector in the direction of the target attribute based on the current attribute editing parameter to obtain a target latent encoding vector carrying the target attribute includes: reading the current attribute editing parameter; randomly sampling an attribute transformation amplitude from the set of attribute transformation amplitudes to obtain an attribute transformation amplitude; and transforming the latent encoding vector in the direction of the target attribute according to the current attribute editing parameter and the attribute transformation amplitude to obtain a target latent encoding vector carrying the target attribute.
[0119] Among them, the attribute transformation amplitude is used to control the amplitude of the change of the target attribute.
[0120] In one embodiment, the computer device maps the original data into a latent encoding vector through a feature mapping network, randomly samples an attribute transformation amplitude from the set of attribute transformation amplitudes to obtain an attribute transformation amplitude, and performs an attribute editing operation on the latent encoding vector in the latent encoding vector space through the current attribute editing parameter and the attribute transformation amplitude. After obtaining the target latent encoding vector carrying the target attribute, a target image corresponding to the target latent encoding vector is generated through a feature synthesis network. This process can be represented by the following formula:
[0121]
[0122] Among them, θ represents the attribute editing parameter; represents the target latent encoding vector carrying the target attribute; x f represents the target image; represents the attribute conversion amplitude,
[0123] In this embodiment, by participating in the training process of the attribute editing parameter with the attribute conversion amplitude, it helps to improve the training effect and training efficiency of the attribute editing parameter.
[0124] In one embodiment, based on the correlation degree of the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained, a target attribute loss is constructed, including: determining, through the image attribute discriminator, the first deviation degree of the correlation degree of the target attribute of the sample image relative to the correlation degree of the target attribute of the target image, and the second deviation degree of the correlation degree of the target attribute of the target image relative to the correlation degree of the target attribute of the sample image; constructing the target attribute loss based on the first deviation degree and the second deviation degree.
[0125] In one embodiment, the target attribute loss term in the target loss for optimizing the attribute editing parameter and the target attribute loss for updating the network parameters of the image attribute discriminator can adopt the same or different loss functions. It can be understood that a general loss function can meet the requirements for the target attribute loss term and the target attribute loss in the embodiments of the present application, and the embodiments of the present application do not limit the type of the loss function adopted by the target attribute loss term and the target attribute loss.
[0126] For example, the relative average loss function (RaHingeGAN) can be used to construct the target attribute loss term and the target attribute loss, which is represented by the following formula:
[0127]
[0128] Among them, L adv represents the target attribute loss or the target attribute loss term in the target loss; x represents the sample image; x f represents the target image; E represents taking the mean of a set of training data; x f ~Q is used to represent that the distribution of x f follows Q; x~R is used to represent that the distribution of x follows R; D adv (x f ) represents the evaluation score of the image attribute discriminator on the degree of the target image having the target attribute; D adv (x) represents the evaluation score of the image attribute discriminator on the degree of the sample image having the target attribute; Represents the magnitude of attribute conversion; f(γ) represents a scalar-to-scalar function, such as f1(γ)=ReLU(1+γ), f2(γ)=ReLU(1-γ).
[0129] Referring to the above formula, The first offset degree of the target attribute correlation degree of the sample image relative to the target attribute correlation degree of the target image may be represented, that is, the possibility that the sample image has a higher probability of having the target attribute than the target image; The second deviation degree of the target attribute correlation degree of the target image relative to the target attribute correlation degree of the sample image may be represented, that is, the possibility that the target image has a higher target attribute than the sample image.
[0130] In this embodiment, the attribute editing parameters are updated by reducing the target loss, wherein the value of the target attribute loss term is also reduced accordingly, so that the update of the attribute editing parameters takes into account the target attributes; the network parameters of the image attribute discriminator are updated by reducing the target attribute loss, so that the image attribute discriminator constrains the attributes of the target image generated by the image generator based on the attribute editing parameters, so that the target image converted based on the attribute editing parameters obtained by the final training has accurate target attributes.
[0131] In one embodiment, the network parameters and attribute editing parameters of the image attribute discriminator are updated according to the target attribute loss, including: updating the network parameters of the image attribute discriminator according to the target attribute loss; determining the image authenticity corresponding to the sample image and the target image respectively through the image authenticity discriminator of the image generator, and constructing the image authenticity loss based on the image authenticity corresponding to the sample image and the target image respectively; updating the attribute editing parameters according to the target loss determined by the image authenticity loss and the target attribute loss.
[0132] The image authenticity discriminator is used to identify the authenticity of the input image. Optionally, the discriminator provided by the image generator can be used as the image authenticity discriminator. For example, the image generator is a network structure for generating comic images, and the image authenticity discriminator can judge the possibility that the input image is a comic image.
[0133] In one embodiment, referring to Figure 5 It can be seen that the image generator, the image authenticity discriminator and the image attribute discriminator form an adversarial framework. The image attribute discriminator constrains the attributes of the target image generated by the image generator based on the attribute editing parameters. The image authenticity discriminator constrains the authenticity of the target image generated by the image generator based on the attribute editing parameters to ensure the image quality of the target image.
[0134] In one embodiment, the computer device updates the network parameters of the image attribute discriminator according to the target attribute loss, updates the network parameters of the image authenticity discriminator according to the image authenticity loss, and updates the attribute editing parameters according to the target loss determined by the image authenticity loss and the target attribute loss.
[0135] In one embodiment, the network parameters of the image generator and the image authenticity discriminator included in the image generator are not updated. The computer device updates the network parameters of the image attribute discriminator according to the target attribute loss determined by the image attribute discriminator, and at the same time determines the image authenticity loss through the image authenticity discriminator, constructs the target loss according to the image authenticity loss and the target attribute loss, and updates the attribute editing parameters according to the target loss. The training process of this embodiment can be expressed by the following formula:
[0136] ρ(θ,D attr )=argminL(θ,D attr )
[0137] Among them, ρ(θ,D attr ) represents the optimization function; L(θ) represents the loss function used to optimize the attribute editing parameters; L(D attr ) represents the loss function used to optimize the network parameters of the image attribute discriminator.
[0138] In one embodiment, the loss used to update the attribute editing parameters may include: an image authenticity loss item provided by the image authenticity discriminator and a target attribute loss item provided by the image attribute discriminator. The target loss may be expressed by the following formula:
[0139] L(θ)=λ1L adv +λ3L dis
[0140] Where L(θ) represents the loss function used to optimize the attribute editing parameters; L adv represents the target attribute loss term provided by the image attribute discriminator; L dis represents the image authenticity loss term provided by the image authenticity discriminator; λ1 and λ3 represent the weights of the target attribute loss term and the image authenticity loss term, respectively.
[0141] During the training process, by reducing L(θ) and L(D attr) value to update the attribute editing parameters and the network parameters of the image attribute discriminator. As the training progresses, the attribute editing parameters become more and more capable of target attribute editing. The target image generated by the image generator based on the attribute editing parameters becomes more and more in line with the target attributes, and the recognition accuracy of the image attribute discriminator for images with different attributes becomes higher and higher. In this way, through the iterative adversarial training of the image generator, the image attribute discriminator, and the image authenticity discriminator, the finally trained attribute editing parameters have accurate target attribute editing capabilities, and the authenticity and image quality of the target image converted based on the attribute editing parameters are also guaranteed.
[0142] In this embodiment, an adversarial framework is formed by the image generator, the image attribute discriminator, and the image authenticity discriminator to train the attribute editing parameters. The image attribute discriminator constrains the attributes of the target image generated by the image generator based on the attribute editing parameters, and the image authenticity discriminator constrains the authenticity of the target image generated by the image generator based on the attribute editing parameters, so that the finally trained attribute editing parameters have accurate target attribute editing capabilities, and the authenticity and image quality of the target image converted based on the attribute editing parameters are also guaranteed.
[0143] In one embodiment, updating the network parameters and the attribute editing parameters of the image attribute discriminator according to the target attribute loss includes: updating the network parameters of the image attribute discriminator according to the target attribute loss; determining the identity categories corresponding to the target image and the original image corresponding to the original data through the image identity discriminator to be trained, and constructing an identity classification loss based on the identity categories corresponding to the target image and the original image; updating the network parameters of the image identity discriminator according to the identity classification loss; and updating the attribute editing parameters according to the target loss determined by the identity classification loss and the target attribute loss.
[0144] Among them, the image identity discriminator is used to identify the identity category of the person in the input image.
[0145] In one embodiment, referring to Figure 6 , it can be seen that an adversarial framework is formed by the image generator, the image identity discriminator, and the image attribute discriminator. The image attribute discriminator constrains the attributes of the target image generated by the image generator based on the attribute editing parameters, and the image identity discriminator constrains the identity of the person in the target image generated by the image generator based on the attribute editing parameters to ensure that the identity of the person in the target image remains the same before and after the attribute conversion. For example, the identity of the person in the image to be processed is a certain star A. After training, the image attribute converter changes the gender and facial expression of the person in the image to be processed to obtain the target image. However, due to the constraint of the image identity discriminator during the training process of the attribute editing parameters, the identity of the person in the target image is still a certain star A.
[0146] In one embodiment, the computer device updates the network parameters of the image attribute discriminator according to the target attribute loss, updates the network parameters of the image identity discriminator according to the identity classification loss, and updates the attribute editing parameters according to the target loss determined by the identity classification loss and the target attribute loss. The training process of this embodiment can be represented by the following formula:
[0147] ρ(θ,D attr ,C id )=argminL(θ,D attr ,C id )
[0148] where ρ(θ,D attr ,C id ) represents the optimization function; L(θ) represents the loss function for optimizing the attribute editing parameters; L(D attr ) represents the loss function for optimizing the network parameters of the image attribute discriminator; L(C id ) represents the loss function for optimizing the network parameters of the image identity discriminator.
[0149] In one embodiment, the loss for updating the attribute editing parameters may include: an identity classification loss term provided by the image identity discriminator and a target attribute loss term provided by the image attribute discriminator. The target loss can be represented by the following formula:
[0150] L(θ)=λ1L adv +λ2L id
[0151] where L(θ) represents the loss function for optimizing the attribute editing parameters; L adv represents the target attribute loss term provided by the image attribute discriminator; L id represents the identity classification loss term provided by the image identity discriminator; λ1 and λ2 respectively represent the weights of the target attribute loss term and the identity classification loss term.
[0152] During the training process, by reducing L(θ), L(C id ) and L(D attr) value to update the attribute editing parameters, the network parameters of the image attribute discriminator, and the network parameters of the image identity discriminator. As the training progresses, the attribute editing parameters become increasingly capable of target attribute editing. The target images generated by the image generator based on the attribute editing parameters increasingly conform to the target attributes. The recognition accuracy of the image attribute discriminator for images with different attributes becomes higher and higher, and the recognition accuracy of the image identity discriminator for the identity categories of people in different images becomes higher and higher. In this way, through the iterative adversarial training of the image generator, the image identity discriminator, and the image attribute discriminator, the finally trained attribute editing parameters have accurate target attribute editing capabilities, and the identity of the people in the target images converted based on the attribute editing parameters is consistent with the identity of the people in the original images.
[0153] In one embodiment, the method further includes: inputting the original data into the feature mapping network in the image generator; mapping the original data into a latent coding vector through the feature mapping network; and outputting an original image corresponding to the original data according to the latent coding vector through the feature synthesis network in the image generator.
[0154] Specifically, the computer device inputs the original data into the feature mapping network in the image generator, maps the original data into a latent coding vector through the feature mapping network, and outputs an original image corresponding to the latent coding vector through the feature synthesis network. Since the latent coding vector has not undergone attribute editing operations, the identity of the people in the original image has a reference meaning. By using the image identity discriminator to constrain the identity of the people in the target images generated by the image generator based on the attribute editing parameters, the target images converted using the attribute editing parameters still retain the identity characteristics of the people in the original images.
[0155] In this embodiment, an adversarial framework is formed by the image generator, the image attribute discriminator, and the image identity discriminator to train the attribute editing parameters. The image attribute discriminator constrains the attributes of the target images generated by the image generator based on the attribute editing parameters, and the image identity discriminator constrains the identity of the people in the target images generated by the image generator based on the attribute editing parameters, so that the finally trained attribute editing parameters have accurate target attribute editing capabilities, and the identity of the people in the target images converted based on the attribute editing parameters is consistent with the identity of the people in the original images.
[0156] In one embodiment, the identity classification loss includes a first identity classification loss and a second identity classification loss; updating the network parameters of the image identity discriminator according to the identity classification loss includes: updating the network parameters of the image identity discriminator according to the first identity classification loss; updating the attribute editing parameters according to the target loss determined by the identity classification loss and the target attribute loss includes: updating the attribute editing parameters according to the target loss determined by the second identity classification loss and the target attribute loss.
[0157] In one embodiment, the identity classification loss term in the target loss and the identity classification loss for updating the network parameters of the image identity discriminator may adopt the same or different loss functions. It can be understood that a general loss function can meet the requirements for the identity classification loss term and the identity classification loss in the embodiments of the present application, and the embodiments of the present application do not limit the types of loss functions adopted for the identity classification loss term and the identity classification loss.
[0158] In one embodiment, the computer device updates the network parameters of the image identity discriminator according to the first identity classification loss, and updates the attribute editing parameters according to the target loss determined by the second identity classification loss and the target attribute loss.
[0159] For example, the cross-entropy loss function can be adopted to construct the first identity classification loss, which is represented by the following formula:
[0160] L(C id ) = φ(x r , k) + λφ(x f , k)
[0161] Where:
[0162] Where, L(C id ) represents the loss function for optimizing the network parameters of the image identity discriminator; x r represents the original image; λ is the loss weight coefficient, used to balance the weights of different generated images; x f represents the target image; x represents the target image or the original image; C id (x) j represents the predicted probability that the input image x belongs to the j-th class by the image identity discriminator; the training objective of the embodiments of the present application is to predict the identities of the people in the target image and the original image as the same class, which is represented by k here. If the predicted class of the input image x is k, then {yx} j is 1, otherwise it is 0.
[0163] In this embodiment, the network parameters of the image identity discriminator are updated by reducing the first identity classification loss, so that the image identity discriminator restricts the identity category of the people in the target image, and makes the identity of the people in the target image converted based on the finally trained attribute editing parameters consistent with the identity of the people in the original image.
[0164] For example, the cosine function can be adopted to construct the second identity classification loss, which is represented by the following formula:
[0165] L id = d cos (fr id ,f f id )
[0166] Among them, L id represents the identity classification loss term in the target loss; f r id Represents the high-dimensional features of the original image; f f id Represents the high-dimensional features of the target image.
[0167] In this embodiment, the attribute editing parameters are updated by reducing the target loss, wherein the value of the second identity classification loss is also reduced accordingly, so that the update of the attribute editing parameters takes into account the preservation of the identity features in the original image.
[0168] In one embodiment, the method further includes: determining the degree of image authenticity corresponding to each of the sample image and the target image through an image authenticity discriminator of the image generator, and constructing an image authenticity loss based on the degree of image authenticity corresponding to each of the sample image and the target image; updating attribute editing parameters according to the target loss determined by the identity classification loss and the target attribute loss, including: updating attribute editing parameters according to the target loss determined by the identity classification loss, the image authenticity loss and the target attribute loss.
[0169] Specifically, refer to Figure 7 It can be seen that the image generator forms an adversarial framework with the image identity discriminator, the image authenticity discriminator and the image attribute discriminator. The image identity discriminator constrains the attributes of the target image generated by the image generator based on the attribute editing parameters. The image authenticity discriminator constrains the authenticity of the target image generated by the image generator based on the attribute editing parameters. The image identity discriminator constrains the identity of the person in the target image generated by the image generator based on the attribute editing parameters.
[0170] In one embodiment, the computer device updates the network parameters of the image attribute discriminator according to the target attribute loss, updates the network parameters of the image authenticity discriminator according to the image authenticity loss, updates the network parameters of the image identity discriminator according to the identity classification loss, and updates the attribute editing parameters according to the target loss determined by the image authenticity loss, the identity classification loss and the target attribute loss.
[0171] In one embodiment, the network parameters of the image generator and the image authenticity discriminator included in the image generator are not updated. The computer device updates the network parameters of the image attribute discriminator according to the target attribute loss, updates the network parameters of the image identity discriminator according to the identity classification loss, and updates the attribute editing parameters according to the target loss determined by the identity classification loss and the target attribute loss. The training process of this embodiment can be expressed by the following formula:
[0172] ρ(θ, C id , D attr ) = argmin L(θ, C id , D attr )
[0173] where ρ(θ, D attr , C id ) represents the optimization function; L(θ) represents the loss function for optimizing the attribute editing parameters; L(D attr ) represents the loss function for optimizing the network parameters of the image attribute discriminator; L(C id ) represents the loss function for optimizing the network parameters of the image identity discriminator.
[0174] In one embodiment, the loss for updating the attribute editing parameters may include: an identity classification loss term provided by the image identity discriminator, an image authenticity loss term provided by the image authenticity discriminator, and a target attribute loss term provided by the image attribute discriminator. The target loss can be represented by the following formula:
[0175] L(θ) = λ1L adv + λ2L id + λ3L dis
[0176] where L(θ) represents the loss function for optimizing the attribute editing parameters; L adv represents the target attribute loss term provided by the image attribute discriminator; L id represents the identity classification loss term provided by the image identity discriminator; L dis represents the image authenticity loss term provided by the image authenticity discriminator; λ1, λ2, and λ3 respectively represent the weights of the target attribute loss term, the identity classification loss term, and the image authenticity loss term.
[0177] During the training process, by reducing L(θ), L(C id ) and L(D attr) value to update the attribute editing parameters, the network parameters of the image attribute discriminator, and the network parameters of the image identity discriminator. As the training progresses, the attribute editing parameters become more and more capable of target attribute editing. The target image generated by the image generator based on the attribute editing parameters becomes more and more in line with the target attribute. The recognition accuracy of the image attribute discriminator for images with different attributes becomes higher and higher, and the recognition accuracy of the image identity discriminator for the identity categories of people in different images becomes higher and higher. In this way, through the iterative adversarial training of the image generator with the image identity discriminator, the image authenticity discriminator, and the image attribute discriminator, the finally trained attribute editing parameters have accurate target attribute editing capabilities, and the identity of the person in the target image converted based on the attribute editing parameters is consistent with the identity of the person in the original image, and the authenticity of the target image is also guaranteed.
[0178] In one embodiment, through the image authenticity discriminator of the image generator, the image authenticity degrees corresponding to the sample image and the target image are determined respectively. Based on the image authenticity degrees corresponding to the sample image and the target image respectively, an image authenticity loss is constructed, including: through the image authenticity discriminator, determining a third deviation degree of the image authenticity degree of the sample image relative to the image authenticity degree of the target image, and a fourth deviation degree of the image authenticity degree of the target image relative to the image authenticity degree of the sample image; based on the third deviation degree and the fourth deviation degree, constructing the image authenticity loss.
[0179] Among them, the image authenticity degree can be the image authenticity level or the image falsity level. In this embodiment, the image authenticity level is taken as an example for illustration.
[0180] In one embodiment, the image authenticity loss term in the target loss and the image authenticity loss used to update the network parameters of the image authenticity discriminator can adopt the same or different loss functions. It can be understood that a general loss function can meet the requirements for the image authenticity loss term and the image authenticity loss in the embodiments of the present application. The embodiments of the present application do not limit the types of loss functions adopted for the image authenticity loss term and the image authenticity loss.
[0181] For example, the relative average loss function (RaHingeGAN) can be used to construct the image authenticity loss term, which is represented by the following formula:
[0182]
[0183] Among them, L dis represents the image authenticity loss term in the target loss; x represents the sample image; x f represents the target image; E represents taking the mean value of a group of training data; x f ~Q is used to indicate that the distribution of x f obeys Q; x~R is used to indicate that the distribution of x obeys R; Dsyn (x f ) represents the evaluation score of the authenticity degree of the target image by the image authenticity discriminator; D syn (x) represents the evaluation score of the authenticity degree of the sample image by the image authenticity discriminator; g(γ) represents a scalar-to-scalar function, for example, g1(γ) = ReLU(1 + γ), g2(γ) = ReLU(1 - γ).
[0184] Referring to the above formula, where can represent the third deviation degree of the image authenticity degree of the sample image relative to the image authenticity degree of the target image, that is, the possibility that the sample image is more authentic than the target image; g2(D syn (x f ) - E x~R (D syn (x))) can represent the fourth deviation degree of the image authenticity degree of the target image relative to the image authenticity degree of the sample image, that is, the possibility that the target image is more authentic than the sample image.
[0185] In this embodiment, the attribute editing parameter is updated by reducing the target loss, and the value of the image authenticity loss term also decreases accordingly, so that the update of the attribute editing parameter takes into account preserving the original distribution of the image, thereby ensuring the authenticity of the image.
[0186] In one embodiment, the target attribute is the first target attribute, the sample image is the first sample image of the first target attribute, and the attribute editing parameter is the first attribute editing parameter obtained by training the image generator using the first sample image. The method further includes: obtaining a second sample image of the second target attribute; training the attribute editing parameter through the second sample image of the second target attribute and the image generator to determine the second attribute editing parameter corresponding to the second target attribute; obtaining an image attribute converter corresponding to the first target attribute and the second target attribute according to the image generator and the first attribute editing parameter and the second attribute editing parameter.
[0187] Specifically, the computer device trains the attribute editing parameter respectively using sample images of more than one and different attributes through the image generator. According to the image generator and the attribute editing parameters obtained by training using sample images of different attributes respectively, an image attribute converter corresponding to more than one attribute can be obtained.
[0188] It can be understood that the attributes of more than one sample image used to train the attribute editing parameters can be binary attributes or non-binary attributes. The computer device uses the binary attribute sample images to train the model of the attribute editing parameters, and can obtain the attribute editing parameters with the binary attribute editing ability; uses the non-binary attribute sample images to train the model of the attribute editing parameters, and can obtain the attribute editing parameters with the non-binary attribute editing ability; uses the binary attribute sample images and the non-binary attribute sample images to train the model of the attribute editing parameters, and can obtain the attribute editing parameters with both the binary attribute editing ability and the non-binary attribute editing ability.
[0189] For example, the computer device uses the binary attributes "female" and "open eyes" and the non-binary attribute "supermodel style" respectively to train the model of the attribute editing parameters through the image generator. The finally trained attribute editing parameters have the attribute editing abilities of "female", "open eyes" and "supermodel style". Through the image generator and the finally trained attribute editing parameters, the input image can be converted into a female, open eyes and / or supermodel style.
[0190] In this embodiment, through the image generator, the model of the attribute editing parameters is trained respectively using more than one sample image with different attributes, and an image attribute converter corresponding to more than one attribute is obtained. This training method weakens the entanglement between attributes and can easily achieve multi-attribute learning, such as learning binary attributes and / or non-binary attributes, thereby improving the applicability of the attribute editing task.
[0191] In one embodiment, the method further includes: obtaining a to-be-processed image to be converted to a target attribute; determining a latent encoding vector corresponding to the to-be-processed image; converting the latent encoding vector corresponding to the to-be-processed image in the direction of the target attribute through the attribute editing parameters corresponding to the target attribute in the image attribute converter to obtain a target latent encoding vector corresponding to the to-be-processed image and carrying the target attribute; generating, according to the target latent encoding vector corresponding to the to-be-processed image, a target image corresponding to the to-be-processed image and carrying the target attribute through the image generator in the image attribute converter.
[0192] Among them, the to-be-processed image is an image to be subjected to attribute conversion by the method provided in the embodiments of the present application. The to-be-processed image can be an image collected by an image acquisition device, an image generated by an image generation model, a video frame extracted from a video, an image downloaded from a website, an image output by a terminal with a painting function, and so on. The to-be-processed image can be a real image or a virtual image such as an anime.
[0193] In one embodiment, the computer device obtains a to-be-processed image to be converted to a target attribute, determines the category of the to-be-processed image, and converts the to-be-processed image into a latent encoding vector according to the category of the to-be-processed image. Optionally, the categories of the images are distinguished according to the acquisition methods of the images. For example, the categories of the images may include a first category and a second category, where the first category is an image generated by an image generation model, and the second category is an image obtained by other methods other than the acquisition method of the first category. The image generation model is, for example, a GAN (Generative Adversarial Network) model.
[0194] In one embodiment, when the category of the to-be-processed image is the first category, the computer device can directly obtain the latent encoding vector corresponding to the to-be-processed image. When the category of the to-be-processed image is the second category, the computer device determines the latent encoding vector corresponding to the to-be-processed image according to the inverse mapping strategy. Among them, inverse mapping (GAN inversion) is to convert the input image into the latent encoding vector space of the pre-trained GAN model. It can be understood that a general inverse mapping strategy can meet the requirements of the embodiments of the present application for inverse mapping. Therefore, a general inverse mapping strategy can be used to determine the latent encoding vector corresponding to the to-be-processed image of the second category.
[0195] In one embodiment, after the computer device obtains the latent encoding vector corresponding to the to-be-processed image, it uses the attribute editing parameter corresponding to the target attribute in the image attribute converter to convert the latent encoding vector corresponding to the to-be-processed image in the direction of the target attribute to obtain a target latent encoding vector carrying the target attribute, and uses the image generator in the image attribute converter to generate a target image corresponding to the to-be-processed image and carrying the target attribute according to the target latent encoding vector corresponding to the to-be-processed image.
[0196] In this embodiment, since the attribute editing parameters trained in the embodiments of the present application have accurate target attribute editing capabilities, the accuracy of the target attribute in the target image can be improved.
[0197] In one embodiment, referring to Figure 8 , a processing method for an image generator is provided. The processing method for the image generator can be applied to the training scenario of attribute editing parameters and includes the following steps:
[0198] Step S802, obtain a sample image of the target attribute and a trained image generator.
[0199] Step S804, initialize the latent vector space, randomly sample from the latent vectors in the latent vector space to obtain the original latent vectors for generating images, input the original latent vectors into the feature mapping network in the image generator, and map the original latent vectors into latent encoding vectors through the feature mapping network.
[0200] Step S806: Through the feature synthesis network in the image generator, output the original image corresponding to the original data according to the latent encoding vector.
[0201] Step S808: Read the current attribute editing parameters, randomly sample the attribute conversion amplitude from the set of attribute conversion amplitudes to obtain the attribute conversion amplitude. According to the current attribute editing parameters and the attribute conversion amplitude, convert the latent encoding vector in the direction of the target attribute to obtain the target latent encoding vector carrying the target attribute.
[0202] Step S810: Input the target latent encoding vector into the feature synthesis network in the image generator, and output the target image corresponding to the target latent encoding vector through the feature synthesis network.
[0203] Step S812: Based on the correlation degree of the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained, construct the target attribute loss.
[0204] In one embodiment, through the image attribute discriminator, determine the first offset degree of the correlation degree of the target attribute of the sample image relative to the correlation degree of the target attribute of the target image, and the second offset degree of the correlation degree of the target attribute of the target image relative to the correlation degree of the target attribute of the sample image; based on the first offset degree and the second offset degree, construct the target attribute loss.
[0205] Step S814: Through the image identity discriminator to be trained, determine the identity categories corresponding to the target image and the original image corresponding to the original data respectively. Based on the identity categories corresponding to the target image and the original image respectively, construct the first identity classification loss and the second identity classification loss.
[0206] Step S816: Through the image authenticity discriminator of the image generator, determine the image authenticity degrees corresponding to the sample image and the target image respectively. Based on the image authenticity degrees corresponding to the sample image and the target image respectively, construct the image authenticity loss.
[0207] In one embodiment, through the image authenticity discriminator, determine the third offset degree of the image authenticity degree of the sample image relative to the image authenticity degree of the target image, and the fourth offset degree of the image authenticity degree of the target image relative to the image authenticity degree of the sample image; based on the third offset degree and the fourth offset degree, construct the image authenticity loss.
[0208] Step S818: Update the attribute editing parameters according to the target loss determined by the second identity classification loss, the image authenticity loss, and the target attribute loss. Update the network parameters of the image attribute discriminator according to the target attribute loss, and update the network parameters of the image identity discriminator according to the first identity classification loss.
[0209] Step S820: Return to the step of obtaining the sample image of the target attribute and continue training until the training ends. At the end of the training, according to the image generator and the attribute editing parameters corresponding to the target attribute obtained at the end of the training, obtain the image attribute converter corresponding to the target attribute.
[0210] For the above processing method of the image generator, an adversarial framework is formed by the image generator, the image attribute discriminator, the image authenticity discriminator, and the image identity discriminator to train the attribute editing parameters. The image attribute discriminator constrains the attributes of the target image generated by the image generator based on the attribute editing parameters. The image identity discriminator constrains the identity of the person in the target image generated by the image generator based on the attribute editing parameters, so that the finally trained attribute editing parameters have accurate target attribute editing capabilities, and the identity of the person in the target image converted based on the attribute editing parameters is consistent with the identity of the person in the original image. The authenticity and image quality of the target image converted based on the attribute editing parameters are also guaranteed.
[0211] In one embodiment, as Figure 9 shown, an image generation method is provided. In this embodiment, it is mainly exemplified that this method is applied to the above Figure 1 computer device (terminal 102 or server 104), including the following steps:
[0212] Step S902: Obtain the image to be processed that needs to be converted to the target attribute.
[0213] Among them, the image to be processed is the image that needs to be subjected to attribute conversion by the method provided in the embodiments of the present application. The image to be processed can be an image collected by an image acquisition device, an image generated by an image generation model, a video frame extracted from a video, an image downloaded from a website, an image output by a terminal with a painting function, and so on. The image to be processed can be a real image or a virtual image such as an anime.
[0214] It can be understood that the method of the embodiments of the present application can be applied to video editing. The computer device obtains each video frame of the video to be processed, and uses the method of the embodiments of the present application to generate, through the image attribute converter, target images corresponding to each video frame and carrying the target attribute.
[0215] In one embodiment, the computer device obtains the image to be processed that needs to be converted to more than one target attribute, and converts the image to be processed into more than one target attribute through the image attribute converter corresponding to more than one attribute.
[0216] Step S904: Determine the latent encoding vector corresponding to the image to be processed.
[0217] In one embodiment, the computer device obtains the image to be processed that needs to be converted to the target attribute, determines the category of the image to be processed, and converts the image to be processed into a latent encoding vector according to the category of the image to be processed. Optionally, the categories of images are distinguished according to the acquisition method of the images. For example, the categories of images may include a first category and a second category, where the first category is an image generated by an image generation model, and the second category is an image obtained by other methods except the acquisition method of the first category. The image generation model is, for example, a GAN (Generative Adversarial Network) model.
[0218] In one embodiment, when the category of the image to be processed is the first category, the computer device can directly obtain the latent encoding vector corresponding to the image to be processed. When the category of the image to be processed is the second category, the computer device determines the latent encoding vector corresponding to the image to be processed according to the inverse mapping strategy. Among them, inverse mapping (GAN inversion) is to convert the input image into the latent encoding vector space of the pre-trained GAN model. It can be understood that a general inverse mapping strategy can meet the requirements of the embodiments of the present application for inverse mapping. Therefore, a general inverse mapping strategy can be used to determine the latent encoding vector corresponding to the image to be processed in the second category.
[0219] Step S906: Convert the latent encoding vector in the direction of the target attribute through the attribute editing parameter corresponding to the target attribute in the trained image attribute converter, and obtain a target latent encoding vector corresponding to the image to be processed and carrying the target attribute; wherein, the attribute editing parameter corresponding to the target attribute in the image attribute converter is determined according to the target attribute loss constructed when training the model of the attribute editing parameter with the sample image of the target attribute and the trained image generator; the target attribute loss is constructed based on the correlation degree of the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained; the target image is generated by the image generator according to the target latent encoding vector after the original data for generating the image is mapped into a latent encoding vector, and then the latent encoding vector is converted in the direction of the target attribute based on the current attribute editing parameter to obtain a target latent encoding vector carrying the target attribute.
[0220] In one embodiment, after the computer device obtains the latent encoding vector corresponding to the image to be processed, it converts the latent encoding vector corresponding to the image to be processed in the direction of the target attribute through the attribute editing parameter corresponding to the target attribute in the image attribute converter, and obtains a target latent encoding vector carrying the target attribute.
[0221] Regarding the training steps of the attribute editing parameters, reference may be made to the above embodiments, and details thereof will not be elaborated in this embodiment.
[0222] Step S908: Through the image generator in the image attribute converter, generate a target image corresponding to the image to be processed and carrying the target attribute according to the target latent coding vector corresponding to the image to be processed.
[0223] In one embodiment, the computer device generates a target image corresponding to the image to be processed and carrying the target attribute through the image generator in the image attribute converter according to the target latent coding vector corresponding to the image to be processed.
[0224] In the above image generation method, obtain the image to be processed to be converted to the target attribute, determine the latent coding vector corresponding to the image to be processed, and through the attribute editing parameters corresponding to the target attribute in the trained image attribute converter, convert the latent coding vector in the direction of the target attribute to obtain a target latent coding vector corresponding to the image to be processed and carrying the target attribute. Then, through the image generator in the image attribute converter, generate a target image corresponding to the image to be processed and carrying the target attribute according to the target latent coding vector corresponding to the image to be processed. Since the attribute editing parameters trained in the embodiments of the present application have accurate target attribute editing capabilities, the accuracy of image attribute conversion can be improved.
[0225] The attribute editing parameters trained by the method of the embodiments of the present application have excellent attribute editing capabilities. The embodiments of the present application use 9 anime attributes and 7 face attributes for experiments. The 9 anime attributes are: open mouth, straight bangs, hair length, black hair, blond hair, pink hair, Itomugi-Kun style, manga style, and Cherry Blossom style. The 7 face attributes are: pose, age, gender, smile, wearing / not wearing glasses, supermodel style, and Chinese star style. For the convenience of understanding, reference is made to Figure 10 , Figure 10 which shows the Itomugi-Kun style, manga style, and Cherry Blossom style.
[0226] The embodiments of the present application are compared with the image attribute conversion methods in the prior art. Reference is made to Figure 11 , Figure 11 which shows the comparison of the attribute conversion effects of binary attributes. Among them, the first row of each group of images is the image attribute conversion effect of the embodiments of the present application, and the second row of each group of images is the image attribute conversion effect of the traditional image attribute conversion method. The left and right images in the same row respectively represent weakening attributes and strengthening attributes. For example, Figure 11In it, the traditional image attribute conversion method makes the aging attribute and the wearing glasses attribute entangled, while the embodiment of the present application specifically changes the aging attribute. It can be seen that in the binary attribute conversion, the embodiment of the present application can reduce the entanglement between attributes, specifically implement the conversion of binary attributes, and can keep the identity of the person in the image unchanged.
[0227] Referring to Figure 12 , Figure 12 shows a comparison of the attribute conversion effects of non-binary attributes. Among them, the first row of each group of images is the image attribute conversion effect of the embodiment of the present application, and the second row of each group of images is the image attribute conversion effect of the traditional image attribute conversion method. The left and right images in the same row respectively represent the weakening attribute and the strengthening attribute. For example, the traditional image attribute conversion method does not capture the unique features of the style, while the embodiment of the present application can capture the representative features of the styles of both anime and real human faces at the same time. It can be seen that in the non-binary attribute conversion, the embodiment of the present application can capture the unique features of non-binary attributes and can keep the identity of the person in the image unchanged.
[0228] Referring to Figure 13 , Figure 13 shows a comparison of the attribute conversion effects of real pictures. Among them, the last row is the image attribute conversion effect of the embodiment of the present application, and the other rows are the image attribute conversion effects of the traditional image attribute conversion method. It can be seen that when performing attribute conversion on real images, the embodiment of the present application can reduce the entanglement between attributes, specifically implement the conversion of attributes, and can keep the identity of the person in the image unchanged.
[0229] It should be understood that although Figure 2 , 8 -9, the steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 2 , 8 -9, at least a part of the steps may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily in sequence either, but can be executed alternately or in turn with at least a part of the steps or stages in other steps or other steps.
[0230] In one embodiment, as Figure 14As shown, a processing device for an image generator is provided. This device can be a software module, a hardware module, or a combination of both to form part of a computer device. Specifically, the device includes: an acquisition module 1402, a feature mapping module 1404, an attribute conversion module 1406, a loss construction module 1408, and a training module 1410, where:
[0231] The acquisition module 1402 is used to acquire a sample image of the target attribute and a pre-trained image generator;
[0232] The feature mapping module 1404 is used to map the original data for generating an image into a latent encoding vector through the image generator;
[0233] The attribute conversion module 1406 is used to convert the latent encoding vector in the direction of the target attribute based on the current attribute editing parameters. After obtaining the target latent encoding vector carrying the target attribute, it generates a target image corresponding to the target latent encoding vector through the image generator;
[0234] The loss construction module 1408 is used to construct a target attribute loss based on the degree of correlation between the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained;
[0235] The training module 1410 is used to update the network parameters of the image attribute discriminator and the attribute editing parameters according to the target attribute loss, and then return to the step of acquiring the sample image of the target attribute to continue training. When the training ends, according to the image generator and the attribute editing parameters obtained at the end of training corresponding to the target attribute, an image attribute converter corresponding to the target attribute is obtained.
[0236] In one embodiment, the feature mapping module 1404 is further used to: initialize the latent vector space; randomly sample a latent vector from the latent vectors in the latent vector space to obtain the original latent vector for generating an image; input the original latent vector into the feature mapping network in the image generator; map the original latent vector into a latent encoding vector through the feature mapping network.
[0237] In one embodiment, the attribute conversion module 1406 is further used to: read the current attribute editing parameters; randomly sample an attribute conversion amplitude from the set of attribute conversion amplitudes to obtain the attribute conversion amplitude; convert the latent encoding vector in the direction of the target attribute according to the current attribute editing parameters and the attribute conversion amplitude to obtain the target latent encoding vector carrying the target attribute.
[0238] In one embodiment, the attribute conversion module 1406 is further used to: input the target latent encoding vector into the feature synthesis network in the image generator; output a target image corresponding to the target latent encoding vector through the feature synthesis network.
[0239] In one embodiment, the loss construction module 1408 is further configured to: determine, via an image attribute discriminator, a first deviation degree of the target attribute relevance degree of the sample image relative to the target attribute relevance degree of the target image, and a second deviation degree of the target attribute relevance degree of the target image relative to the target attribute relevance degree of the sample image; and construct a target attribute loss based on the first deviation degree and the second deviation degree.
[0240] In one embodiment, the training module 1410 is further configured to: update the network parameters of the image attribute discriminator according to the target attribute loss; determine, via an image authenticity discriminator of the image generator, the image authenticity degrees corresponding to the sample image and the target image respectively, and construct an image authenticity loss based on the image authenticity degrees corresponding to the sample image and the target image respectively; and update the attribute editing parameters according to the target loss determined by the image authenticity loss and the target attribute loss.
[0241] In one embodiment, the training module 1410 is further configured to: update the network parameters of the image attribute discriminator according to the target attribute loss; determine, via a to-be-trained image identity discriminator, the identity categories corresponding to the target image and the original image corresponding to the original data respectively, and construct an identity classification loss based on the identity categories corresponding to the target image and the original image respectively; update the network parameters of the image identity discriminator according to the identity classification loss; and update the attribute editing parameters according to the target loss determined by the identity classification loss and the target attribute loss.
[0242] In one embodiment, the identity classification loss includes a first identity classification loss and a second identity classification loss; the training module 1410 is further configured to: update the network parameters of the image identity discriminator according to the first identity classification loss; and update the attribute editing parameters according to the target loss determined by the second identity classification loss and the target attribute loss.
[0243] In one embodiment, the loss construction module 1408 is further configured to: determine, via an image authenticity discriminator of the image generator, the image authenticity degrees corresponding to the sample image and the target image respectively, and construct an image authenticity loss based on the image authenticity degrees corresponding to the sample image and the target image respectively; the training module 1410 is further configured to: update the attribute editing parameters according to the target loss determined by the identity classification loss, the image authenticity loss and the target attribute loss.
[0244] In one embodiment, the loss construction module 1408 is further configured to: determine, via an image authenticity discriminator, a third deviation degree of the image authenticity degree of the sample image relative to the image authenticity degree of the target image, and a fourth deviation degree of the image authenticity degree of the target image relative to the image authenticity degree of the sample image; and construct an image authenticity loss based on the third deviation degree and the fourth deviation degree.
[0245] In one embodiment, the feature mapping module 1404 is further configured to: input the original data into the feature mapping network in the image generator; map the original data into a latent encoding vector through the feature mapping network; the processing device of the image generator further includes a feature synthesis module, and the feature synthesis module is configured to: output the original image corresponding to the original data according to the latent encoding vector through the feature synthesis network in the image generator.
[0246] In one embodiment, the target attribute is the first target attribute, the sample image is the first sample image of the first target attribute, and the attribute editing parameter is the first attribute editing parameter obtained by training the model of the image generator using the first sample image; the obtaining module 1402 is further configured to: obtain a second sample image of a second target attribute, and the second target attribute and the first target attribute are non-binary attributes; the training module 1410 is further configured to: perform model training on the attribute editing parameter through the second sample image of the second target attribute and the image generator to determine a second attribute editing parameter corresponding to the second target attribute; according to the image generator and the first attribute editing parameter and the second attribute editing parameter, obtain an image attribute converter corresponding to the first target attribute and the second target attribute.
[0247] In one embodiment, the obtaining module 1402 is further configured to: obtain a to-be-processed image to be converted to the target attribute; the feature mapping module 1404 is further configured to: determine a latent encoding vector corresponding to the to-be-processed image; the attribute conversion module 1406 is further configured to: convert the latent encoding vector corresponding to the to-be-processed image in the direction of the target attribute through the attribute editing parameter corresponding to the target attribute in the image attribute converter to obtain a target latent encoding vector corresponding to the to-be-processed image and carrying the target attribute; the processing device of the image generator further includes an image generation module, and the image generation module is further configured to: generate a target image corresponding to the to-be-processed image and carrying the target attribute according to the target latent encoding vector corresponding to the to-be-processed image through the image generator in the image attribute converter.
[0248] For the specific definition of the processing device of the image generator, reference may be made to the definition of the processing method of the image generator in the foregoing text, which will not be elaborated herein. Each module in the above-mentioned processing device of the image generator can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above-mentioned modules.
[0249] In the processing device of the above image generator, the original data for generating an image is mapped into a latent encoding vector by the image generator, and the latent encoding vector is transformed in the direction of the target attribute based on the current attribute editing parameter. After obtaining the target latent encoding vector carrying the target attribute, the image generator generates a target image corresponding to the target latent encoding vector. Based on the degree of correlation of the target attributes corresponding to the sample image and the target image determined by the image attribute discriminator to be trained, a target attribute loss is constructed. After updating the network parameters of the image attribute discriminator and the attribute editing parameter according to the target attribute loss, the step of obtaining the sample image with the target attribute is returned for continued training. In this way, through iterative adversarial training of the image generator and the image attribute discriminator, the image attribute discriminator constrains the attributes of the target image generated by the image generator based on the attribute editing parameter, so that the finally trained attribute editing parameter has an accurate target attribute editing ability. When performing image attribute conversion through the image generator and the trained attribute editing parameter, the accuracy of image attribute conversion can be improved.
[0250] In one embodiment, as Figure 15 shown, an image generation device is provided. The device can be a software module, a hardware module, or a combination of the two to form a part of a computer device. The device specifically includes: an acquisition module 1502, a feature mapping module 1504, an attribute conversion module 1506, and an image generation module 1508, where:
[0251] The acquisition module 1502 is configured to acquire a to-be-processed image to be converted to a target attribute;
[0252] The feature mapping module 1504 is configured to determine a latent encoding vector corresponding to the to-be-processed image;
[0253] The attribute conversion module 1506 is configured to transform the latent encoding vector in the direction of the target attribute by using the attribute editing parameter corresponding to the target attribute in the trained image attribute converter, and obtain a target latent encoding vector carrying the target attribute and corresponding to the to-be-processed image;
[0254] Among them, the attribute editing parameter corresponding to the target attribute in the image attribute converter is determined according to the target attribute loss constructed during the model training of the attribute editing parameter using the sample image of the target attribute and the trained image generator; the target attribute loss is constructed based on the correlation degree of the target attribute corresponding to the sample image and the target image determined by the image attribute discriminator to be trained; the target image is obtained by mapping the original data for generating the image into a latent encoding vector by the image generator, then converting the latent encoding vector in the direction of the target attribute based on the current attribute editing parameter to obtain a target latent encoding vector carrying the target attribute, and finally generating the target image by the image generator according to the target latent encoding vector.
[0255] The image generation module 1508 is configured to generate, by means of the image generator in the image attribute converter, a target image corresponding to the image to be processed and carrying the target attribute according to the target latent encoding vector corresponding to the image to be processed.
[0256] For the specific limitations of the image generation device, reference can be made to the limitations of the image generation method in the foregoing text, which will not be elaborated herein. Each module in the above image generation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in the form of hardware or be independent of it, or can be stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0257] In the above image generation device, an image to be processed that needs to be converted to the target attribute is obtained, the latent encoding vector corresponding to the image to be processed is determined, the latent encoding vector is converted in the direction of the target attribute by using the attribute editing parameter corresponding to the target attribute in the trained image attribute converter, a target latent encoding vector carrying the target attribute and corresponding to the image to be processed is obtained, and a target image corresponding to the image to be processed and carrying the target attribute is generated by means of the image generator in the image attribute converter according to the target latent encoding vector corresponding to the image to be processed. Since the attribute editing parameter obtained by training in the embodiments of the present application has an accurate target attribute editing ability, the accuracy of image attribute conversion can be improved.
[0258] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 16As shown in the figure. The computer device includes a processor, a memory, and a network interface connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the processing data and / or image generation data of the image generator. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a processing method and / or an image generation method of an image generator.
[0259] In one embodiment, a computer device is provided. The computer device can be a terminal or a face acquisition device, and its internal structure diagram can be as Figure 17 shown in the figure. The computer device includes a processor, a memory, a communication interface, and an image acquisition device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner. The wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a processing method and / or an image generation method of an image generator.
[0260] Those skilled in the art can understand that Figure 16 and Figure 17 the structures shown in the figure are only block diagrams of some structures related to the solution of this application, and do not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0261] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps in the above method embodiments.
[0262] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0263] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.
[0264] Those of ordinary skill in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include at least one of non-volatile and volatile memories. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0265] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0266] The above-described embodiments merely represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A processing method for an image generator, characterized in that, The method includes: Obtaining a sample image of a target attribute and a pre-trained image generator; Mapping, by the image generator, the original data for generating an image into a latent encoding vector; Based on current attribute editing parameters, converting the latent encoding vector in the direction of the target attribute to obtain a target latent encoding vector carrying the target attribute, and then, by the image generator, generating a target image corresponding to the target latent encoding vector; Determining, by a to-be-trained image attribute discriminator, a first deviation degree of the correlation degree of the target attribute of the sample image relative to the correlation degree of the target attribute of the target image, and a second deviation degree of the correlation degree of the target attribute of the target image relative to the correlation degree of the target attribute of the sample image; constructing a target attribute loss based on the first deviation degree and the second deviation degree; Determining, by a to-be-trained image identity discriminator, the identity categories corresponding to the target image and the original image corresponding to the original data respectively, and constructing an identity classification loss based on the identity categories corresponding to the target image and the original image respectively; Updating the network parameters of the image attribute discriminator according to the target attribute loss, updating the network parameters of the image identity discriminator according to the identity classification loss, updating the attribute editing parameters according to the target loss determined by the identity classification loss and the target attribute loss, and then returning to the step of obtaining the sample image of the target attribute to continue training. Until the training ends, according to the image generator and the attribute editing parameters obtained at the end of the training corresponding to the target attribute, an image attribute converter corresponding to the target attribute is obtained. The obtained image attribute converter is used to convert the attribute of the image to be processed into the target attribute to obtain a target image, and the identity of the person in the obtained target image remains the same before and after the attribute conversion.
2. The method according to claim 1, characterized in that The mapping, by the image generator, of the original data for generating an image into a latent encoding vector includes: Initializing a latent vector space; Randomly sampling a latent vector from the latent vectors in the latent vector space to obtain an original latent vector for generating an image; Inputting the original latent vector into a feature mapping network in the image generator; Mapping the original latent vector into the latent encoding vector by the feature mapping network.
3. The method according to claim 1, characterized in that The converting of the latent encoding vector in the direction of the target attribute based on current attribute editing parameters to obtain a target latent encoding vector carrying the target attribute includes: Reading the current attribute editing parameters; Randomly sampling an attribute conversion amplitude from an attribute conversion amplitude set to obtain an attribute conversion amplitude; Converting the latent encoding vector in the direction of the target attribute according to the current attribute editing parameters and the attribute conversion amplitude to obtain a target latent encoding vector carrying the target attribute.
4. The method according to claim 1, wherein The generating, by the image generator, of a target image corresponding to the target latent encoding vector includes: Inputting the target latent encoding vector into a feature synthesis network in the image generator; Outputting, by the feature synthesis network, a target image corresponding to the target latent encoding vector.
5. The method according to claim 1, characterized in that The identity classification loss includes a first identity classification loss and a second identity classification loss; Updating the network parameters of the image identity discriminator according to the identity classification loss includes: Updating the network parameters of the image identity discriminator according to the first identity classification loss; Updating the attribute editing parameters according to the target loss determined by the identity classification loss and the target attribute loss includes: Updating the attribute editing parameters according to the target loss determined by the second identity classification loss and the target attribute loss.
6. The method according to claim 1, wherein The method further includes: Determining the authenticity degrees of the sample image and the target image respectively through the image authenticity discriminator of the image generator, and constructing an image authenticity loss based on the authenticity degrees of the sample image and the target image respectively; Updating the attribute editing parameters according to the target loss determined by the identity classification loss and the target attribute loss includes: Updating the attribute editing parameters according to the target loss determined by the identity classification loss, the image authenticity loss and the target attribute loss.
7. The method according to claim 6, wherein Determining the authenticity degrees of the sample image and the target image respectively through the image authenticity discriminator of the image generator, and constructing an image authenticity loss based on the authenticity degrees of the sample image and the target image respectively includes: Determining a third deviation degree of the authenticity degree of the sample image relative to the authenticity degree of the target image, and a fourth deviation degree of the authenticity degree of the target image relative to the authenticity degree of the sample image through the image authenticity discriminator; Constructing the image authenticity loss based on the third deviation degree and the fourth deviation degree.
8. The method according to claim 1, wherein The method further includes: Inputting the original data into the feature mapping network in the image generator; Mapping the original data into a latent code vector through the feature mapping network; Outputting an original image corresponding to the original data according to the latent code vector through the feature synthesis network in the image generator.
9. The method according to claim 1, wherein The target attribute is a first target attribute, the sample image is a first sample image of the first target attribute, and the attribute editing parameter is a first attribute editing parameter obtained by training the image generator using the first sample image. The method further includes: Obtaining a second sample image of a second target attribute, where the second target attribute and the first target attribute are non-binary attributes; Training the attribute editing parameter through the second sample image of the second target attribute and the image generator to determine a second attribute editing parameter corresponding to the second target attribute; Obtaining an image attribute converter corresponding to the first target attribute and the second target attribute according to the image generator, the first attribute editing parameter, and the second attribute editing parameter.
10. The method according to any one of claims 1 to 8, characterized in that, The method further includes: Obtaining a to-be-processed image to be converted to a target attribute; Determining a latent code vector corresponding to the to-be-processed image; Convert the hidden encoding vector corresponding to the image to be processed in the direction of the target attribute by using the attribute editing parameter corresponding to the target attribute in the image attribute converter, so as to obtain a target hidden encoding vector corresponding to the image to be processed and carrying the target attribute; Generate a target image corresponding to the image to be processed and carrying the target attribute according to the target hidden encoding vector corresponding to the image to be processed by using the image generator in the image attribute converter.
11. An image generation method, characterized in that, The method includes: Obtain an image to be processed that needs to be converted to a target attribute; Determine the hidden encoding vector corresponding to the image to be processed; Convert the hidden encoding vector in the direction of the target attribute by using the attribute editing parameter corresponding to the target attribute in the trained image attribute converter, so as to obtain a target hidden encoding vector corresponding to the image to be processed and carrying the target attribute; Among them, the image attribute converter is used to convert the attribute of the image to be processed into the target attribute to obtain a target image, and the identity of the person in the obtained target image remains the same before and after the attribute conversion. The attribute editing parameter corresponding to the target attribute in the image attribute converter is determined according to the target attribute loss constructed when the attribute editing parameter is trained by using the sample image of the target attribute and the trained image generator. The target attribute loss is constructed based on the degree of correlation between the target attributes of the sample image and the target image determined by the image attribute discriminator to be trained. The target image is obtained by mapping the original data for generating the image into a hidden encoding vector by using the image generator, then converting the hidden encoding vector in the direction of the target attribute based on the current attribute editing parameter to obtain a target hidden encoding vector carrying the target attribute, and then generating the target image by using the image generator according to the target hidden encoding vector. Among them, the degree of correlation between the target attributes is constructed by the image attribute discriminator to be trained according to the first deviation degree of the degree of correlation between the target attributes of the sample image relative to the degree of correlation between the target attributes of the target image, and the second deviation degree of the degree of correlation between the target attributes of the target image relative to the degree of correlation between the target attributes of the sample image. When training the attribute editing parameter by using the sample image of the target attribute and the trained image generator, the identity discriminator to be trained is also used to determine the identity categories corresponding to the target image and the original image corresponding to the original data respectively. Based on the identity categories corresponding to the target image and the original image respectively, an identity classification loss is constructed, the network parameters of the image identity discriminator are updated according to the identity classification loss, and the attribute editing parameter is updated according to the target loss determined by the identity classification loss and the target attribute loss; Generate a target image corresponding to the image to be processed and carrying the target attribute according to the target hidden encoding vector corresponding to the image to be processed by using the image generator in the image attribute converter.
12. A processing device for an image generator, characterized in that, The device includes: An acquisition module, configured to acquire a sample image of a target attribute and a trained image generator; A feature mapping module, configured to map the original data for generating an image into a latent encoding vector through the image generator; An attribute conversion module, configured to convert the latent encoding vector in the direction of the target attribute based on the current attribute editing parameter. After obtaining the target latent encoding vector carrying the target attribute, generate a target image corresponding to the target latent encoding vector through the image generator; A loss construction module, configured to determine, through a to-be-trained image attribute discriminator, a first deviation degree of the target attribute relevance degree of the sample image relative to the target attribute relevance degree of the target image, and a second deviation degree of the target attribute relevance degree of the target image relative to the target attribute relevance degree of the sample image; construct a target attribute loss based on the first deviation degree and the second deviation degree; The loss construction module is further configured to determine, through a to-be-trained image identity discriminator, the identity categories corresponding to the target image and the original image corresponding to the original data respectively, and construct an identity classification loss based on the identity categories corresponding to the target image and the original image respectively; A training module, configured to update the network parameters of the image attribute discriminator according to the target attribute loss, update the network parameters of the image identity discriminator according to the identity classification loss, update the attribute editing parameter according to the target loss determined by the identity classification loss and the target attribute loss, and then return to the step of obtaining the sample image with the target attribute to continue training. Until the training ends, according to the image generator and the attribute editing parameter obtained at the end of the training corresponding to the target attribute, obtain an image attribute converter corresponding to the target attribute. The obtained image attribute converter is used to convert the attribute of the to-be-processed image into the target attribute to obtain a target image, and the identity of the person in the obtained target image remains the same before and after the attribute conversion.
13. The device according to claim 12, characterized in that, The feature mapping module is further configured to initialize the latent vector space; randomly sample a latent vector from the latent vectors in the latent vector space to obtain an original latent vector for generating an image; input the original latent vector into the feature mapping network in the image generator; map the original latent vector into the latent encoding vector through the feature mapping network.
14. The device according to claim 12, characterized in that, The attribute conversion module is further configured to read the current attribute editing parameter; randomly sample an attribute conversion amplitude from the set of attribute conversion amplitudes to obtain an attribute conversion amplitude; convert the latent encoding vector in the direction of the target attribute according to the current attribute editing parameter and the attribute conversion amplitude to obtain a target latent encoding vector carrying the target attribute.
15. The device according to claim 12, characterized in that, The attribute conversion module is further configured to input the target latent encoding vector into the feature synthesis network in the image generator; output a target image corresponding to the target latent encoding vector through the feature synthesis network.
16. The device according to claim 12, characterized in that, The identity classification loss includes a first identity classification loss and a second identity classification loss; the training module is further configured to update the network parameters of the image identity discriminator according to the first identity classification loss; and update the attribute editing parameters according to the target loss determined by the second identity classification loss and the target attribute loss.
17. The device according to claim 12, characterized in that, The loss construction module is further configured to determine the authenticity degrees of the sample image and the target image respectively through the image authenticity discriminator of the image generator, and construct an image authenticity loss based on the authenticity degrees of the sample image and the target image respectively; the training module is further configured to update the attribute editing parameters according to the target loss determined by the identity classification loss, the image authenticity loss and the target attribute loss.
18. The device according to claim 17, characterized in that, The loss construction module is further configured to determine a third deviation degree of the authenticity degree of the sample image relative to the authenticity degree of the target image and a fourth deviation degree of the authenticity degree of the target image relative to the authenticity degree of the sample image through the image authenticity discriminator. Based on the third deviation degree and the fourth deviation degree, the image authenticity loss is constructed.
19. The device according to claim 12, characterized in that, The feature mapping module is further configured to input the original data into the feature mapping network in the image generator; map the original data into a latent coding vector through the feature mapping network; and output an original image corresponding to the original data according to the latent coding vector through the feature synthesis network in the image generator.
20. The device according to claim 12, wherein The target attribute is a first target attribute, the sample image is a first sample image of the first target attribute, and the attribute editing parameter is a first attribute editing parameter obtained by training the model of the image generator using the first sample image. The acquisition module is further configured to acquire a second sample image of a second target attribute, and the second target attribute and the first target attribute are non-binary attributes. The training module is further configured to perform model training on the attribute editing parameters through the second sample image of the second target attribute and the image generator, and determine a second attribute editing parameter corresponding to the second target attribute. According to the image generator, the first attribute editing parameter, and the second attribute editing parameter, an image attribute converter corresponding to the first target attribute and the second target attribute is obtained.
21. The device according to any one of claims 12 to 19, characterized in that, The acquisition module is further configured to acquire a to-be-processed image to be converted to a target attribute; the feature mapping module is further configured to determine a latent coding vector corresponding to the to-be-processed image; the attribute conversion module is further configured to convert the latent coding vector corresponding to the to-be-processed image in the direction of the target attribute through the attribute editing parameter corresponding to the target attribute in the image attribute converter, and obtain a target latent coding vector carrying the target attribute and corresponding to the to-be-processed image. The device further includes an image generation module, which is configured to generate a target image corresponding to the image to be processed and carrying the target attribute through an image generator in the image attribute converter according to a target latent coding vector corresponding to the image to be processed.
22. An image generation device, characterized in that, The device includes: An acquisition module, configured to acquire an image to be processed that is to be converted to a target attribute; A feature mapping module, configured to determine a latent coding vector corresponding to the image to be processed; An attribute conversion module, configured to convert the latent coding vector in the direction of the target attribute through attribute editing parameters corresponding to the target attribute in a trained image attribute converter, so as to obtain a target latent coding vector corresponding to the image to be processed and carrying the target attribute; Wherein, the image attribute converter is configured to convert the attribute of the image to be processed into the target attribute to obtain a target image, and the identity of the person in the obtained target image remains the same before and after the attribute conversion. The attribute editing parameters corresponding to the target attribute in the image attribute converter are determined according to a target attribute loss constructed when training the attribute editing parameters by using a sample image of the target attribute and a trained image generator. The target attribute loss is constructed based on the degree of correlation of the target attributes corresponding to the sample image and the target image determined by a to-be-trained image attribute discriminator. The target image is obtained by mapping the original data for generating the image into a latent coding vector through the image generator, then converting the latent coding vector in the direction of the target attribute based on the current attribute editing parameters to obtain a target latent coding vector carrying the target attribute, and finally generating the target image by the image generator according to the target latent coding vector. Wherein, the degree of correlation of the target attributes is constructed by the to-be-trained image attribute discriminator according to a first deviation degree of the degree of correlation of the target attributes of the sample image relative to the degree of correlation of the target attributes of the target image, and a second deviation degree of the degree of correlation of the target attributes of the target image relative to the degree of correlation of the target attributes of the sample image. When training the attribute editing parameters by using the sample image of the target attribute and the trained image generator, the identity categories corresponding to the target image and the original image corresponding to the original data are also determined through a to-be-trained image identity discriminator. Based on the identity categories corresponding to the target image and the original image, an identity classification loss is constructed, the network parameters of the image identity discriminator are updated according to the identity classification loss, and the attribute editing parameters are updated according to the target loss determined by the identity classification loss and the target attribute loss; An image generation module, configured to generate a target image corresponding to the image to be processed and carrying the target attribute through an image generator in the image attribute converter according to a target latent coding vector corresponding to the image to be processed.
23. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
24. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.
25. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 11.