An image generation method, device, equipment and storage medium
By extracting person features from style images and noise images, the method generates accurate and versatile style and realism images without additional training, addressing the inconsistency in existing models.
Patent Information
- Application Number
- CN202510013100.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-01-06
AI Technical Summary
In the prior art, independent training of the first diffusion model and the second diffusion model leads to the low accuracy of generating realistic images and stylized images of the same character, and the model is poorly scalable. Repeated training is required for each new style or character, which is high in training.
By obtaining pure noise images and stylized images of the target characters, extracting the character characteristics of the target characters, and removing noise information that does not meet the correlation conditions from the pure noise images, generating target realistic images, combining the image features of multiple sample realistic images, and using a de-stylized model for iterative training, realizing zero-shot to generate realistic images of the same character.
It improves the accuracy of generating realistic images and stylized images of the same character, enhances the scalability and generalization ability of the model, avoids resource consumption caused by repeated training, and improves the accuracy of training samples.
Smart Images

Figure CN119540389B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this application relate to the field of artificial intelligence technology, and in particular, to an image generation method, apparatus, device, and storage medium. Background Art
[0002] The human stylization technology is a technology that can convert real human images (also known as realistic images) into stylized images, and can be used in various fields such as artistic creation, video entertainment, virtual reality, and games. The human stylization technology is mainly achieved by training a human stylization model. Before training the human stylization model, a large number of training samples need to be prepared, and each training sample is a realistic image and a stylized image of the same person.
[0003] In order to obtain a realistic image and a stylized image of the same person, the related technology first obtains a pre-trained latent diffusion model (Latent Diffusion Model, abbreviated as LDM). The latent diffusion model can generate an image with the same meaning as the input text information according to the input text information. Then, multiple stylized images are used to fine-tune and train the latent diffusion model to obtain a first diffusion model. The first diffusion model can generate a stylized image with the same meaning as the input text information according to the input text information. Multiple realistic images are used to fine-tune and train the latent diffusion model to obtain a second diffusion model. The second diffusion model can generate a realistic image with the same meaning as the input text information according to the input text information. Finally, taking the description text of a person as the input, the stylized image of this person is generated through the first diffusion model; similarly, taking the description text of this person as the input, the realistic image of this person is generated through the second diffusion model, so as to obtain a realistic image and a stylized image of the same person.
[0004] Since the first diffusion model and the second diffusion model are two models independently trained, the degrees of understanding of the text information by these two models are different. In this case, for the same person, the human features in the stylized image generated by the first diffusion model may be quite different from the human features in the realistic image generated by the second diffusion model, that is, the realistic image and the stylized image do not correspond to the same person. Therefore, the accuracy of generating a realistic image and a stylized image of the same person by using the above scheme is relatively low. Summary of the Invention
[0005] Embodiments of this application provide an image generation method, apparatus, device, and storage medium, which are used to improve the accuracy of generating a realistic image and a stylized image of the same person.
[0006] On the one hand, embodiments of this application provide an image generation method, and the method includes:
[0007] Obtain a pure noise image and a stylized image of the target person;
[0008] Extract the personal characteristics of the target person from the stylized image of the target person;
[0009] Extract target noise information from the pure noise image that does not meet the correlation condition with the image characteristics of the realistic type image and the personal characteristics of the target person, and the image characteristics of the realistic type image are learned from multiple sample realistic images;
[0010] Remove the target noise information from the pure noise image to obtain the target realistic image of the target person.
[0011] On the one hand, an embodiment of the present application provides an image generation device, and the device includes:
[0012] An acquisition module, configured to acquire a pure noise image and a stylized image of a target person;
[0013] An extraction module, configured to extract the personal characteristics of the target person from the stylized image of the target person;
[0014] A prediction module, configured to extract target noise information from the pure noise image that does not meet the correlation condition with the image characteristics of the realistic type image and the personal characteristics of the target person, and the image characteristics of the realistic type image are learned from multiple sample realistic images;
[0015] A denoising module, configured to remove the target noise information from the pure noise image to obtain the target realistic image of the target person.
[0016] Optionally, the extraction module is specifically configured to:
[0017] Perform the following operations through a trained destylization model:
[0018] Add noise to the stylized image of the target person according to a preset noise addition mode to obtain a target noisy image, and the preset noise addition mode is obtained by training the destylization model;
[0019] Extract the personal characteristics of the target person from the target noisy image.
[0020] Optionally, the preset noise addition mode includes: a plurality of noise addition time steps and first noises respectively corresponding to the plurality of noise addition time steps; the intensities of the first noises respectively corresponding to the plurality of noise addition time steps are different;
[0021] The extraction module is specifically configured to:
[0022] During the multiple noise-adding time steps, iteratively add noise to the stylized image to obtain the target noisy image; wherein, each round of iterative noise addition includes the following steps:
[0023] Obtain the first noise of one noise-adding time step corresponding to the current round of iterative noise addition;
[0024] Use the first noise of the one noise-adding time step to add noise to the stylized image after the previous round of iterative noise addition to obtain the stylized image after the current round of iterative noise addition.
[0025] Optionally, the prediction module is specifically configured to:
[0026] Perform the following operations through the trained destylization model:
[0027] Concatenate the image features of the realistic-type image and the person features of the target person to obtain target concatenated features;
[0028] Use the target concatenated features to predict the probability that each pixel in the pure noise image contains noise information;
[0029] Based on the probability that each pixel in the pure noise image contains noise information, obtain the noise distribution in the pure noise image;
[0030] Based on the noise distribution in the pure noise image, extract the target noise information in the pure noise image.
[0031] Optionally, the target noise information includes: the second noise corresponding to each of the multiple denoising time steps; the intensities of the second noises corresponding to each of the multiple denoising time steps are different;
[0032] The denoising module is specifically configured to:
[0033] Perform the following operations through the trained destylization model:
[0034] During the multiple denoising time steps, iteratively denoise the pure noise image to obtain the target realistic image of the target person; wherein, each round of iterative denoising includes the following steps:
[0035] Obtain the second noise of one denoising time step corresponding to the current round of iterative denoising;
[0036] Remove the second noise of the one denoising time step from the pure noise image after the previous round of iterative denoising to obtain the pure noise image after the current round of iterative denoising.
[0037] Optionally, the image features of the realistic-type image include: sub-image features corresponding to each of the multiple denoising time steps;
[0038] The denoising module is specifically configured to:
[0039] Concatenate the person characteristics of the target person and the sub-image characteristics corresponding to the one denoising time step to obtain the sub-concatenated characteristics corresponding to the one denoising time step;
[0040] Based on the sub-concatenated characteristics corresponding to the one denoising time step, predict the probability that each pixel in the pure noise image after denoising in the previous iteration contains noise information;
[0041] Based on the probability that each pixel in the pure noise image after denoising in the previous iteration contains noise information, obtain the noise distribution in the pure noise image after denoising in the previous iteration;
[0042] Based on the noise distribution in the pure noise image after denoising in the previous iteration, obtain the second noise corresponding to the one denoising time step.
[0043] Optionally, it further includes a module training module;
[0044] The module training module is specifically configured to:
[0045] Obtain a sample set containing multiple training samples, and each training sample includes: a sample style image and a sample realistic image of the same sample person;
[0046] Use the sample set to perform multiple rounds of iterative training on the de-stylization model to be trained until the training termination condition is met, and obtain the trained de-stylization model. Wherein, each round of iterative training process includes the following steps:
[0047] Through the de-stylization model used in this round, convert each training sample input in this round of iteration to obtain the predicted realistic image of each training sample respectively;
[0048] According to the obtained predicted realistic images and the corresponding sample realistic images, determine the model loss value of the de-stylization model used in this round, and after adjusting the parameters of the de-stylization model used in this round according to the model loss value, use the de-stylization model with adjusted parameters to enter the next round of iterative training.
[0049] Optionally, the module training module is specifically configured to:
[0050] For each of the training samples, perform the following operations respectively:
[0051] Add noise to the sample style image in a training sample to obtain a first noise-added image; and, add noise to the sample realistic image in the one training sample to obtain a second noise-added image;
[0052] Extract the personal features of the sample person from the first noise-added image; and, extract the image features of the realistic type image from the second noise-added image;
[0053] Based on the personal features of the sample person and the image features of the realistic type image, obtain the predicted realistic image of the sample person.
[0054] Optionally, the module training module is specifically configured to:
[0055] Perform splicing based on the personal features of the sample person and the image features of the realistic type image to obtain sample splicing features;
[0056] Based on the sample splicing features, obtain the predicted noise information in the second noise-added image;
[0057] Remove the predicted noise information from the second noise-added image to obtain the predicted realistic image of the sample person.
[0058] On the one hand, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above image generation method is implemented, including:
[0059] The processor obtains a pure noise image and a stylized image of the target person;
[0060] The processor extracts the personal features of the target person from the stylized image of the target person; and extracts the target noise information from the pure noise image that does not meet the correlation condition with the image features of the realistic type image and the personal features of the target person, where the image features of the realistic type image are learned from multiple sample realistic images and stored in the memory.
[0061] The processor removes the target noise information from the pure noise image to obtain the target realistic image of the target person.
[0062] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is enabled to execute the steps of the above image generation method.
[0063] On the one hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is enabled to execute the steps of the above image generation method.
[0064] In the embodiments of the present application, taking the personal characteristics of the target person in the stylized image as reference information, and combining the image characteristics of the realistic type images learned from multiple sample realistic images, a target realistic image of the target person is generated, which ensures that the obtained stylized image and the target realistic image correspond to the same person, thereby improving the accuracy of generating realistic images and stylized images of the same person.
[0065] Secondly, extracting the personal characteristics of the target person from the stylized image to guide the generation of the target realistic image of the target person does not require extracting the style characteristics of the stylized image. That is, the process of the de-stylization model generating a realistic image of a person has nothing to do with the style of the input stylized image. Therefore, for any style or any person, a corresponding realistic image can be generated, realizing zero-shot generation of realistic images and stylized images of the same person, improving the scalability and generalization ability of the de-stylization model, and also avoiding the resource consumption caused by repeated training.
[0066] In addition, the technical solution of the present application ensures that the obtained stylized image and the realistic image correspond to the same person. Therefore, when using the obtained stylized image and realistic image as training samples for subsequent model training, the accuracy of the training samples is correspondingly improved. Since the de-stylization model of the present application has zero-shot ability, without repeating the training of the de-stylization model, a large number of diverse training samples can also be generated through the de-stylization model for subsequent model training, thereby improving the performance of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0068] Figure 1 It is a schematic structural diagram of a system architecture provided by an embodiment of the present application;
[0069] Figure 2 It is a schematic diagram of an application scenario provided by an embodiment of the present application Figure 1 ;
[0070] Figure 3 It is a schematic diagram of an application scenario provided by an embodiment of the present application Figure 2 ;
[0071] Figure 4 It is a schematic structural diagram of a de-stylization model provided by an embodiment of the present application;
[0072] Figure 5Flow schematic of a method for training a de-stylization model provided by an embodiment of the present application Figure 1 ;
[0073] Figure 6 Style schematic diagram of a style image provided by an embodiment of the present application;
[0074] Figure 7 Schematic diagram of a Cross-Attention layer provided by an embodiment of the present application;
[0075] Figure 8 Flow schematic diagram of a fine-tuning training method for a UNet network provided by an embodiment of the present application;
[0076] Figure 9 Flow schematic diagram of a method for feature splicing of a denoising network and a reference network provided by an embodiment of the present application;
[0077] Figure 10 Flow schematic of an image generation method provided by an embodiment of the present application Figure 1 ;
[0078] Figure 11A Flow schematic of a method for training a de-stylization model provided by an embodiment of the present application Figure 2 ;
[0079] Figure 11B Flow schematic of an image generation method provided by an embodiment of the present application Figure 2 ;
[0080] Figure 11C Structural schematic diagram of an image generation device provided by an embodiment of the present application;
[0081] Figure 12 Structural schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0082] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0083] For the convenience of understanding, the terms involved in the embodiments of the present invention will be explained below.
[0084] The embodiments of the present application relate to artificial intelligence (AI) technology and are mainly designed based on computer vision (CV) technology in artificial intelligence technology.
[0085] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems that can perceive the environment, acquire knowledge, and use knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0086] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0087] Computer vision technology is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as target recognition and measurement in machine vision, and further performing image processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in various visual fields can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0088] Diffusion models refer to models that perform image drawing based on the text or images input by users. Diffusion models include two main processes: forward diffusion and reverse diffusion. In the forward diffusion stage, the image is gradually contaminated by the introduced noise until the image becomes completely random noise. In the reverse process, a series of Markov chains are used to gradually remove the predicted noise at each time step to recover the data from the Gaussian noise. Due to the ability of diffusion models to retain the semantic structure of the data, they have been proven not to be affected by mode collapse.
[0089] The Latent Diffusion Model (LDM) is a deep learning model for image generation. Its core idea is to generate images by performing a diffusion process in the latent space. LDM decomposes the generation task into a conversion process from noise to data, enabling the model to efficiently generate high-quality images.
[0090] zero-shot: A learning method aimed at enabling the model to recognize, classify, or understand categories or concepts not seen during training. This learning method is crucial for solving the common long-tail distribution problems in the real world, that is, for some rare or unknown class samples, traditional supervised learning methods may be difficult to handle. Specifically, the model is trained using a large amount of training set data, and without further training, the generalization ability of the model is utilized to transfer the model's capabilities to the test set. There is no intersection between the training set categories and the test set categories, and since no further training is required, the training cost is significantly reduced.
[0091] The LoRA model (Low-Rank Adaptation of Large Language Models) is a fine-tuning method for large language models. By introducing low-rank matrices, it reduces the number of parameters and the fine-tuning cost while maintaining the model performance. The core idea of LoRA is to freeze the pre-trained model weight parameters and only train the parameters of the newly added network layers. This method is achieved by introducing low-rank matrices, which have a small number of parameters, thus significantly reducing the fine-tuning cost. Specifically, LoRA adapts to specific tasks by injecting low-rank matrices into large pre-trained language models, retaining the generalization ability of the original model while significantly reducing the computational resources and data volume required for fine-tuning.
[0092] Unet Network: A denoising network mainly composed of two structures: Residual Network (ResNet) and Transformer (a neural network model based on self-attention mechanism). ResNet is mainly composed of convolutional networks and is responsible for encoding image features. Transformer is mainly composed of multiple Attention layers, and the calculation formula of the Attention layer is shown in formula (1):
[0093]
[0094] Among them, Q is the query matrix, K is the key matrix, and V is the value matrix.
[0095] The Attention layer is responsible for encoding image features and introduced conditional features (such as text, image, and audio features). The Attention layer in Transformer mainly includes two forms: Self-Attention and Cross-Attention. When Q, K, and V are the same (for example, all are image features), the Attention layer is in the form of Self-Attention and is responsible for calculating the relationship between image features. When Q is an image feature and K and V are other conditional features, the Attention layer is in the form of Cross-Attention and is responsible for calculating the relationship between image features and other features.
[0096] Realistic image: An image that is closer to the real world. For example, images taken by mobile phones or cameras.
[0097] Stylized image: An image that leans towards artistic expression. For example, watercolor-style images, pixel-style images, line-style images, cartoon-rendered style images, etc.
[0098] Stylized data pair: The realistic image and stylized image of the same person.
[0099] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common ones include smart home, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, Artificial Intelligence Generated Content (AIGC for short), conversational interaction, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0100] The solution provided by the embodiments of this application mainly relates to the application of artificial intelligence technology in the scenario of image generation. In the specific process, the personal characteristics of the target person are extracted from the stylized image of the target person. Then, the noise information irrelevant to the image characteristics of the realistic image and the personal characteristics of the target person in the pure noise image is predicted. Next, the noise information is removed from the pure noise image to obtain the target realistic image of the target person.
[0101] Specifically, in the embodiments of this application, the image generation process can be divided into two parts, including a training part and an application part. Among them, the training part involves the technical field of machine learning. In the training part, the embodiments of this application use stylized data pairs (that is, the realistic image and the stylized image of the same person) as training data, perform iterative training on the de-stylization model to be trained, and continuously adjust the model parameters through an optimization algorithm until the model converges, so that the trained de-stylization model has the ability to generate the realistic image of any person with zero-shot.
[0102] In the application part, the de-stylization model trained in the training part is used to convert the stylized image of a person input during actual use to obtain the realistic image of this person.
[0103] In addition, it should be noted that the artificial neural network model in the embodiments of this application can be trained online or offline, and no specific limitation is made here. In this article, an example of offline training is used for illustration.
[0104] It can be understood that in the specific implementation of this application, relevant data such as the stylized image of a person and the realistic image of a person are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0105] In practical applications, in order to implement the person stylization technology, it is often necessary to prepare a large amount of stylized data pairs to train the person stylization model. However, it is relatively simple to obtain the stylized images of multiple individuals separately, or to obtain the real-person images (i.e., realistic images) of multiple individuals separately. However, it is relatively difficult to obtain a large number of stylized data pairs (the realistic image and the stylized image of the same person).
[0106] To obtain realistic images and stylized images of the same person, related technologies first obtain a pre-trained latent diffusion model. The latent diffusion model can generate images with the same meaning as the input text information according to the input text information. Then, the Lora fine-tuning scheme is used to fine-tune and train the latent diffusion model based on multiple stylized images to obtain a first diffusion model. The first diffusion model can generate stylized images with the same meaning as the input text information according to the input text information. The Lora fine-tuning scheme is used to fine-tune and train the latent diffusion model based on multiple realistic images to obtain a second diffusion model. The second diffusion model can generate realistic images with the same meaning as the input text information according to the input text information. Finally, taking the description text of a person as the input, a stylized image of the person is generated through the first diffusion model; similarly, taking the description text of the person as the input, a realistic image of the person is generated through the second diffusion model, so as to obtain realistic images and stylized images of the same person.
[0107] Since the first diffusion model and the second diffusion model are two models independently trained, the degrees of understanding of the text information by these two models are different. In this case, for the same person, the characteristics of the person in the stylized image generated by the first diffusion model may be quite different from the characteristics of the person in the realistic image generated by the second diffusion model, that is, the realistic image and the stylized image do not correspond to the same person. Therefore, the accuracy of generating realistic images and stylized images of the same person by using the above scheme is relatively low.
[0108] If the first diffusion model and the second diffusion model are fused into the same model, at this time, since the positions of introducing additional parameter matrices in the LDM model are the same during the process of separately training the first diffusion model and the second diffusion model by using the Lora fine-tuning scheme, there is a problem of position conflict when fusing the first diffusion model and the second diffusion model. In this case, only the fusion weights of the first diffusion model and the second diffusion model can be adjusted. Then, the first diffusion model generates a stylized image separately according to the fusion weight, and the second diffusion model generates a realistic image separately according to the fusion weight. This scheme is essentially still that the first diffusion model and the second diffusion model independently generate stylized images and realistic images. Therefore, the accuracy of this scheme in generating realistic images and stylized images of the same person is also relatively low.
[0109] Secondly, the scalability of the first diffusion model and the second diffusion model is poor; specifically, the first diffusion model and the second diffusion model can only generate images with the same person or style as the training samples, and cannot expand styles or people that have not been trained. In this way, each time a new style or person is added, fine-tuning training needs to be carried out again, and each time fine-tuning training requires dozens or even hundreds of training samples and takes dozens of minutes, resulting in high training costs and large training time overheads.
[0110] In view of this, the embodiments of the present application provide an image generation method. In this method, a pure noise image and a stylized image of a target person are used as the input of a de-stylization model. The de-stylization model extracts the personal features of the target person from the stylized image of the target person. Then, target noise information that does not meet the correlation condition with the image features of the realistic type image and the personal features of the target person is extracted from the pure noise image. After that, the target noise information is removed from the pure noise image to obtain the target realistic image of the target person.
[0111] That is to say, in the embodiments of the present application, taking the personal features of the target person in the stylized image as reference information and combining the image features of the realistic type image learned from multiple sample realistic images, the target realistic image of the target person is generated. This ensures that the obtained stylized image and the target realistic image correspond to the same person, thereby improving the accuracy of generating the realistic image and the stylized image of the same person.
[0112] Secondly, extracting the personal features of the target person from the stylized image to guide the generation of the target realistic image of the target person does not require extracting the style features of the stylized image. That is, the process of the de-stylization model generating the realistic image of a person has nothing to do with the style of the input stylized image. Therefore, for any style or any person, the corresponding realistic image can be generated, realizing zero-shot generation of the realistic image and the stylized image of the same person, improving the scalability and generalization ability of the de-stylization model, and also avoiding the resource consumption caused by repeated training.
[0113] In addition, the technical solution of the present application ensures that the obtained stylized image and the realistic image correspond to the same person. Therefore, when the obtained stylized image and the realistic image are used as training samples for subsequent model training, the accuracy of the training samples is correspondingly improved. Since the de-stylization model of the present application has the zero-shot ability, without repeating the training of the de-stylization model, a large number of diverse training samples can also be generated through the de-stylization model for subsequent model training, thereby improving the performance of model training.
[0114] Next, a simple introduction is made to the system architecture diagram applicable to the technical solution of the embodiments of the present application. It should be noted that the system architecture diagram introduced below is only used to illustrate the embodiments of the present application and is not a limitation.
[0115] Refer to Figure 1 , which is a system architecture diagram applicable to the embodiments of the present application. The system architecture at least includes a terminal device 101 and a server 102. The number of terminal devices 101 can be one or more, and the number of servers 102 can also be one or more. The present application does not make specific limitations on the number of terminal devices 101 and servers 102.
[0116] The terminal device 101 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, smart home appliances, a smart voice interaction device, a smart vehicle-mounted device, etc., but is not limited thereto.
[0117] The server 102 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, but is not limited thereto.
[0118] It should be noted that the method in the embodiments of the present application can be executed alone by the terminal device 101 or the server 102, or can be jointly executed by the terminal device 101 and the server 102.
[0119] When executed alone by the terminal device 101 or the server 102, the de-stylization model is deployed on the terminal device 101 or the server 102. When executed alone by the terminal device 101 or the server 102, the model training and application processes can both be independently implemented by the terminal device 101 or the server 102. For example, on the terminal device 101, the de-stylization model can be trained using the collected training data. Correspondingly, after training, the terminal device 101 can use the trained de-stylization model to generate realistic images of people, or the above process can also be executed by the server 102.
[0120] When jointly executed by the server 102 and the terminal device 101, the de-stylization model can be trained by the server 102 and then the trained de-stylization model can be deployed to the terminal device 101 to generate realistic images of people by the terminal device 101. Or, part of the model training or application process is implemented by the terminal device 101 and part is implemented by the server 102, and the two cooperate to implement the model training or application process. In actual applications, specific configurations can be made according to the situation, and the present application does not make specific limitations here.
[0121] Among them, both the server 102 and the terminal device 101 may include one or more processors, memories, and an interactive I / O interface, etc. In addition, the server 102 may also be configured with a database, which can be used to store model parameters of the de-styling model, etc. Among them, the memory of the server 102 and the terminal device 101 may also store program instructions required to be executed respectively in the image generation method provided by the embodiments of the present application. When these program instructions are executed by the processor, they can be used to implement the image generation process provided by the embodiments of the present application.
[0122] It should be noted that when the image generation method provided by the embodiments of the present application is executed independently by the server 102 or the terminal device 101, the system architecture of the present application may also only include a single device of the server 102 or the terminal device 101, or it can also be considered that the server 102 and the terminal device 101 are the same device. Of course, in actual applications, when the image generation method provided by the embodiments of the present application is executed jointly by the server 102 and the terminal device 101, the server 102 and the terminal device 101 may also be the same device, that is, the server 102 and the terminal device 101 may be different functional modules of the same device, or virtual devices virtualized by the same physical device.
[0123] In some embodiments, the user can initiate the image generation process by providing a stylized image of the target person through the terminal device 101. The server 102 can then receive the stylized image of the target person provided by the user, and then adopt the image generation method of the embodiments of the present application to generate a target realistic image of the target person based on the pure noise image and the stylized image of the target person, and return it to the terminal device 101 for presentation.
[0124] In the embodiments of the present application, the terminal device 101 and the server 102 can be directly or indirectly communicatively connected through one or more networks. The network can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network or a Wireless-Fidelity (WIFI) network. Of course, it can also be other possible networks, and the embodiments of the present application do not limit this.
[0125] The following briefly introduces some application scenarios applicable to the technical solutions of the embodiments of the present application. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application rather than to limit them. In the specific implementation process, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.
[0126] The solution provided in the embodiments of the present application can be applied to image generation in various application scenarios. For example, training data augmentation scenarios, game scenarios, Extended Reality (XR) scenarios, etc. This solution can also be used as a basic technology in various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, audio and video, and other scenarios.
[0127] The following provides an exemplary description of the application scenarios applicable to the solution provided in the present application.
[0128] Application scenario 1: Training data augmentation scenario.
[0129] Refer to Figure 2 , and set the target application as an application with the function of generating personalized character avatars. The target application realizes the function of generating personalized character avatars through a deployed stylized model. Before training the stylized model, a large number of stylized data pairs (i.e., realistic images and stylized images of the same person) need to be prepared. The stylized data pairs are obtained in the following specific way:
[0130] The terminal device 1 sends a set of stylized images of multiple people to the server, and the server is deployed with a trained destylized model. In the server, for each stylized image in the set of stylized images (such as the stylized image of person A), a pure noise image that satisfies the standard normal distribution and the stylized image of person A are input into the destylized model. Through the destylized model, the person characteristics of person A are extracted from the stylized image of person A; and the noise information that is irrelevant to the image characteristics of the realistic type image and the person characteristics of person A in the pure noise image is predicted. Then, the noise information is removed from the pure noise image to obtain the realistic image of person A. A training sample is obtained based on the realistic image of person A and the stylized image of person A.
[0131] It should be noted that by adjusting the pure noise image input into the destylized model, different realistic images of person A can be obtained, and then different training samples can be constructed.
[0132] The server can directly use the obtained multiple training samples to train and obtain the stylized model, and then deploy the trained stylized model on the server or the terminal device 2. The terminal device 1 and the terminal device 2 can be the same device or different devices.
[0133] The server can also send the obtained multiple training samples to the terminal device 2. The terminal device 2 uses the obtained multiple training samples to train and obtain the stylized model, and then deploys the trained stylized model on the server or the terminal device 2.
[0134] In the application stage, it is assumed that the trained stylized model is deployed on the server. When the user uses the target application to generate a personalized portrait, the target application is triggered on the terminal device 2 to display an input interface, and the user's realistic image (such as an ID photo) is submitted on the input interface. The terminal device 2 sends the user's realistic image to the server. The server inputs the user's realistic image into the trained stylized model to obtain the user's personalized portrait, and returns the user's personalized portrait to the terminal device 2. The terminal device 2 displays the user's personalized portrait in the conversion result interface of the target application.
[0135] Application scenario two: game scenario.
[0136] See Figure 3 , in order to make the game characters more realistic and thus improve the game experience, the stylized image of the game character can be converted into a realistic image of the game character through the de-stylization model of the present application.
[0137] Specifically, on the terminal device or the server, the collected pure noise image that satisfies the standard normal distribution and the stylized image of game character B are input into the de-stylization model. The de-stylization model extracts the character features of game character B from the stylized image of game character B; and predicts the noise information in the pure noise image that is irrelevant to the image features of the realistic type image and the character features of game character B. Then, the noise information is removed from the pure noise image to obtain the realistic image of game character B.
[0138] It should be noted that the embodiments of the present application are not limited to the several application scenarios exemplified above, and may also be other application scenarios. In this regard, the present application does not make specific limitations.
[0139] In the embodiments of the present application, the image generation process can be divided into two parts, including a training part and an application part. The following will specifically introduce the training part and the application part respectively.
[0140] To more clearly introduce the content of the training part and the application part of the present application, first introduce the de-stylization model structure involved in both the training part and the application part. See Figure 4 , which is a schematic structural diagram of a de-stylization model provided by the embodiments of the present application. The de-stylization model includes:
[0141] A first encoding module 401, a first noise addition module 402, a reference network 403, a second encoding module 404, a second noise addition module 405, a denoising network 406, and a decoding module 407.
[0142] The first encoding module 401 is used to map the stylized image of a person from the pixel space to the latent space. The first noise-adding module 402 is used to add noise to the stylized image in the latent space, where the added noise is noise that satisfies the standard normal distribution, such as Gaussian noise.
[0143] The second encoding module 404 is used to map the realistic image of a person from the pixel space to the latent space. The second noise-adding module 405 is used to add noise to the realistic image in the latent space.
[0144] The reference network 403 is used to extract the person features from the stylized image after adding noise and transfer the person features to the denoising network 406. The reference network 403 and the denoising network 406 have the same network structure.
[0145] The denoising network 406 is used to denoise the input noisy image with the person features as reference information to obtain a denoising result.
[0146] The decoding module 407 is used to map the denoising result from the latent space to the pixel space to obtain the realistic image of the person.
[0147] The following specifically introduces the training process of the de-stylization model in combination with the structure of the de-stylization model. See Figure 5 , including the following steps:
[0148] Step 501, obtain a sample set containing multiple training samples, and each training sample includes: a sample stylized image and a sample realistic image of the same sample person.
[0149] In this application, the characteristics representing style are abstracted into two aspects: structure and texture. Based on this, representative styles in each aspect are selected to achieve the purpose of expanding to other different styles. In terms of structure, it mainly focuses on 2D, 3D, and facial features, etc.; in terms of texture, it mainly focuses on light, color, and material, etc.
[0150] For example, see Figure 6 , which are 8 different styles provided by the embodiments of this application, as well as the corresponding characteristics and example diagrams. Among them, the 8 different styles are 3D real scene cartoon style, ink painting style, 3D cartoon style, ice sculpture style, 2D anime style, single-line drawing style, 2D comic style, and magic style.
[0151] In practical applications, use a text-to-image model (such as the Flux1 dev model) to generate a specific resolution (such as Stylized images of characters in ( ). Specifically, the input text template includes 3 adjustable placeholders, namely placeholder Token1, placeholder Token2, and placeholder Token3. At placeholder Token1, the gender can be selected, for example, {"female", "male"}; at placeholder Token2, the age can be selected, for example, {5, 10, 20, 30, 40, 50, 60, 70}; at placeholder Token3, different styles can be selected, for example, 8 different styles. Under this text template, by changing the noise random number initialized as the input of the model, stylized images of multiple age groups, different genders, and different styles are generated.
[0152] After generating stylized images of multiple age groups, different genders, and different styles, select stylized images with consistent styles and high quality. 80 stylized images can be retained for each style. These 80 stylized images are a combination of 2 genders, 8 age groups, and 5 stylized images, that is, for each of the 2 genders, the corresponding stylized images of that gender at 8 age groups are retained, and 5 stylized images are retained for each age group.
[0153] Next, obtain the large text-to-image model (StableDiffusion XL, SDXL) trained with publicly available text-image pairs of data. The SDXL model itself has the ability to generate realistic images of specific characters. For example, the training data for training the SDXL model contains a large amount of information about celebrities and stars. Therefore, the SDXL model itself has the ability to generate realistic images of celebrities and stars.
[0154] On this basis, fine-tune the pre-trained SDXL model with multiple stylized images obtained from the above text-to-image model to obtain a Lora-style model capable of generating stylized images of specific characters (such as celebrities and stars).
[0155] Specifically, the SDXL model includes a UNet network for denoising. During the fine-tuning training process, additional trainable parameters are introduced into the linear layer of Cross-Attention in the UNet network. Specifically, additional trainable parameters are introduced into the linear projection part of QKV in the Cross-Attention layer, and the linear part of the subsequent Feed Forward Networks (FFN).
[0156] To reduce the number of parameters required for fine-tuning, the number of trainable parameters can be reduced by matrix decomposition. Specifically, refer to Figure 7 , x represents the input of the Cross-Attention layer, the input includes matrix Q from the image and matrices K and V from the text information, and h represents the output of the Cross-Attention layer. represents the original parameters after pre-training of the Cross-Attention layer, represents the additional trainable parameters (i.e., the inserted layer), and are matrices of the same size.
[0157] To minimize the number of parameters of the newly added trainable parameters as much as possible and improve efficiency, can be decomposed into the product of two matrices B and A. For example, if is matrix, then matrix B can be matrix, and matrix A is matrix, where r can be much smaller than d. Theoretically, the smaller the rank r of the product between matrices B and A, the smaller the number of parameters of the inserted layer.
[0158] At this time, after introducing additional trainable parameters in the Cross-Attention layer, in the Cross-Attention layer, through the Attention operation, under the condition of text information, the encoded result of the image is calculated, and the specific calculation formula is shown in the following formula (2):
[0159]
[0160] During the fine-tuning training process, refer to Figure 8 and freeze the pre-trained model weight parameters in the SDXL model (i.e., ), and use the multiple stylized images obtained above to train the additional trainable parameters (i.e., the inserted layer) newly added in the UNet network to obtain the Lora-style model. Since the Cross-Attention in the UNet network plays an important role in the SDXL model, even without fine-tuning the parameters of the entire model, only fine-tuning the newly added parameter matrix of this part is sufficient to achieve good performance.
[0161] Since the fine-tuning training only adjusts the parameter information of the inserted layer and the originally pre-trained model weight parameters in the SDXL model remain unchanged, the trained Lora-style model is essentially a fine-tuned inserted layer (i.e., a Lora module) mounted on the basis of the pre-trained SDXL model.
[0162] In this way, after the fine-tuning training is completed, using the text template and the noise random number as the input of the LoRA style model, stylized images of specific characters (such as famous celebrities) can be generated. For example, the input text template includes two placeholders, namely placeholder Token1 and placeholder Token2. Among them, at placeholder Token1, it is a required option to fill in the name of a famous celebrity; at placeholder Token2, it is an optional option with 8 different styles to choose from. When no style is specified, realistic images of the corresponding famous celebrity are generated. Under this text template, by changing the noise random number of the initial input of the model, stylized images of different styles of the specified famous celebrity can be generated.
[0163] In addition, using the text template and the noise random number as the input of the pre-trained SDXL model (that is, the LoRA style model without the LoRA module), realistic images of specific characters (such as famous celebrities) are obtained. Finally, by combining the LoRA style model and the pre-trained SDXL model, stylized images and realistic images of the same specific character are obtained.
[0164] In the embodiments of the present application, in the case that the SDXL model itself has the ability to generate realistic images of specific characters, the SDXL model is fine-tuned to obtain a LoRA style model that can generate stylized images of specific characters, avoiding the fusion of multiple models and solving the problem of conflicts in the fusion of character features and styles. Secondly, when obtaining the LoRA style model by means of fine-tuning training, only a small amount of training data (more than a dozen images) is required to stably generate a fixed style. Compared with directly using the text-to-image of the basic model, the cost of image generation is significantly reduced, and the style consistency is also improved.
[0165] In some embodiments, in order to further improve the consistency of the generated stylized images with the character features of the corresponding specific characters (such as famous celebrities), the present application screens the stylized images and realistic images of specific characters.
[0166] Specifically, the stylized image is encoded into a first feature vector with multiple dimensions (such as 512 dimensions), and the realistic image is encoded into a second feature vector with multiple dimensions (such as 512 dimensions). Then, the cosine similarity between the first feature vector and the second feature vector is calculated. This cosine similarity is the face similarity between the face in the stylized image and the face in the realistic image. The cosine value range is [-1, 1], and after linear transformation, the cosine value range is [0, 1]. Among them, the greater the cosine similarity, the greater the probability that the stylized image and the realistic image correspond to the same character. The specific calculation process is shown in formula (3):
[0167]
[0168] Among them, Represents the facial similarity between the stylized image X and the realistic image Y.
[0169] After obtaining the facial similarity of each group of stylized images and realistic images using the above formula (3), retain the stylized images and realistic images with facial similarity greater than a preset threshold (such as 0.8) as training samples.
[0170] Step 502: Use the sample set to perform multiple rounds of iterative training on the de-stylization model to be trained until the training termination condition is met, and obtain the trained de-stylization model.
[0171] Specifically, by continuously optimizing the model parameters, the model can meet the accuracy requirements. When this round of iteration is the first iteration, the model used in this round is the initial model; when this round of iteration is not the first iteration, the model used in this round is the model after parameter adjustment in the previous round. In each round of iterative training, some or all of the training samples can be extracted from the sample set and input into the model used in this round. For example, a random selection method can be used, or the sample set can be pre-divided into batches in advance, and each time a batch of training samples is input, and each training is based on the input training samples, that is, the model used in this round can perform forward inference of the model on each training sample.
[0172] During the training process of the de-stylization model, since each round of iterative training process is similar, the following takes one iteration process as an example for introduction. Among them, in each round of iterative training process, the following steps are executed:
[0173] Step 5021: Use the de-stylization model used in this round to transform each training sample input in this round of iteration, and obtain the predicted realistic image of each training sample respectively.
[0174] In some embodiments, for each training sample, the following operations are respectively performed:
[0175] Add noise to the sample style image in a training sample to obtain a first noise-added image; and add noise to the sample realistic image in the training sample to obtain a second noise-added image. Extract the character features of the sample character from the first noise-added image; and extract the image features of the realistic type image from the second noise-added image. Then, based on the character features of the sample character and the image features of the realistic type image, obtain the predicted realistic image of the sample character.
[0176] Specifically, the sample style image is mapped from the pixel space to the latent space through the first encoding module, and then the sample style image is added noise in the latent space through the first noise-adding module to obtain the first noise-added image.
[0177] In practical applications, the first noise addition module iteratively adds noise to the sample style image within multiple noise addition time steps to obtain the first noise-added image. The intensities of the noises corresponding to the multiple noise addition time steps can be the same or different (for example, the intensity of the noise decreases successively within the multiple noise addition time steps). The noise can be Gaussian noise that satisfies the standard normal distribution. Each round of iterative noise addition includes the following steps:
[0178] Obtain the noise of a noise addition time step corresponding to the current round of iterative noise addition. Use the noise of this noise addition time step to add noise to the sample style image after the previous round of iterative noise addition to obtain the sample style image after the current round of iterative noise addition.
[0179] It should be noted that the first round of noise addition is to add noise to the input sample style image. The sample style image after the last round of noise addition is the first noise-added image, and the first noise-added image is a pure noise image.
[0180] In addition, the sample realistic image is mapped from the pixel space to the latent space through the second encoding module, and then the sample realistic image is added noise in the latent space through the second noise addition module to obtain the second noise-added image.
[0181] In practical applications, the second noise addition module iteratively adds noise to the sample realistic image within multiple noise addition time steps to obtain the second noise-added image. The intensities of the noises corresponding to the multiple noise addition time steps can be the same or different (for example, the intensity of the noise decreases successively within the multiple noise addition time steps). The noise can be Gaussian noise that satisfies the standard normal distribution. Each round of iterative noise addition includes the following steps:
[0182] Obtain the noise of a noise addition time step corresponding to the current round of iterative noise addition. Use the noise of this noise addition time step to add noise to the sample realistic image after the previous round of iterative noise addition to obtain the sample realistic image after the current round of iterative noise addition.
[0183] It should be noted that the first round of noise addition is to add noise to the input sample realistic image. The sample realistic image after the last round of noise addition is the second noise-added image, and the second noise-added image is a pure noise image. The number of noise addition time steps corresponding to the first noise addition module and the second noise addition module is the same. The intensities of the noises of each noise addition time step used by the first noise addition module and the second noise addition module can be the same or different.
[0184] The reference network extracts the person features of the sample person from the first noisy image and transmits the person features of the sample person to the denoising network. The denoising network extracts the image features of the realistic type image from the second noisy image and denoises the second noisy image based on the person features of the sample person and the image features of the realistic type image to obtain a sample denoising result. The sample denoising result is mapped from the latent space to the pixel space through a decoding module to obtain the predicted realistic image of the sample person.
[0185] In the embodiments of the present application, the sample style image and the sample realistic image of the same sample person are used as training samples to perform multiple rounds of iterative training on the de-styling model, so that the de-styling model learns to generate a realistic image of the corresponding person based on the person features in the sample style image, thereby obtaining the ability to generate a realistic image of the same person in a zero-shot manner, and improving the scalability and generalization ability of the model.
[0186] In some embodiments, the person features of the sample person and the image features of the realistic type image are spliced to obtain sample splicing features. Based on the sample splicing features, the predicted noise information in the second noisy image is obtained. The predicted noise information is removed from the second noisy image to obtain the predicted realistic image of the sample person.
[0187] Specifically, the denoising network iteratively denoises the second noisy image in multiple denoising time steps to obtain a sample denoising result. The intensities of the noises corresponding to the multiple denoising time steps can be the same or different (for example, the intensity of the noise increases sequentially in multiple denoising time steps). The noise can be Gaussian noise that satisfies the standard normal distribution. Each round of iterative denoising includes the following steps:
[0188] Obtain the noise of a denoising time step corresponding to the current round of iterative denoising, and remove the noise of this denoising time step from the second noisy image after the previous round of iterative denoising to obtain the second noisy image after the current round of iterative denoising.
[0189] In practical applications, in order to generate a realistic image consistent with the person features of the stylized image, in addition to using the denoising network for denoising, a reference network (ReferenceNet) is introduced, and the network structure of the reference network is the same as that of the denoising network.
[0190] In each denoising time step, the person features of the sample person are extracted from the first noisy image through the reference network; the sub-image features of the realistic type image are extracted from the second noisy image after the previous round of iterative denoising through the denoising network. Then, the person features of the sample person generated by the reference network are transmitted to the denoising network and spliced with the sub-image features of the realistic type image generated at the same position in the denoising network to obtain sub-splicing features. The noise of this denoising time step is estimated using the sub-splicing features.
[0191] In a specific implementation, when the denoising network is a Unet network, the reference network is also a Unet network, that is, both the denoising network and the reference network contain multiple Attention layers. Considering that the Attention layer in the form of Cross-Attention involves not only image features but also text features, while the role of the reference network is to encode human features and help the denoising network restore the realistic image of the human, which only involves image features, so the Attention layers in the denoising network and the reference network of this application are both in the form of Self-Attention.
[0192] The input features of the Attention layer encoded by the reference network (including the human features of the sample human) are passed to the Attention layer at the same position in the denoising network in a vector splicing manner and spliced with the input features of the Attention layer at the same position in the denoising network (including the sub-image features of the realistic type image) to obtain a splicing result. Then, based on the splicing result, the noise at this denoising time step is estimated.
[0193] For example, referring to Figure 9 , the denoising network includes Attention layer 1, Attention layer 2, and Attention layer 3. Since the network structures of the denoising network and the reference network are the same, the reference network includes Attention layer 1, Attention layer 2, and Attention layer 3.
[0194] Taking Attention layer 1 as an example, the input features of Attention layer 1 in the reference network include: matrix , matrix , and matrix ; the input features of Attention layer 1 in the denoising network include: matrix , matrix , and matrix . The reference network passes matrix to Attention layer 1 of the denoising network and splices it with matrix ; the reference network passes matrix to Attention layer 1 of the denoising network and splices it with matrix . In this case, the calculation formula of Attention layer 1 of the denoising network is as shown in formula (4):
[0195]
[0196] where y represents the output of the Attention layer.
[0197] Similarly, the input features of other Attention layers in the reference network can also be transmitted to the Attention layer at the same position in the denoising network and concatenated with the input of the Attention layer at the same position. Details are not described herein again.
[0198] In the embodiments of the present application, the character features of the sample character generated by the reference network are transmitted to the denoising network and concatenated with the sub-image features of the realistic type image generated at the same position in the denoising network to obtain sub-concatenated features. Then, the sub-concatenated features are used to estimate the noise at the denoising time step. In this way, the character features of the sample character can fully play a role in the denoising process, enabling the de-stylization model to learn to generate realistic images of the sample character by gradually denoising, thereby improving the training effect of the de-stylization model.
[0199] Step 5022: Determine the model loss value of the de-stylization model used in this round according to each obtained predicted realistic image and the corresponding sample realistic image.
[0200] Specifically, when calculating the model loss value, the model loss value of the de-stylization model is obtained based on the difference between each predicted realistic image and the corresponding sample realistic image.
[0201] In the embodiments of the present application, the model loss value can adopt loss functions such as cross-entropy loss function (Cross Entropy Loss, celoss), mean squared error (Mean Squared Error, MSE) loss function, squared absolute error loss function, maximum likelihood loss (Likelihood Loss, LHL) function, etc. to obtain the model loss value of the de-stylization model. Of course, it can also be other possible loss functions, and the embodiments of the present application do not limit this.
[0202] Step 5023: Determine whether the de-stylization model used in this round reaches the iteration termination condition. If so, end; otherwise, execute step 5024.
[0203] In the embodiments of the present application, the iteration termination condition may include at least one of the following conditions:
[0204] (1) The number of iterations reaches the set number threshold.
[0205] (2) The model loss value is less than the set loss threshold, or the model loss value reaches the minimum value.
[0206] Step 5024: Adjust the parameters of the de-stylization model used in this round according to the model loss value, and use the de-stylization model with adjusted parameters to enter the next round of iterative training.
[0207] In an embodiment of the present application, when the number of iterations does not exceed a preset number threshold and the model loss value is not less than a set loss threshold, the determination process of step 5023 is negative, that is, it is considered that the current model does not meet the iteration termination condition, so it is necessary to adjust the model parameters and continue training. After the parameter adjustment, the next round of iterative training process is entered, that is, it jumps to step 5021.
[0208] In a possible implementation manner, when the model still does not meet the convergence condition, the model weight parameters can be updated through optimization algorithms such as the gradient descent method and the stochastic gradient descent algorithm to minimize the above loss function, and continue training with the updated model weight parameters.
[0209] When the number of iterations has exceeded the preset number threshold, or the model loss value is less than the set loss threshold or reaches the minimum value, the determination process of step 5023 is positive, that is, it is considered that the current model has met the convergence condition, the model training ends, and a trained de-stylized model is obtained.
[0210] In an embodiment of the present application, the sample style image and the sample realistic image of the same sample person are used as training samples to perform multiple rounds of iterative training on the de-stylized model, so that the de-stylized model fully understands the character features in the sample style image and the image features of the realistic type images in the sample realistic image, so as to obtain the ability to generate realistic images consistent with the character features of the stylized image in a zero-shot manner. Therefore, when using the de-stylized model to generate realistic images consistent with the character features of the stylized image in the application stage, not only the scalability and generalization ability are improved, but also the accuracy of the realistic images and stylized images of the same person obtained is improved.
[0211] After introducing the training process of the de-stylized model, the following is based on the trained de-stylized model and Figure 1 the system architecture diagram shown, to introduce the process of an image generation method provided by an embodiment of the present application. As Figure 10 shown, the process of this method is executed by a computer device, and the computer device can be Figure 1 the terminal device 101 and / or the server 102 shown, including the following steps:
[0212] Step 1001, obtain a pure noise image and a stylized image of a target person.
[0213] Specifically, the pure noise image follows a standard normal distribution and has the same size as the stylized image of the target person. It should be noted that in the application stage of the de-stylization model, there is no longer a need to input the realistic image of a person into the de-stylization model. Correspondingly, there is no longer a need for the second encoding module and the second noise addition module in the de-stylization model to add noise to the realistic image of a person to obtain a pure noise image. Instead, a pure noise image that follows a standard normal distribution is directly collected as the input of the de-stylization model.
[0214] The style corresponding to the stylized image of the target person can be the style included in the training samples when training the de-stylization model, or a new style not included in the training samples. The person corresponding to the stylized image of the target person can be the person included in the training samples when training the de-stylization model, or a new person not included in the training samples.
[0215] Step 1002: Extract the personal features of the target person from the stylized image of the target person.
[0216] Specifically, the personal features of the target person are related to the characteristics of the target person himself / herself and have nothing to do with the style corresponding to the stylized image. The personal features of the target person can be the characteristics of the person himself / herself, such as gender, age, eye size, nose bridge height, etc.
[0217] In some embodiments, the following operations are performed through the de-stylization model to extract the personal features of the target person:
[0218] Add noise to the stylized image of the target person according to a preset noise addition pattern to obtain a target noisy image, where the preset noise addition pattern is obtained by training the de-stylization model. Extract the personal features of the target person from the target noisy image.
[0219] In practical applications, the stylized image of the target person is input into the first encoding module in the de-stylization model, and the first encoding module maps the stylized image of the target person from the pixel space to the latent space. The first noise addition module adds noise to the stylized image of the target person in the latent space to obtain a target noisy image, where the target noisy image is a pure noise image that follows a standard normal distribution.
[0220] In the embodiments of the present application, the personal features of the target person are extracted from the stylized image to guide the generation of the target realistic image of the target person. In this process, the style features of the stylized image are not concerned. Therefore, for any style or any person, the corresponding realistic image can be generated, realizing zero-shot generation of a realistic image consistent with the personal features of the stylized image, significantly reducing the training cost and improving the scalability.
[0221] In practical applications, the stylized image of the target person can be denoised once using a preset denoising mode to directly obtain the target denoised image; alternatively, the stylized image of the target person can be denoised through multiple rounds of iteration using the preset denoising mode to obtain the target denoised image.
[0222] In some embodiments, when the stylized image of the target person is denoised through multiple rounds of iteration using the preset denoising mode, the preset denoising mode includes: multiple denoising time steps and the first noise corresponding to each of the multiple denoising time steps. The first noise can be Gaussian noise that satisfies the standard normal distribution.
[0223] The intensities of the first noise corresponding to each of the multiple denoising time steps can be the same or different. For example, the intensities of the first noise corresponding to each of the multiple denoising time steps decrease sequentially within the multiple denoising time steps. Or, for example, the intensities of the first noise corresponding to each of the multiple denoising time steps increase sequentially within the multiple denoising time steps. Of course, the intensities of the first noise corresponding to each of the multiple denoising time steps can also be in other forms, which are not specifically limited here.
[0224] Within the multiple denoising time steps, the stylized image is denoised iteratively to obtain the target denoised image; wherein, each round of iterative denoising includes the following steps:
[0225] Obtain the first noise of one denoising time step corresponding to the current round of iterative denoising. Use the first noise of this denoising time step to denoise the stylized image after the previous round of iterative denoising to obtain the stylized image after the current round of iterative denoising.
[0226] Specifically, in the first round of denoising (corresponding to the first denoising time step), input the information of the first denoising time step and the stylized image of the target person into the first denoising module, predict the first noise of the first denoising time step, and use the first noise of the first denoising time step to denoise the input stylized image of the target person to obtain the stylized image after the first round of denoising (i.e., the denoised image in the intermediate state).
[0227] For any other round of denoising, input the information of the denoising time step corresponding to the current round of denoising and the stylized image after the previous round of denoising into the first denoising module, predict the first noise of the denoising time step corresponding to the current round of denoising, and use the obtained first noise to denoise the stylized image after the previous round of denoising to obtain the stylized image after the current round of denoising (i.e., the denoised image in the intermediate state).
[0228] The stylized image after the last round of denoising (corresponding to the last denoising time step) is the obtained target denoised image.
[0229] In the embodiments of the present application, within multiple noise-adding time steps, the stylized image of the target person is iteratively noise-added in multiple rounds, gradually converting the obtained stylized image of the target person into a target noise image of pure noise, so as to subsequently reversely restore the realistic image of the target person from the pure noise image step by step, thereby improving the quality of the restored realistic image of the target person.
[0230] Step 1003: Extract the image features of the realistic type image and the personal features of the target person from the pure noise image, and the target noise information that does not meet the correlation condition. The image features of the realistic type image are learned from multiple sample realistic images.
[0231] Step 1004: Remove the target noise information from the pure noise image to obtain the target realistic image of the target person.
[0232] Specifically, not meeting the correlation condition may be irrelevant to the image features of the realistic type image and the personal features of the target person, or the correlation degree with the image features of the realistic type image and the personal features of the target person is less than a preset threshold.
[0233] Extract the personal features of the target person from the target noise image through a reference network, and transmit the personal features of the target person to a denoising network. Extract the image features of the realistic type image from the pure noise image through the denoising network, then combine the personal features of the target person transmitted by the reference network, predict the target noise information in the pure noise image that is irrelevant to the image features of the realistic type image and the personal features of the target person, and remove the target noise information from the pure noise image to obtain a target denoising result. Map the target denoising result from the latent space to the pixel space through a decoding module to obtain the target realistic image of the target person.
[0234] In the embodiments of the present application, taking the personal features of the target person in the stylized image as reference information, and combining the image features of the realistic type image learned from multiple sample realistic images, generate the target realistic image of the target person, which ensures that the obtained stylized image and the target realistic image correspond to the same person, thereby improving the accuracy of generating the realistic image and the stylized image of the same person.
[0235] Secondly, extract the personal features of the target person from the stylized image to guide the generation of the target realistic image of the target person, and it is not necessary to extract the style features of the stylized image. That is, the process of the de-stylization model generating the realistic image of a person has nothing to do with the style of the input stylized image. Therefore, for any style or any person, the corresponding realistic image can be generated, realizing zero-shot generation of the realistic image and the stylized image of the same person, improving the scalability and generalization ability of the de-stylization model, and also avoiding the resource consumption caused by repeated training.
[0236] In addition, the technical solution of this application ensures that the obtained stylized image and realistic image correspond to the same person. Therefore, when using the obtained stylized image and realistic image as training samples for subsequent model training, the accuracy of the training samples is correspondingly improved. Since the de-stylization model of this application has the zero-shot ability, without retraining the de-stylization model, a large number of diverse training samples can also be generated through the de-stylization model for subsequent model training, thereby improving the performance of model training.
[0237] In some embodiments, a denoising network can be used to perform denoising on the input pure noise image once to directly obtain the target realistic image of the target person. Specifically, the following operations are performed through the de-stylization model:
[0238] The image features of the realistic type image and the person features of the target person are concatenated to obtain target concatenated features; then, the target concatenated features are used to predict the probability that each pixel in the pure noise image contains noise information. Based on the probability that each pixel in the pure noise image contains noise information, the noise distribution in the pure noise image is obtained. Based on the noise distribution in the pure noise image, the target noise information in the pure noise image is extracted.
[0239] Specifically, the reference network transmits the person features of the generated target person to the denoising network and concatenates them with the image features of the realistic type image generated at the same position in the denoising network to obtain target concatenated features. When using the target concatenated features to predict the target noise information in the pure noise image, the obtained target noise information has a low correlation not only with the person features of the target person but also with the image features of the realistic type image. Then, after removing the target noise information from the pure noise image, the useful information with a high correlation with the person features of the target person and the image features of the realistic type image is retained. Therefore, based on this useful information, the target realistic image of the target person is restored, which can effectively improve the quality of the restored realistic image.
[0240] In some embodiments, a denoising network can be used to perform multi-round iterative denoising on the pure noise image to obtain the target realistic image of the target person. In this case, the target noise information includes: the second noise corresponding to each of the multiple denoising time steps. During the denoising process, the following operations are performed through the de-stylization model:
[0241] Within multiple denoising time steps, iterative denoising is performed on the pure noise image to obtain the target realistic image of the target person, where each round of iterative denoising includes the following steps:
[0242] Obtain the second noise of a denoising time step corresponding to this round of iteration; remove the second noise of this denoising time step from the pure noise image after the previous round of iterative denoising to obtain the pure noise image after this round of iterative denoising.
[0243] Specifically, the second noises corresponding to multiple denoising time steps are all obtained by predicting through a denoising network. The intensities of the second noises corresponding to multiple denoising time steps can be the same or different. For example, the intensities of the second noises corresponding to multiple denoising time steps decrease sequentially within the multiple denoising time steps. Or, for another example, the intensities of the second noises corresponding to multiple denoising time steps increase sequentially within the multiple denoising time steps. Of course, the intensities of the second noises corresponding to multiple denoising time steps can also be in other forms, which are not specifically limited herein.
[0244] In some cases, the number of noise addition time steps adopted by the first noise addition module is the same as the number of denoising time steps adopted by the denoising network, and the variation trend of the intensities of the first noises corresponding to multiple noise addition time steps is symmetric with the variation trend of the intensities of the second noises corresponding to multiple denoising time steps.
[0245] For example, the first noise addition module includes 4 noise addition time steps, and the intensities of the first noises corresponding to the 4 noise addition time steps decrease sequentially within the 4 noise addition time steps; the denoising network includes 4 denoising time steps, and the intensities of the second noises corresponding to the 4 denoising time steps increase sequentially within the 4 denoising time steps.
[0246] In the embodiments of the present application, within multiple denoising time steps, multi-round iterative denoising is performed on the pure noise image. In each round of denoising process, noise information with a low correlation with the personal characteristics of the target person and the image characteristics of the realistic type image is removed, and the input pure noise image is gradually converted into the target realistic image of the target person, thereby improving the quality of the obtained target realistic image of the target person.
[0247] In some embodiments, the image characteristics of the realistic type image include: sub-image characteristics corresponding to multiple denoising time steps. The sub-image characteristics corresponding to each denoising time step are extracted from the pure noise image after the previous round of iterative denoising.
[0248] In each round of denoising process, based on the personal characteristics of the target person transmitted by the reference network and the sub-image characteristics of the corresponding denoising time step, the denoising network predicts the second noise required for this round of iterative denoising. The following takes the process of obtaining the second noise of a denoising time step corresponding to this round of iteration as an example for specific description:
[0249] The personal characteristics of the target person are spliced with the sub-image characteristics corresponding to this denoising time step to obtain the sub-splicing characteristics corresponding to this denoising time step. Based on the sub-splicing characteristics corresponding to this denoising time step, the probability that each pixel in the pure noise image after the previous round of iterative denoising contains noise information is predicted.
[0250] Based on the probability that each pixel in the pure noise image after denoising in the previous iteration contains noise information, the noise distribution in the pure noise image after denoising in the previous iteration is obtained. Then, based on the noise distribution in the pure noise image after denoising in the previous iteration, the second noise corresponding to this denoising time step is obtained.
[0251] Specifically, in the first round of denoising (corresponding to the first denoising time step), the information of the first denoising time step and the original pure noise image are input into the denoising network to obtain the sub-image features of the realistic type image corresponding to the first denoising time step. The reference network transmits the character features of the generated target character to the denoising network and splices them with the sub-image features of the realistic type image generated by the denoising network to obtain the sub-spliced features. The denoising network uses the sub-spliced features to predict the second noise of the first denoising time step. Further, the denoising network uses the obtained second noise to denoise the original pure noise image to obtain the pure noise image after the first round of denoising (i.e., the denoised image in the intermediate state).
[0252] For any other round of denoising, the information of this round of denoising time step and the pure noise image after denoising in the previous round are input into the denoising network to obtain the sub-image features of the realistic type image corresponding to this round of denoising time step. The reference network transmits the character features of the generated target character to the denoising network and splices them with the sub-image features of the realistic type image generated by the denoising network to obtain the sub-spliced features. The denoising network uses the sub-spliced features to predict the second noise of this round of denoising time step. Further, the denoising network uses the obtained second noise to denoise the pure noise image after denoising in the previous round to obtain the pure noise image after this round of denoising (i.e., the denoised image in the intermediate state).
[0253] The pure noise image after the last round of denoising (corresponding to the last denoising time step) is the obtained target denoising result. The target denoising result is mapped from the latent space to the pixel space through the decoding module to obtain the target realistic image of the target character.
[0254] In the embodiments of the present application, during each round of denoising, the sub-image features of the realistic type image used for denoising in this round of iteration are extracted from the pure noise image after denoising in the previous iteration. The character features of the target character are spliced with the sub-image features of the realistic type image used for denoising in this round of iteration to obtain the sub-spliced features used for denoising in this round of iteration. Then, based on the obtained sub-spliced features, the second noise used for denoising in this round of iteration is predicted. In this way, the second noise is more matched with the noise image to be denoised in this round, thereby improving the effect of multi-round iterative denoising and further improving the quality of the generated target realistic image.
[0255] In some embodiments, a large number of stylized data pairs (i.e., realistic images and stylized images of the same person) can be obtained by using the de-stylization model of the embodiments of the present application. Then, the obtained large number of stylized data pairs can be used to train the stylization model.
[0256] Specifically, to obtain multi-style picture data, the LAION-5B data set is selected as the pre-training data to pre-train the stylization model. Among them, the LAION-5B data set is a large text-to-image data set, and the LAION-5B data set contains multi-domain data. To improve the quality of the pre-training data, the present application screens the obtained LAION-5B data set. The specific screening methods include: aesthetic quality scoring and face ratio screening.
[0257] Specifically, the aesthetic quality score is a picture quality evaluation index used in the Laion data set, with a range of [0, 10]. The higher the score, the better the quality of the picture. The present application uses the aesthetic quality score to screen the LAION-5B data set and retains the pictures with a score value greater than 6.5.
[0258] The face ratio screening refers to screening images according to the picture size and the face ratio. Specifically, the shortest side of the image is greater than 768. Since the smaller the face ratio, the greater the possibility of face distortion and blurring. Therefore, the present application proposes to retain only the pictures with a face ratio above 1 / 16. Taking an image of size as an example, the square face size with a ratio of 1 / 16 is
[0259] which basically ensures the generation of high-quality faces.
[0260] In the embodiments of the present application, a large number of stylized data pairs (i.e., realistic images and stylized images of the same person) are obtained by using the de-stylization model and used as training samples to train the stylization model, greatly reducing the difficulty of obtaining training samples.
[0261] To better explain the embodiments of the present application, the following introduces an image generation method provided by the embodiments of the present application in combination with a specific implementation scenario. The flow of this method can be Figure 1The execution of the terminal device 101 shown can also be performed by the server 102, or by the interaction between the terminal device 101 and the server 102. The process of this method mainly includes a training phase and an application phase.
[0262] First, the training phase is introduced. Refer to Figure 11A , it is set that both the first encoding module and the second encoding module in the de-styling model are variational auto-encoders (VAEs), the reference network and the denoising network are both UNet networks, and the decoding module is the decoder corresponding to the VAE.
[0263] Obtain a sample set containing multiple training samples. Each training sample includes: a realistic image and a stylized image of the same person. Using the sample set, perform multiple rounds of iterative training on the de-styling model to be trained until the training termination condition is met, and obtain the trained de-styling model. The following takes a training sample in one iteration process as an example for specific introduction, including the following steps:
[0264] Set that a training sample includes: the sample realistic image of person C and the sample stylized image of person C, and input a training sample into the de-styling model to be trained.
[0265] In the de-styling model to be trained, the first encoding module maps the sample stylized image of person C from the pixel space to the latent space. The first noise addition module gradually adds Gaussian noise to the sample stylized image of person C in the latent space through multiple noise addition time steps to obtain a pure noise image 1.
[0266] The second encoding module maps the sample realistic image of person C from the pixel space to the latent space. The second noise addition module gradually adds Gaussian noise to the sample realistic image of person C in the latent space through multiple noise addition time steps to obtain a pure noise image 2.
[0267] In the first denoising time step, the reference network extracts the person feature of person C from the pure noise image 1 and transfers the person feature of person C to the denoising network. The denoising network extracts the image feature of the realistic type image based on the first denoising time step and the pure noise image 2, and splices it with the received person feature of person C to obtain a first splicing feature. The denoising network uses the first splicing feature to predict the Gaussian noise at the first denoising time step and removes the Gaussian noise at the first denoising time step from the pure noise image 2 to obtain the intermediate state image at the first time step.
[0268] At the second denoising time step, the reference network extracts the person feature of person C from the pure noise image 1 and transfers the person feature of person C to the denoising network. The denoising network extracts the image feature of the realistic type image based on the intermediate state image of the second denoising time step and the first time step, and splices it with the received person feature of person C to obtain the second splicing feature. The denoising network uses the second splicing feature to predict the Gaussian noise of the second denoising time step, and removes the Gaussian noise of the second denoising time step from the intermediate state image of the first time step to obtain the intermediate state image of the second time step.
[0269] And so on until the intermediate state image of the last denoising time step is obtained.
[0270] The decoding module maps the intermediate state image of the last denoising time step from the latent space to the pixel space to obtain the predicted realistic image of person C.
[0271] The difference between each obtained predicted realistic image and the corresponding sample realistic image (including the difference between the sample realistic image of person C and the predicted realistic image of person C) is used to obtain the model loss value of the de-stylization model, and the de-stylization model used in this round is tuned according to the model loss value.
[0272] After introducing the training stage, the application stage will be introduced next. See Figure 11B , and input the stylized image of person D and the pure noise image 3 collected into the trained de-stylization model.
[0273] In the trained de-stylization model, the first encoding module maps the stylized image of person D from the pixel space to the latent space. The first noise addition module gradually adds Gaussian noise to the stylized image of person D through multiple noise addition time steps in the latent space to obtain a noise-added image.
[0274] At the first denoising time step, the reference network extracts the person feature of person D from the noise-added image and transfers the person feature of person D to the denoising network. The denoising network extracts the image feature of the realistic type image based on the first denoising time step and the input pure noise image 3, and splices it with the received person feature of person D to obtain the first splicing feature. The denoising network uses the first splicing feature to predict the Gaussian noise of the first denoising time step, and removes the Gaussian noise of the first denoising time step from the pure noise image 3 to obtain the intermediate state image of the first time step.
[0275] At the second denoising time step, the reference network extracts the person features of person D from the noisy image and transfers the person features of person D to the denoising network. The denoising network extracts the image features of the realistic type image based on the intermediate state image of the second denoising time step and the first time step, and splices them with the received person features of person D to obtain the second splicing feature. The denoising network uses the second splicing feature to predict the Gaussian noise at the second denoising time step, and removes the Gaussian noise at the second denoising time step from the intermediate state image of the first time step to obtain the intermediate state image of the second time step.
[0276] And so on until the intermediate state image of the last denoising time step is obtained.
[0277] The decoding module maps the intermediate state image of the last denoising time step from the latent space to the pixel space to obtain the realistic image of person D.
[0278] In the embodiment of the present application, the sample style image and the sample realistic image of the same sample person are used as training samples to perform multiple rounds of iterative training on the de-styling model, so that the de-styling model fully understands the person features in the sample style image and the image features of the realistic type image in the sample realistic image, so as to obtain the ability to generate a realistic image consistent with the person features of the stylized image in a zero-shot manner. Therefore, when using the de-styling model to generate a realistic image consistent with the person features of the stylized image in the application stage, not only the scalability and generalization ability are improved, but also the accuracy of the realistic image and the stylized image of the same person obtained is improved. Secondly, by using this method, a large number of stylized data pairs can be constructed for the training of the stylization model, which greatly reduces the difficulty of obtaining training samples.
[0279] Based on the same technical concept, the embodiment of the present application provides a schematic structural diagram of an image generation device, as Figure 11C shown. The image generation device 1100 includes:
[0280] An acquisition module 1101, configured to acquire a pure noise image and a stylized image of a target person;
[0281] An extraction module 1102, configured to extract the person features of the target person from the stylized image of the target person;
[0282] A prediction module 1103, configured to extract target noise information from the pure noise image that does not meet the correlation condition with the image features of the realistic type image and the person features of the target person, and the image features of the realistic type image are learned from multiple sample realistic images;
[0283] A denoising module 1104, configured to remove the target noise information from the pure noise image to obtain a target realistic image of the target person.
[0284] Optionally, the extraction module 1102 is specifically configured to:
[0285] Perform the following operations through a trained destylization model:
[0286] Add noise to the stylized image of the target person according to a preset noise addition mode to obtain a target noisy image, where the preset noise addition mode is obtained by training the destylization model;
[0287] Extract the person features of the target person from the target noisy image.
[0288] Optionally, the preset noise addition mode includes: a plurality of noise addition time steps and first noises respectively corresponding to the plurality of noise addition time steps; the intensities of the first noises respectively corresponding to the plurality of noise addition time steps are different;
[0289] The extraction module 1102 is specifically configured to:
[0290] Perform iterative noise addition on the stylized image within the plurality of noise addition time steps to obtain the target noisy image; wherein, each round of iterative noise addition includes the following steps:
[0291] Obtain the first noise of one noise addition time step corresponding to the current round of iterative noise addition;
[0292] Use the first noise of the one noise addition time step to add noise to the stylized image after the previous round of iterative noise addition to obtain the stylized image after the current round of iterative noise addition.
[0293] Optionally, the prediction module 1103 is specifically configured to:
[0294] Perform the following operations through the trained destylization model:
[0295] Concatenate the image features of the realistic type image and the person features of the target person to obtain target concatenated features;
[0296] Use the target concatenated features to predict the probability that each pixel in the pure noise image contains noise information;
[0297] Based on the probability that each pixel in the pure noise image contains noise information, obtain the noise distribution in the pure noise image;
[0298] Based on the noise distribution in the pure noise image, extract the target noise information in the pure noise image.
[0299] Optionally, the target noise information includes: second noises corresponding to multiple denoising time steps; the intensities of the second noises corresponding to the multiple denoising time steps are different;
[0300] The denoising module 1104 is specifically configured to:
[0301] Perform the following operations through the trained de-stylization model:
[0302] Iteratively denoise the pure noise image within the multiple denoising time steps to obtain a target realistic image of the target person; wherein, each round of iterative denoising includes the following steps:
[0303] Obtain the second noise of a denoising time step corresponding to the current round of iterative denoising;
[0304] Remove the second noise of the denoising time step from the pure noise image after the previous round of iterative denoising to obtain the pure noise image after the current round of iterative denoising.
[0305] Optionally, the image features of the realistic type image include: sub-image features corresponding to the multiple denoising time steps;
[0306] The denoising module 1104 is specifically configured to:
[0307] Concatenate the character features of the target person with the sub-image features corresponding to a denoising time step to obtain sub-concatenated features corresponding to the denoising time step;
[0308] Predict the probability that each pixel in the pure noise image after the previous round of iterative denoising contains noise information based on the sub-concatenated features corresponding to the denoising time step;
[0309] Obtain the noise distribution in the pure noise image after the previous round of iterative denoising based on the probability that each pixel in the pure noise image after the previous round of iterative denoising contains noise information;
[0310] Obtain the second noise corresponding to the denoising time step based on the noise distribution in the pure noise image after the previous round of iterative denoising.
[0311] Optionally, it further includes a module training module 1105;
[0312] The module training module 1105 is specifically configured to:
[0313] Obtain a sample set including multiple training samples, and each training sample includes: a sample style image and a sample realistic image of the same sample person;
[0314] Using the sample set, perform multiple rounds of iterative training on the de-stylization model to be trained until the training termination condition is met, to obtain the trained de-stylization model, where each round of iterative training process includes the following steps:
[0315] Through the de-stylization model used in this round, transform each training sample input in this round of iteration to respectively obtain the predicted realistic images of the respective training samples;
[0316] Based on the obtained predicted realistic images and the corresponding sample realistic images, determine the model loss value of the de-stylization model used in this round, and after adjusting the parameters of the de-stylization model used in this round according to the model loss value, use the de-stylization model with adjusted parameters to enter the next round of iterative training.
[0317] Optionally, the module training module 1105 is specifically configured to:
[0318] For each of the training samples, respectively perform the following operations:
[0319] Add noise to the sample style image in a training sample to obtain a first noise-added image; and, add noise to the sample realistic image in the training sample to obtain a second noise-added image;
[0320] Extract the character features of the sample character from the first noise-added image; and, extract the image features of the realistic type image from the second noise-added image;
[0321] Based on the character features of the sample character and the image features of the realistic type image, obtain the predicted realistic image of the sample character.
[0322] Optionally, the module training module 1105 is specifically configured to:
[0323] Perform splicing based on the character features of the sample character and the image features of the realistic type image to obtain sample splicing features;
[0324] Based on the sample splicing features, obtain the predicted noise information in the second noise-added image;
[0325] Remove the predicted noise information from the second noise-added image to obtain the predicted realistic image of the sample character.
[0326] In the embodiments of the present application, taking the character features of the target character in the stylized image as reference information, and combining the image features of the realistic type image learned from multiple sample realistic images, generate the target realistic image of the target character, which ensures that the obtained stylized image and the target realistic image correspond to the same character, thereby improving the accuracy of generating the realistic image and the stylized image of the same character.
[0327] Secondly, the personal features of the target person are extracted from the stylized image to guide the generation of the target realistic image of the target person, and it is not necessary to extract the style features of the stylized image. That is, the process of the de-stylization model generating the realistic image of the person has nothing to do with the style of the input stylized image. Therefore, for any style or any person, the corresponding realistic image can be generated, realizing zero-shot generation of the realistic image and the stylized image of the same person, improving the scalability and generalization ability, and also avoiding the resource consumption caused by repeated training.
[0328] In addition, the technical solution of this application ensures that the obtained stylized image and realistic image correspond to the same person. Therefore, when using the obtained stylized image and realistic image as training samples, the accuracy of the training samples is correspondingly improved. Moreover, the de-stylization model of this application has zero-shot ability. Therefore, a large number of training samples can be generated through the de-stylization model for model training, thereby improving the performance of model training.
[0329] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0330] Based on the same technical concept, the embodiments of this application provide a computer device, which can be Figure 1 the terminal device and / or server shown in the figure, such as Figure 12 shown, including at least one processor 1201 and a memory 1202 connected to at least one processor. In the embodiments of this application, the specific connection medium between the processor 1201 and the memory 1202 is not limited. Figure 12 Taking the connection between the processor 1201 and the memory 1202 through a bus as an example. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0331] In the embodiments of this application, the memory 1202 stores instructions executable by at least one processor 1201. By executing the instructions stored in the memory 1202, at least one processor 1201 can execute the steps of the above image generation method.
[0332] As an embodiment, the processor 1201 obtains a pure noise image and a stylized image of a target person; the processor 1201 extracts the personal features of the target person from the stylized image of the target person; and extracts target noise information from the pure noise image, where the image features of the realistic type image and the personal features of the target person do not meet the correlation condition. The image features of the realistic type image are learned from multiple sample realistic images and stored in the memory 1202. The processor 1201 removes the target noise information from the pure noise image to obtain a target realistic image of the target person.
[0333] Among them, the processor 1201 is the control center of the computer device, and can connect various parts of the computer device through various interfaces and lines. By running or executing the instructions stored in the memory 1202 and calling the data stored in the memory 1202, the realistic image of the person can be obtained. Optionally, the processor 1201 may include one or more processing units. The processor 1201 may integrate an application processor and a modulation / demodulation processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modulation / demodulation processor mainly processes wireless communication. It can be understood that the above modulation / demodulation processor may not be integrated into the processor 1201. In some embodiments, the processor 1201 and the memory 1202 can be implemented on the same chip. In some embodiments, they can also be separately implemented on independent chips.
[0334] The processor 1201 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0335] The memory 1202, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 1202 can include at least one type of storage medium. For example, it can include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical disks, and so on. The memory 1202 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited thereto. The memory 1202 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0336] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium storing a computer program executable by a computer device. When the computer program runs on the computer device, it causes the computer device to execute the steps of the above image generation method.
[0337] Based on the same inventive concept, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer device, cause the computer device to execute the steps of the above image generation method.
[0338] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0339] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer device or other programmable data processing device generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.
[0340] These computer program instructions can also be stored in a computer-readable memory that can direct a computer device or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.
[0341] These computer program instructions can also be loaded onto a computer device or other programmable data processing device, such that a series of operational steps are executed on the computer device or other programmable device to produce a process implemented by the computer device, so that the instructions executed on the computer device or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.
[0342] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0343] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. An image generation method, characterized in that, Including: Obtain a pure noise image and a stylized image of a target person; Add noise to the stylized image of the target person according to a preset noise addition mode to obtain a target noise-added image; Extract the person features of the target person from the target noise-added image through a reference network, and the person features are independent of the style corresponding to the stylized image; Transmit the person features of the target person to a denoising network, and splice them with the image features of a realistic type image generated at the same position in the denoising network to obtain target splicing features; the image features of the realistic type image are extracted by the denoising network from the pure noise image; the network structures of the reference network and the denoising network are the same; Use the target splicing features to predict the target noise information in the pure noise image that does not meet the correlation condition with the image features of the realistic type image and the person features of the target person, and the image features of the realistic type image are learned from multiple sample realistic images; Remove the target noise information from the pure noise image to obtain the target realistic image of the target person, and the stylized image and the target realistic image correspond to the same person.
2. The method according to claim 1, wherein The preset noise addition mode is obtained by training a destylization model.
3. The method according to claim 2, characterized in that, The preset noise addition mode includes: a plurality of noise addition time steps and the first noise corresponding to each of the plurality of noise addition time steps; the intensities of the first noise corresponding to each of the plurality of noise addition time steps are different; The adding noise to the stylized image of the target person according to the preset noise addition mode to obtain a target noise-added image includes: During the plurality of noise addition time steps, iteratively add noise to the stylized image of the target person to obtain the target noise-added image; wherein, each round of iterative noise addition includes the following steps: Obtain the first noise of one noise addition time step corresponding to the current round of iterative noise addition; Use the first noise of the one noise addition time step to add noise to the stylized image after the previous round of iterative noise addition to obtain the stylized image after the current round of iterative noise addition.
4. The method according to claim 2, wherein The using the target splicing features to predict the target noise information in the pure noise image that does not meet the correlation condition with the image features of the realistic type image and the person features of the target person includes: Perform the following operations through a trained destylization model: Use the target splicing features to predict the probability that each pixel in the pure noise image contains noise information; Based on the probability that each pixel in the pure noise image contains noise information, obtain the noise distribution in the pure noise image; Based on the noise distribution in the pure noise image, extract the target noise information in the pure noise image.
5. The method according to claim 2, characterized in that The target noise information includes: the second noise corresponding to each of the plurality of denoising time steps; the intensities of the second noise corresponding to each of the plurality of denoising time steps are different; The removing the target noise information from the pure noise image to obtain the target realistic image of the target person includes: Perform the following operations through a trained destylization model: During the plurality of denoising time steps, iteratively denoise the pure noise image to obtain the target realistic image of the target person; wherein, each round of iterative denoising includes the following steps: Obtain the second noise for a denoising time step corresponding to the current iteration of denoising; Remove the second noise of the one denoising time step from the pure noise image after denoising in the previous iteration to obtain the pure noise image after denoising in the current iteration.
6. The method according to claim 5, wherein The image features of the realistic type image include: sub-image features corresponding to each of the multiple denoising time steps; The obtaining the second noise for a denoising time step corresponding to the current iteration of denoising includes: Concatenate the character features of the target person and the sub-image features corresponding to the one denoising time step to obtain the sub-concatenated features corresponding to the one denoising time step; Based on the sub-concatenated features corresponding to the one denoising time step, predict the probability that each pixel in the pure noise image after denoising in the previous iteration contains noise information; Based on the probability that each pixel in the pure noise image after denoising in the previous iteration contains noise information, obtain the noise distribution in the pure noise image after denoising in the previous iteration; Based on the noise distribution in the pure noise image after denoising in the previous iteration, obtain the second noise corresponding to the one denoising time step.
7. The method according to any one of claims 2 to 6, characterized in that, The training process of the de-stylization model includes the following steps: Obtain a sample set including multiple training samples, and each training sample includes: a sample style image and a sample realistic image of the same sample person; Use the sample set to perform multiple rounds of iterative training on the de-stylization model to be trained until the training termination condition is met, and obtain the trained de-stylization model. Among them, each round of iterative training process includes the following steps: Through the de-stylization model used in the current round, transform each training sample input in the current iteration to obtain the predicted realistic image of each training sample respectively; According to the obtained predicted realistic images and the corresponding sample realistic images, determine the model loss value of the de-stylization model used in the current round, and after adjusting the parameters of the de-stylization model used in the current round according to the model loss value, use the de-stylization model with adjusted parameters to enter the next round of iterative training.
8. The method according to claim 7, wherein The through the de-stylization model used in the current round, transform each training sample input in the current iteration to obtain the predicted realistic image of each training sample respectively, includes: For each of the training samples, perform the following operations respectively: Add noise to the sample style image in a training sample to obtain a first noise-added image; and, add noise to the sample realistic image in the one training sample to obtain a second noise-added image; Extract the character features of the sample person from the first noise-added image; and, extract the image features of the realistic type image from the second noise-added image; Based on the character features of the sample person and the image features of the realistic type image, obtain the predicted realistic image of the sample person.
9. The method according to claim 8, wherein The based on the character features of the sample person and the image features of the realistic type image, obtain the predicted realistic image of the sample person, includes: Concatenate based on the character features of the sample person and the image features of the realistic type image to obtain sample concatenated features; Based on the sample concatenated features, obtain the predicted noise information in the second noise-added image; From the second noisy image, remove the predicted noise information to obtain the predicted realistic image of the sample person.
10. An image generation device, characterized in that, Including: An acquisition module, configured to acquire a pure noise image and a stylized image of a target person; An extraction module, configured to add noise to the stylized image of the target person according to a preset noise addition mode to obtain a target noisy image; extract the person features of the target person from the target noisy image through a reference network, and the person features are independent of the style corresponding to the stylized image; A prediction module, configured to transmit the person features of the target person to a denoising network, and splice the person features with the image features of a realistic type image generated at the same position in the denoising network to obtain target splicing features; the image features of the realistic type image are extracted by the denoising network from the pure noise image; the reference network has the same network structure as the denoising network; use the target splicing features to predict the target noise information in the pure noise image that does not satisfy the correlation condition with the image features of the realistic type image and the person features of the target person, and the image features of the realistic type image are learned from multiple sample realistic images; A denoising module, configured to remove the target noise information from the pure noise image to obtain the target realistic image of the target person, and the stylized image and the target realistic image correspond to the same person.
11. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium, characterized in that, It stores a computer program executable by a computer device. When the computer program runs on the computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Image processing method, device and equipment and readable storage medium
CN117788273A
Image processing method and device, program product, computer equipment and storage medium
CN119228665A