Image processing method and device, computer equipment, storage medium and program product
By obtaining the prompt information and outline feature map of the initial rendered image, combining the style and outline constraint model, the target style constraint model is trained, and the problem of poor quality of the three-rendered and two-rendered image is solved, achieving higher quality two-dimensional rendering effect.
Patent Information
- Application Number
- CN202410123053.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-07-29
AI Technical Summary
In the traditional three-rendering and two-dimensional technology, the image quality is poor in the process of rendering a cartoon three-dimensional model to a two-dimensional image, making it difficult to achieve the desired effect.
By obtaining multiple initial rendered images of the target virtual character, obtaining their prompt information and outline feature maps, using the pre-trained style constraint model and outline constraint model, generate a style rendered image, and train the target style constraint model to perform two-dimensional rendering of the three-dimensional character model to obtain the style matching the reference rendered image.
Improve the quality of the three-rendered two-rendered image, making its style and details more match the reference rendered image, and improving the visual effect of the image.
Smart Images

Figure CN120388111A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and particularly to an image processing method, apparatus, computer device, storage medium, and computer program product. Background Art
[0002] With the development of computer technologies, the Cel shading technology has emerged. Cel shading, namely 3D modeling and 2D rendering, is a special rendering technique of de-realization. By parsing the planar color and contour on the basic appearance of a three-dimensional object, the object can present a two-dimensional effect while having a three-dimensional perspective, and finally endow the 3D graphics with the texture of 2D hand-drawing.
[0003] In traditional technologies, anime Cel shading is usually based on UE. After the 3D modeling is completed, shader parameters are constructed for the 3D scene model to render the 3D model scene into a 2D animation. However, there are often problems with poor image quality of the rendered images. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide an image processing method, apparatus, computer device, computer-readable storage medium, and computer program product, which can improve the image quality obtained by Cel shading rendering.
[0005] In a first aspect, the present application provides an image processing method. The method includes: obtaining a plurality of initial rendered images of a target virtual character, where the initial rendered images are obtained by performing 2D rendering on a three-dimensional character model of the target virtual character; for each initial rendered image, obtaining the hint information of the initial rendered image being targeted, and obtaining the contour feature map of the initial rendered image being targeted; obtaining a plurality of pre-trained style constraint models and contour constraint models, where the style combination indicated by the plurality of style constraint models fits a reference rendered image of the target virtual character; based on the hint information and the contour feature map, generating a style rendered image through the plurality of style constraint models and the contour constraint models, where the contour of the style rendered image matches the initial rendered image being targeted, and the style matches the reference rendered image; training a target style constraint model by using the style rendered image, where the target style constraint model is used to re-render a first rendered image obtained by performing 2D rendering on the three-dimensional character model to obtain a second rendered image with a style matching the reference rendered image.
[0006] In a second aspect, the present application also provides an image processing device. The device includes: an initial rendering image acquisition module, configured to acquire a plurality of initial rendering images of a target virtual character, where the initial rendering images are obtained by performing two-dimensional rendering on a three-dimensional character model of the target virtual character; a prompt information acquisition module, configured to, for each initial rendering image, acquire the prompt information of the initial rendering image targeted thereby, and acquire the contour feature map of the initial rendering image targeted thereby; a model acquisition module, configured to acquire a plurality of pre-trained style constraint models and contour constraint models, where the style combinations indicated by the plurality of style constraint models fit a reference rendering image of the target virtual character; a style rendering image generation module, configured to, based on the prompt information and the contour feature map, generate a style rendering image through the plurality of style constraint models and the contour constraint models, where the contour of the style rendering image matches that of the initial rendering image targeted thereby, and the style matches that of the reference rendering image; a model training module, configured to train a target style constraint model by using the style rendering image, where the target style constraint model is configured to perform secondary rendering on a first rendering image obtained by performing two-dimensional rendering on the three-dimensional character model, to obtain a second rendering image having a style matching that of the reference rendering image.
[0007] In some embodiments, the model training module is further configured to determine the style description information of the style rendering image; based on the style description information, determine a graphic and text data pair including the style rendering image; and use the graphic and text data pair to train the style constraint model to be trained, to obtain the target style constraint model.
[0008] In some embodiments, the model training module is further configured to determine the first expression description information of the style rendering image, where the first expression description information is used to describe the expression of the target virtual character in the style rendering image; and based on the style description information and the expression description information, determine a graphic and text data pair including the style rendering image.
[0009] In some embodiments, the data acquisition module is further configured to: acquire a trained label model, input the initial rendering image targeted thereby into the label model, to obtain a plurality of initial labels of the initial rendering image targeted thereby; screen the plurality of initial labels, to obtain a plurality of scene element prompt words of the initial rendering image targeted thereby; acquire a plurality of positive prompt words of the initial rendering image targeted thereby, where the plurality of positive prompt words are used to describe the image quality of the initial rendering image targeted thereby; and based on the plurality of scene element prompt words and the plurality of positive prompt words, determine the prompt information of the initial rendering image targeted thereby.
[0010] In some embodiments, the data acquisition module is further configured to: acquire a plurality of negative prompt words for the targeted initial rendering image, where the plurality of negative prompt words are used to describe the expected missing features of the targeted initial rendering image; determine the prompt information for the targeted initial rendering image based on the plurality of picture element prompt words, the plurality of positive prompt words, and the plurality of negative prompt words.
[0011] In some embodiments, the above device further includes: a secondary rendering module, configured to, when acquiring a first rendering image obtained by performing two-dimensional rendering based on the three-dimensional character model, determine the prompt information for the first rendering image and determine the contour feature map of the first rendering image; acquire the target style constraint model and the contour constraint model; generate a second rendering image with a style matching that of the reference rendering image through the target style constraint model and the contour constraint model based on the prompt information for the first rendering image and the contour feature map of the first rendering image.
[0012] In some embodiments, the secondary rendering module is further configured to acquire the plurality of pre-trained style constraint models; generate a second rendering image with a style matching that of the reference rendering image through the plurality of pre-trained style constraint models, the target style constraint model, and the contour constraint model based on the prompt information for the first rendering image and the contour feature map of the first rendering image.
[0013] In some embodiments, the secondary rendering module is further configured to respectively inject the model parameters of the plurality of pre-trained style constraint models and the target style constraint model into a pre-trained image generation model with different weights to obtain a target image generation model, where the weight corresponding to the target style constraint model is greater than the weights corresponding to the plurality of pre-trained style constraint models respectively; input the prompt information for the first rendering image into the target image generation model, and input the contour feature map of the first rendering image into the contour constraint model; perform image generation through the target image generation model, and perform contour constraint on the image generated by the target image generation model through the output of the contour constraint model to obtain a second rendering image with a style matching that of the reference rendering image.
[0014] In some embodiments, the secondary rendering module is further configured to acquire the preset initial prompt information for the target virtual character; determine the second expression description information for the first rendering image, where the second expression description information is used to describe the expression of the target virtual character in the first rendering image; add the second expression description information to the initial prompt information to obtain the prompt information for the first rendering image.
[0015] In some embodiments, the above device further includes: a face restoration module, configured to perform face detection on the second rendered image when the second rendered image is obtained; intercept the detected face region from the second rendered image, and enlarge the intercepted face region to obtain a face image; redraw the face image to obtain a redrawn face image; after shrinking the redrawn face image to match the size of the face region, post the obtained face image back to the face region to obtain a restored second rendered image.
[0016] In some embodiments, the multiple style constraint models are respectively obtained by training the low-rank matrix of the StableDiffusion model, and the contour constraint model is obtained by training a trainable copy cloned from the weights of the StableDiffusion model.
[0017] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: obtaining a plurality of initial rendered images of a target virtual character, where the initial rendered images are obtained by performing two-dimensional rendering on a three-dimensional character model of the target virtual character; for each initial rendered image, obtaining the prompt information of the initial rendered image targeted, and obtaining the contour feature map of the initial rendered image targeted; obtaining a plurality of pre-trained style constraint models and contour constraint models, where the style combinations indicated by the plurality of style constraint models fit a reference rendered image of the target virtual character; based on the prompt information and the contour feature map, generating a style rendered image through the plurality of style constraint models and the contour constraint models, where the contour of the style rendered image matches the initial rendered image targeted, and the style matches the reference rendered image; training a target style constraint model by using the style rendered image, where the target style constraint model is used to perform secondary rendering on a first rendered image obtained by performing two-dimensional rendering on the three-dimensional character model to obtain a second rendered image having a style matching the reference rendered image.
[0018] Fourth aspect, the present application further provides a computer-readable storage medium. On the computer-readable storage medium, there is a computer program stored thereon, and when the computer program is executed by a processor, the following steps are implemented: obtaining a plurality of initial rendering images of a target virtual character, where the initial rendering images are obtained by performing two-dimensional rendering on a three-dimensional character model of the target virtual character; for each initial rendering image, obtaining the prompt information of the initial rendering image being targeted, and obtaining the contour feature map of the initial rendering image being targeted; obtaining a plurality of pre-trained style constraint models and contour constraint models, where the style combinations indicated by the plurality of style constraint models fit a reference rendering image of the target virtual character; based on the prompt information and the contour feature map, generating a style rendering image through the plurality of style constraint models and the contour constraint models, where the contour of the style rendering image matches the initial rendering image being targeted and the style matches the reference rendering image; training a target style constraint model by using the style rendering image, and the target style constraint model is used to perform secondary rendering on a first rendering image obtained by performing two-dimensional rendering on the three-dimensional character model, so as to obtain a second rendering image having a style matching the reference rendering image.
[0019] Fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the following steps are implemented: obtaining a plurality of initial rendering images of a target virtual character, where the initial rendering images are obtained by performing two-dimensional rendering on a three-dimensional character model of the target virtual character; for each initial rendering image, obtaining the prompt information of the initial rendering image being targeted, and obtaining the contour feature map of the initial rendering image being targeted; obtaining a plurality of pre-trained style constraint models and contour constraint models, where the style combinations indicated by the plurality of style constraint models fit a reference rendering image of the target virtual character; based on the prompt information and the contour feature map, generating a style rendering image through the plurality of style constraint models and the contour constraint models, where the contour of the style rendering image matches the initial rendering image being targeted and the style matches the reference rendering image; training a target style constraint model by using the style rendering image, and the target style constraint model is used to perform secondary rendering on a first rendering image obtained by performing two-dimensional rendering on the three-dimensional character model, so as to obtain a second rendering image having a style matching the reference rendering image.
[0020] The above image processing method, apparatus, computer device, storage medium, and computer program product obtain multiple initial rendering images of a target virtual character. The initial rendering images are obtained by performing 2D rendering on the 3D character model of the target virtual character. For each initial rendering image, obtain the prompt information of the initial rendering image, and obtain the contour feature map of the initial rendering image. Obtain multiple pre-trained style constraint models and contour constraint models. The style combinations indicated by the multiple style constraint models fit the reference rendering image of the target virtual character. Based on the prompt information and the contour feature map, through the multiple style constraint models and contour constraint models, generate a style rendering image. The contour of the generated style rendering image matches the initial rendering image it is targeted at, and the style matches the reference rendering image. Use the style rendering image to train and obtain a target style constraint model. The target style constraint model is used to re-render the first rendering image obtained by performing 2D rendering on the 3D character model to obtain a second rendering image with a style matching the reference rendering image. Since the style combinations indicated by the multiple style constraint models fit the reference rendering image of the target virtual character, a target style constraint model of the target virtual character can be further trained through the style rendering images generated by the multiple style constraint models and contour constraint models. Furthermore, the target style constraint model can be used to re-render the 2.5D image of the target virtual character with a fixed style, improving the quality of the images obtained by 2.5D rendering. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is an application environment diagram of the image generation method in some embodiments;
[0022] Figure 2 It is a schematic diagram of the effect of the image processing method in some embodiments;
[0023] Figure 3 It is a schematic flowchart of the image generation method in some embodiments;
[0024] Figure 4 It is a schematic diagram of the initial rendering image in some other embodiments;
[0025] Figure 5 It is a schematic diagram of the generated image of the LoRA model in some embodiments;
[0026] Figure 6 It is a schematic diagram of the reference rendering image of the initial rendering image in some embodiments;
[0027] Figure 7 It is a schematic diagram of the style rendering image in some embodiments;
[0028] Figure 8 It is a schematic diagram of the effect of the second rendering image in some embodiments;
[0029] Figure 9 Schematic diagram of the effect of the second rendered image in some other embodiments;
[0030] Figure 10 Schematic diagram of the images generated with and without adding expression description information in some embodiments;
[0031] Figure 11 Schematic diagram of the comparison before and after face restoration in some embodiments;
[0032] Figure 12 Schematic diagram of the process of the image generation method in some other embodiments;
[0033] Figure 13 Block diagram of the structure of the image generation device in some embodiments;
[0034] Figure 14 Internal structure diagram of a computer device in some embodiments;
[0035] Figure 15 Internal structure diagram of a computer device in some other embodiments. Detailed implementation manners
[0036] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0037] The image processing method provided by the embodiments of the present application relates to technologies such as machine learning (ML) and computer vision (CV) in artificial intelligence, where:
[0038] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, a theory, method, technology and application system that perceives the environment, acquires knowledge and uses the knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence is also to study the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning and decision-making.
[0039] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0040] Computer Vision (CV) technology is a science that studies how to enable machines to "see". More specifically, it refers to using cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement in machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multi-dimensional data. The large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.
[0041] Machine Learning (ML) is an interdisciplinary subject involving multiple fields, including probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration. The pre-trained model is the latest development result of deep learning, which integrates the above technologies.
[0042] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0043] The image processing method provided in the embodiments of this application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 and the server 104 can communicate with each other through a network, such as a wired or wireless network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed in the cloud or on other servers. Among them, the terminal 102 can be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc.
[0044] For the image processing method provided in the embodiments of this application, the execution subject of each step can be a computer device, which refers to an electronic device with data calculation, processing, and storage capabilities. Taking Figure 1 the application environment shown as an example, the image processing method can be executed independently by the terminal 102, or independently by the server 104, or executed cooperatively by the interaction between the terminal 102 and the server 104. This application does not make any limitations in this regard.
[0045] Taking the interaction and cooperation between the terminal 102 and the server as an example, the terminal 102 can perform two-dimensional rendering on the three-dimensional character model of the target virtual character to obtain multiple initial rendering images, and send these initial rendering images to the server. Thus, the server 104 can obtain multiple initial rendering images of the target virtual character. For each initial rendering image, obtain the prompt information of the initial rendering image targeted, and obtain the contour feature map of the initial rendering image targeted. The server 104 can also obtain multiple pre-trained style constraint models and contour constraint models. The style combinations indicated by the multiple style constraint models are used to fit the reference rendering image of the target virtual character. Then, based on the prompt information and the contour feature map, through the multiple style constraint models and contour constraint models, generate a style rendering image. The contour of the style rendering image matches the initial rendering image targeted, and the style matches the reference rendering image. The server 104 can further use the style rendering image to train and obtain a target style constraint model. The server can send the target style constraint model to the terminal. The terminal can use the target style constraint model to re-render the first rendering image obtained by two-dimensional rendering based on the three-dimensional character model, and obtain a second rendering image with a style matching the reference rendering image.
[0046] The image processing method provided in the embodiments of the present application aims to provide a 3D-to-2D stylization solution for anime characters based on a generative model, that is, by introducing a generative model of a target style constraint model, to perform secondary rendering with a fixed style on a single three-dimensional character model, which can be used to improve the efficiency of animation studios. In practical applications, the studio provides relevant files obtained by preliminary 3D-to-2D rendering, and then can use the target style constraint model obtained by the image processing method provided in the embodiments of the present application for rendering efficiency improvement, and enhance aspects such as details and styles.
[0047] Reference Figure 2 , for some embodiments, it is a schematic diagram of the effect after processing by the image processing method provided in the embodiments of the present application. Among them, Figures (a) and (b) are the original images, Figure (c) is obtained by processing Figure (a) using the image processing method provided in the embodiments of the present application, and Figure (d) is obtained by processing Figure (b) using the image processing method provided in the embodiments of the present application.
[0048] In some embodiments, as Figure 3 shown, an image processing method is provided. Taking the example that this method is applied to Figure 1 the server 104 in
[0049] Step 302, obtain multiple initial rendering images of the target virtual character. The initial rendering images are obtained by performing two-dimensional rendering on the three-dimensional character model of the target virtual character.
[0050] Among them, the target virtual character can be a character in a game or an animation, and can be a virtual human or a virtual animal, etc.
[0051] Specifically, the server can obtain multiple initial rendered images, which are obtained by performing two-dimensional rendering on the three-dimensional character model of the target virtual character, that is, by performing three-to-two initial rendering based on the three-dimensional character model of the target virtual character, the initial rendered images can be obtained.
[0052] It should be noted that, due to the image processing method provided in the embodiments of the present application, the trained target style constraint model can be a generative model. Here, three-to-two can only perform some simple shader configurations without paying attention to details, styles, etc. Through the target style constraint model of the present application, it can be rendered again to generate a rendered image with a style matching the reference rendered image. Refer to Figure 4 , for a schematic diagram of the initial rendered image in some embodiments, by Figure 4 It can be seen that these initial rendered images lack details and have relatively poor light and shadow textures. It is necessary to perform secondary rendering on them to obtain style-rendered images and use them in the training of the next-step style constraint model.
[0053] A generative model is a type of machine learning model that can generate new data samples by learning the latent distribution of data and is widely used in fields such as image generation and text generation. Optionally, the generative model adopted in the embodiments of the present application can be a model based on Stable Diffusion (SD, stable diffusion model). Stable Diffusion is a generative model based on the diffusion process that generates new data samples by adding noise to the data space and gradually removing the noise. Its advantages include high generation quality, stable training, and the ability to be combined with other generative models to improve generation performance, and can generate pictures according to text descriptions, etc.
[0054] Exemplarily, the generative model adopted in the present application can be obtained by training the low-rank matrix of the stable diffusion model, specifically, it can be a LoRA model. The full name of LoRA is Low-Rank Adaptation of Large Language Models, which can be understood as a small patch on the SD model. After adding a certain LoRA, the SD model will have the capabilities corresponding to this LoRA. In the field of image generation, LoRA is often used to create a certain character, painting style, item, etc. Refer to Figure 5 , for examples of generated images of several LoRA models in civitai.
[0055] It can be understood that in some other embodiments, the generative model adopted in the present application can also be other models, such as IP-Adapter.
[0056] Step 304: For each initial rendered image, obtain the hint information for the targeted initial rendered image and obtain the contour feature map of the targeted initial rendered image.
[0057] Among them, the hint information is the information used to prompt the image generation process of the model, which can be the information describing the details, features, elements, styles, etc. of the image. The hint information can specifically be text information, such as the prompt word "prompt", or it can also be a hint image. The contour feature map is used to make the image generated by the model retain the layout, characters, etc. of the original image, and is obtained by extracting the contour of the initial rendered image. In specific applications, at least one of the canny algorithm or the softHED algorithm can be used to generate an edge line drawing of the initial rendered image, and then feature extraction is performed on the obtained edge line drawing to obtain the contour feature map.
[0058] Specifically, for each initial rendered image, the server can obtain the hint information of the rendered image and obtain the contour feature map of the initial rendered image, and then can re-render the initial rendered image according to the hint information and the contour feature map.
[0059] Exemplarily, for each initial rendered image, the server can use one or more of the label, positive prompt, or negative prompt of the initial rendered image as the hint information of the rendered image.
[0060] Step 306: Obtain multiple pre-trained style constraint models and contour constraint models, and the style combinations indicated by the multiple style constraint models fit the reference rendered image of the target virtual character.
[0061] Among them, different style constraint models can indicate different styles, and the style combinations indicated by the multiple style constraint models fit the reference rendered image of the target virtual character. The reference rendered image of the target virtual character can be the finalized portrait of the target virtual character provided by the painter. The reference rendered image is the target for re-rendering the image obtained by three-dimensional rendering to two-dimensional rendering in this application. Refer to Figure 6 , in some embodiments, Figure 4 Schematic diagram of the reference rendered image of the initial rendered image.
[0062] Specifically, the server can obtain multiple pre-trained style constraint models, and these multiple style constraint models respectively indicate different painting styles and are selected from the pre-trained style constraint model set according to the reference rendered image. In order to make the generated image retain the layout, characters, etc. of the original image, the server can also obtain a contour constraint model to make the generated characters as consistent as possible.
[0063] In some embodiments, the contour constraint model can be trained from a trainable copy cloned from the weights of the Stable Diffusion model. Exemplarily, the contour constraint model can be the ControlNet model.
[0064] Step 308: Based on the prompt information and the contour feature map, generate a style-rendered image through multiple style constraint models and the contour constraint model, where the contour of the style-rendered image matches the initial rendered image it targets, and the style matches the reference rendered image.
[0065] Specifically, for each initial rendered image, the server can, through multiple style constraint models and the contour constraint model, in combination with the SD model as the base model, generate a style-rendered model for the initial rendered image based on the prompt information of the initial rendered image and the contour feature map of the initial rendered image. In a specific application, the parameters of the multiple style constraint models can be injected into the SD model with different weights respectively to obtain an SD model with updated parameters. Input the prompt information of the initial rendered image into the SD model, and input the contour feature map of the initial rendered image into the contour control model. During the process of image generation by the SD model, guide the image generated by the SD model through the output of the contour control model, thereby generating a style-rendered image for the initial rendered image. The contour of the generated style-rendered model matches the contour of the initial rendered image, and the style matches the reference rendered image, that is, the obtained style-rendered image is an image obtained by re-rendering the initial rendered image, and the obtained style-rendered image fits the style of the reference rendered image provided by the painter.
[0066] For example, referring to Figure 7 , it is a schematic diagram of a style-rendered image in some embodiments. Among them, figure (a) is the initial rendered image, and figure (b) is the style-rendered image jointly generated by 4 different LoRA models.
[0067] Step 310: Train a target style constraint model using the style-rendered image. The target style constraint model is used to re-render the first rendered image obtained by two-dimensionally rendering a three-dimensional character model to obtain a second rendered image with a style matching the reference rendered image.
[0068] Specifically, the server can generate training samples using the obtained style-rendered images and perform model training through the training samples. When the training stop condition is met, a target style constraint model can be obtained. This style constraint model can, in combination with the contour constraint model and the SD model, re-render the first rendered image obtained by two-dimensionally rendering a three-dimensional character model to obtain a second rendered image with a style matching the reference rendered image.
[0069] Exemplarily, the target style constraint model here can also be a LoRA model. That is, in the embodiments of the present application, multiple LoRA models with different styles can be used to jointly generate images. However, multiple LoRAs often affect each other, resulting in problems such as out-of-control details in the generated images. In the embodiments of the present application, a small number of style-rendered images are generated through multiple LoRAs, and then the LoRA model is trained using these style-rendered images. The trained LoRA can solve the problem of out-of-control details in the image caused by the conflict between multiple LoRAs.
[0070] In the above image processing method, multiple initial rendered images of the target virtual character are obtained. The initial rendered image is obtained by performing two-dimensional rendering on the three-dimensional character model of the target virtual character. For each initial rendered image, the prompt information of the initial rendered image is obtained, and the contour feature map of the initial rendered image is obtained. Multiple pre-trained style constraint models and contour constraint models are obtained. The style combination indicated by the multiple style constraint models fits the reference rendered image of the target virtual character. Based on the prompt information and the contour feature map, through the multiple style constraint models and contour constraint models, a style-rendered image is generated. The contour of the generated style-rendered image matches the initial rendered image it is targeted at, and the style matches the reference rendered image. The target style constraint model is obtained by training using the style-rendered image. The target style constraint model is used to re-render the first rendered image obtained by performing two-dimensional rendering on the three-dimensional character model to obtain a second rendered image with a style matching the reference rendered image. Since the style combination indicated by the multiple style constraint models fits the reference rendered image of the target virtual character, the target style constraint model of the target virtual character can be further trained through the style-rendered images generated by the multiple style constraint models and contour constraint models. Furthermore, the two-and-a-half rendering image of the target virtual character can be re-rendered with a fixed style through the target style constraint model, improving the quality of the image obtained by two-and-a-half rendering.
[0071] In some embodiments, obtaining the target style constraint model by training using the style-rendered image includes: determining the style description information of the style-rendered image; based on the style description information, determining the text-image data pair including the style-rendered image; and training the style constraint model to be trained using the text-image data pair to obtain the target style constraint model.
[0072] Among them, the style description information refers to the information describing the style of the style-rendered image. Exemplarily, the server can input the style-rendered image into a text-image generation model to output the style description information corresponding to the style-rendered image. The text-image generation model can be an open-source text-image generation model. For example, the text-image generation model can be a Large Language Model Enhanced Vision-Language Understanding (miniGPT4) model, and the miniGPT4 model can be a model such as CLIP, BLIP, or BLIP2.
[0073] Specifically, for the obtained style-rendered images, the server can respectively obtain the style description information of these style-rendered images, and then form image-data pairs by combining each style-rendered image with its corresponding style description information. Furthermore, these text-image data pairs can be used as training samples to train the low rank matrices of the Stable Diffusion model. When the training stop condition is met, the target style constraint model is obtained.
[0074] In this embodiment, by determining the style description information of the style-rendered image, based on the style description information, determining the text-image data pair containing the style-rendered image, and using the text-image data pair to train the style constraint model to be trained to obtain the target style constraint model, the model training efficiency can be improved.
[0075] In some embodiments, based on the style description information, determining the text-image data pair containing the style-rendered image includes: determining the first expression description information of the style-rendered image, where the first expression description information is used to describe the expression of the target virtual character in the style-rendered image; and determining the text-image data pair containing the style-rendered image based on the style description information and the expression description information.
[0076] Among them, the expression description information is the information describing the expression of the target virtual character in the style-rendered image. For example, it can be expression description words such as "angry", "sad", "happy", etc.
[0077] Specifically, for each style-rendered image, the server can obtain the first expression description information of the style-rendered image, then combine the first expression description information with the style description information of the style-rendered image to obtain the text information of the style-rendered image, and form a text-image data pair by combining the obtained text information with the style-rendered image.
[0078] Exemplarily, the user can perform expression annotation on the style-rendered image in the terminal by manual annotation and send the annotated expression information to the server, or the server can input the style-rendered image into a pre-trained expression classification model to classify the expression of the style-rendered image to determine the expression category to which the style-rendered image belongs.
[0079] In some embodiments, for each style-rendered image, the server may also obtain other description information of the style-rendered image, such as the image feature information, element description information, etc. of the style-rendered image, and combine this information with the style description information and expression description information to form the text information of the style-rendered image.
[0080] In the above embodiments, adding the expression of the character in the training label can make the corresponding label be displayed after the character is generated in the picture, rather than integrating the expression as the default attribute of the character into the character image, so that it is possible to avoid all generated images of the character having the same expression.
[0081] In some embodiments, obtaining the prompt information for the targeted initial rendered image includes: obtaining the trained label model, inputting the targeted initial rendered image into the label model to obtain multiple initial labels for the targeted initial rendered image; screening the multiple initial labels to obtain multiple picture element prompt words for the targeted initial rendered image; obtaining multiple positive prompt words for the targeted initial rendered image, where the multiple positive prompt words are used to describe the image quality of the targeted initial rendered image; and determining the prompt information for the targeted initial rendered image based on the multiple picture element prompt words and the multiple positive prompt words.
[0082] Among them, the initial label is used to describe the characteristics of the target virtual character in the initial rendered image, and the characteristics of the target virtual character may specifically include the appearance characteristics of the target virtual character. The positive prompt words may include words for describing the image quality of the initial rendered image. The image quality may be, for example, high-definition image quality, high-fidelity image quality, blurred image quality, sketch image quality, hand-painted style, 3D image quality, 4K image quality, VR image quality, watercolor image quality, etc. The positive prompt words may also include words for indicating the style of the initial rendered image. The style may be, for example, realism, watercolor style, oil painting style, sketch style, cartoon style, photography style, abstractionism, surrealism, illustration style, a certain movie style, idealism, etc.
[0083] In this embodiment, the prompt information of the initial rendering image can be generated by combining the label model and manual annotation. Specifically, the server can obtain the trained label model, input the initial rendering image into the label model to obtain multiple initial labels of the initial rendering image, and then send the obtained initial labels to the terminal. After the terminal displays the initial labels, the user screens the initial labels, deletes the obviously incorrect labels, and sends the deleted labels to the server. Thus, the server can filter out these deleted labels and use the remaining labels as the prompt words for the picture elements of the initial rendering image. Further, the user can also add prompt words for describing the image quality of the initial rendering image through the terminal and send these prompt words to the server. Thus, the server can obtain multiple positive prompt words for the initial rendering image. The server combines the picture element prompt words and the positive prompt words to form the prompt information of the initial rendering image.
[0084] In a specific application, the label model can specifically be a label inference model, such as deepboru. Continuing to refer to Figure 7 , for Figure 7 the label model shown in Figure (a) therein, the following initial labels can be generated through the label model: girl, single, long hair, shorts, gray background, red hair, navel, simple background, jacket, boots, standing, full body, upper body, socks, jewelry, ultra-shorts, looking at the audience, yellow jacket, necklace. After the initial labels are screened, they are retained as the prompts for the picture elements.
[0085] In addition, some words indicating style and image quality can be added to improve the picture effect, such as: masterpiece, best quality, detail, red eyes, no bangs, masterpiece, side light, (exquisite detail beautiful eyes: 1.2), high saturation, color, masterpiece, portrait, realistic, shiny skin. These words added for improving the image quality are the positive prompt words.
[0086] In the above embodiment, by obtaining the initial labels and screening to determine the picture prompt words, further obtaining the positive prompt words for describing the image quality, and determining the prompt information of the targeted initial rendering image based on multiple picture element prompt words and multiple positive prompt words, the accuracy of the prompt information can be improved.
[0087] In some embodiments, determining the prompt information of the targeted initial rendering image based on multiple picture element prompt words and multiple positive prompt words includes: obtaining multiple negative prompt words of the targeted initial rendering image, where the multiple negative prompt words are used to describe the expected missing features of the targeted initial rendering image; and determining the prompt information of the targeted initial rendering image based on the multiple picture element prompt words, the multiple positive prompt words, and the multiple negative prompt words.
[0088] Among them, the negative prompt is used to describe the expected missing features of the initial rendered image being targeted. The expected missing features are those that are not desired to appear in the generated image, specifically including features that are generally not desired to appear and features that are not desired to appear in terms of details. Features that are generally not desired to appear include, for example, low resolution, dim, distorted, noisy, overexposed, underexposed, vulgar, deformed, and rigid. Features that are not desired to appear in terms of details include, for example, bad hands, long necks, bad feet, bad heads, bad eyes, bad eyebrows, and missing noses.
[0089] In this embodiment, for each initial rendered image, the user can add a negative prompt to the initial rendered image through the terminal. After sending the negative prompt to the server, the server can obtain the negative prompt of the initial rendered image. Furthermore, the server can combine the element prompt, positive prompt, and negative prompt of the initial rendered image to obtain the prompt information of the initial rendered image.
[0090] For example, taking Figure 7 the initial rendered image shown in Figure (a) as an example, the negative prompt examples are as follows: sketch, (worst quality: 2), (low quality: 2), (normal quality: 2), low quality, normal quality, ((monochrome)), (grayscale), skin spots, acne, skin blemishes, bad anatomy, (long hair: 1.4), depth negative, (obesity: 1.2), back-facing, looking, tilted head, low quality, bad anatomy, bad hands, text, error, missing fingers, extra fingers, fewer fingers, cropped, worst quality, low quality, normal quality, signature, watermark, username, blurred, bad feet, cropped, poorly drawn hands, poorly drawn face, mutation, deformation, worst quality, low quality, normal quality, signature, watermark, extra fingers, fewer digits, extra limbs, extra arms, extra legs, deformed limbs, fused fingers, too many fingers, long neck, cross-eyed, mutated hands, extremely low, bad body, bad proportions, thick proportions, text, error, missing fingers, missing arms, missing legs, extra digits, extra arms, extra legs, extra feet.
[0091] In the above embodiment, since the negative prompt of the initial rendered image is obtained, the prompt information of the initial rendered image determined based on the element prompt, positive prompt, and negative prompt can reduce the lack of the generated image's visual effect.
[0092] In some embodiments, the image processing method of the present application further includes: when obtaining a first rendered image obtained by two-dimensional rendering based on a three-dimensional character model, determining the prompt information of the first rendered image, and determining the contour feature map of the first rendered image; obtaining a target style constraint model and a contour constraint model; generating a second rendered image with a style matching that of a reference rendered image based on the prompt information of the first rendered image and the contour feature map of the first rendered image through the target style constraint model and the contour constraint model.
[0093] Among them, the first rendered image is obtained by two-dimensional rendering of the three-dimensional character model of the target virtual character, that is, through three-to-two initial rendering based on the three-dimensional character model of the target virtual character, an initial rendered image can be obtained. Here, three-to-two can only perform some simple shader parameter configurations without paying attention to details, styles, etc. It can be understood that the first rendered image here is different from the initial rendered image in the training process, and the first rendered image here can be an image obtained during the generation of an animated video.
[0094] Specifically, when obtaining a first rendered image obtained by two-dimensional rendering based on a three-dimensional character model, the server can determine the prompt information of the first rendered image. The prompt information can include picture element prompt words, positive prompt words, and negative prompt words, and at least one of the canny algorithm or the softHED algorithm is used to generate an edge line drawing of the first rendered image, and then feature extraction is performed on the obtained edge line drawing to obtain the contour feature map of the first rendered image. A second rendered image with a style matching that of a reference rendered image is generated based on the prompt information of the first rendered image and the contour feature map of the first rendered image through the target style constraint model and the contour constraint model.
[0095] For example, referring to Figure 8 , assuming that the first rendered image is the second picture in the first row of Figure 4 , then the effect diagram of the generated second rendered image can be as shown in figure (a) of Figure 8 . It can be seen that the second rendered image has added a lot of details and styles compared to the first rendered image.
[0096] In the above embodiments, by obtaining the prompt information and the contour feature map of the first rendered image, and generating a second rendered image with a style matching that of a reference rendered image based on the prompt information of the first rendered image and the contour feature map of the first rendered image through the target style constraint model and the contour constraint model, since the first rendered image can be re-rendered through the target style constraint model, the image quality of the first rendered image can be improved.
[0097] In some embodiments, based on the prompt information of the first rendered image and the contour feature map of the first rendered image, a second rendered image with a style matching that of the reference rendered image is generated through a target style constraint model and a contour constraint model, including: obtaining a plurality of pre-trained style constraint models; based on the prompt information of the first rendered image and the contour feature map of the first rendered image, generating a second rendered image with a style matching that of the reference rendered image through the plurality of pre-trained style constraint models, the target style constraint model, and the contour constraint model.
[0098] Specifically, during the process of re-rendering the initially rendered first rendered image, the server can also obtain a plurality of pre-trained style constraint models. Here, the plurality of style constraint models can be the models used during the training of the target style constraint model. Then, the model parameters of these style constraint models and the model parameters of the target style constraint model are injected into the StableDiffusion model to obtain an updated StableDiffusion model. Furthermore, the prompt information of the first rendered image can be input into this StableDiffusion model, and the contour feature map of the first rendered image can be input into the contour constraint model. Image generation is performed through the StableDiffusion model, and the output of the contour constraint model is used to perform contour constraint on the image generated by the StableDiffusion model, so that the target virtual character in the image generated by the StableDiffusion model is consistent with that of the first rendered image, thereby obtaining a second rendered image corresponding to the first rendered image, and this second rendered image has a style matching that of the reference rendered image.
[0099] In this embodiment, during the process of re-rendering the initially rendered first rendered image, image generation can also be jointly performed in combination with the plurality of style constraint models used during the training process to further improve the quality of the image obtained by 3D-to-2D conversion.
[0100] In some embodiments, based on the prompt information of the first rendered image and the contour feature map of the first rendered image, a second rendered image with a style matching that of the reference rendered image is generated through a plurality of pre-trained style constraint models, a target style constraint model, and a contour constraint model, including: injecting the model parameters of the plurality of pre-trained style constraint models and the target style constraint model into the pre-trained image generation model with different weights respectively to obtain a target image generation model, where the weight corresponding to the target style constraint model is greater than the weights corresponding to the plurality of pre-trained style constraint models respectively; inputting the prompt information of the first rendered image into the target image generation model, and inputting the contour feature map of the first rendered image into the contour constraint model; performing image generation through the target image generation model, and performing contour constraint on the image generated by the target image generation model through the output of the contour constraint model to obtain a second rendered image with a style matching that of the reference rendered image.
[0101] Specifically, different weights can be set for the model parameters of multiple pre-trained style constraint models and the model parameters of the target style constraint model respectively. Among them, the weight corresponding to the target style constraint model is greater than the weights corresponding to the multiple pre-trained style constraint models respectively, so that the target style constraint model plays a dominant role in the process of image generation. Thus, the server can inject the model parameters of the multiple pre-trained style constraint models and the model parameters of the target style constraint model into the pre-trained image generation model according to their respective weights to obtain the target image generation model. Then, input the prompt information of the first rendered image into the target image generation model, and input the contour feature map of the first rendered image into the contour constraint model; perform image generation through the target image generation model, and perform contour constraint on the image generated by the target image generation model through the output of the contour constraint model to obtain the second rendered image with a style matching the reference rendered image. Here, the pre-trained image generation model can be a Stable Diffusion model, that is, the SD model.
[0102] For example, continue to refer to Figure 8 , assuming that the first rendered image is Figure 4 the second picture in the first row of Figure 8 , then the effect diagram of the generated second rendered image can be as shown in Figure (b) of
[0103] Refer to Figure 9 , assuming that the first rendered image is Figure 9 Figure (a) of Figure 9 , and the second rendered image obtained by using multiple pre-trained style constraint models and the target style constraint model is Figure (b) of
[0104] In the above embodiment, by injecting the model parameters of multiple pre-trained style constraint models and the model parameters of the target style constraint model into the pre-trained image generation model with different weights respectively, and the weight corresponding to the target style constraint model is greater than the weights corresponding to the multiple pre-trained style constraint models respectively, it is possible to avoid conflicts among multiple style constraint models while further improving the quality of the image obtained by 2D-rendering from 3D models.
[0105] In some embodiments, determining the prompt information for the first rendered image includes: obtaining the initial prompt information preset for the target virtual character; determining the second expression description information for the first rendered image, where the second expression description information is used to describe the expression of the target virtual character in the first rendered image; and adding the second expression description information to the initial prompt information to obtain the prompt information for the first rendered image.
[0106] In this embodiment, to avoid distortion of the expression in the generated image, for example, an angry expression becoming peaceful, when generating the prompt information for the first rendered image, the server can determine the second expression description information for the first rendered image and add the second expression description information to the initial prompt information to obtain the prompt information for the first rendered image.
[0107] Exemplarily, the user can perform expression annotation on the first rendered image in the terminal by manual annotation and send the annotated expression information to the server, or the server can input the first rendered image into a pre-trained expression classification model to classify the expression of the first rendered image to determine the expression category to which the first rendered image belongs.
[0108] For example, referring to Figure 10 , it is a schematic diagram of the images generated with and without adding expression description information. Among them, figure (a) is the image generated without adding expression description information, and figure (b) is the image generated with the expression description information of "angry and furious" added. It can be seen that the image generated with the added expression description information is more vivid.
[0109] In the above embodiment, by determining the second expression description information for the first rendered image, where the second expression description information is used to describe the expression of the target virtual character in the first rendered image, and adding the second expression description information to the initial prompt information to obtain the prompt information for the first rendered image, it is possible to avoid distortion of the expression in the generated image and further improve the quality of the image obtained by three-dimensional rendering to two-dimensional.
[0110] In some embodiments, the image processing method of the present application further includes: when the second rendered image is obtained, performing face detection on the second rendered image; intercepting the detected face region from the second rendered image and magnifying the intercepted face region to obtain a face image; redrawing the face image to obtain a redrawn face image; and after shrinking the redrawn face image to match the size of the face region, pasting the obtained face image back to the face region to obtain a repaired second rendered image.
[0111] Considering that the generated image may experience face distortion during re - rendering, especially when the face is small. Therefore, in this embodiment, after obtaining the second rendered image, face repair can be performed. Specifically, the server can detect the face in the second rendered image, intercept the detected face area from the second rendered image, enlarge the intercepted face area to obtain a face image, redraw the face image to obtain a redrawn face image, reduce the redrawn face image to match the size of the face area, and then paste the obtained face image back to the face area to obtain the repaired second rendered image.
[0112] For example, referring to Figure 11 , it is a comparison schematic diagram before and after face repair. Among them, Figure (a) is the image before face repair, and Figure (b) is the image after face repair. It can be seen that after repair using the method of this embodiment, the distorted face can be restored to normal.
[0113] In the above - mentioned embodiment, since the face in the second rendered image can be detected, the detected face area can be intercepted from the second rendered image, the intercepted face area can be enlarged to obtain a face image, the face image can be redrawn to obtain a redrawn face image, the redrawn face image can be reduced to match the size of the face area, and then the obtained face image can be pasted back to the face area to obtain the repaired second rendered image, the distorted face can be repaired, further improving the quality of the image obtained by three - dimensional to two - dimensional rendering.
[0114] In some specific embodiments, as Figure 12 shown, an image - processing method is provided. Taking the server 104 in Figure 1 as an example for illustration, it includes the following steps:
[0115] Step 1202, obtain multiple initial rendered images of the target virtual character.
[0116] Among them, the initial rendered image is obtained by performing two - dimensional rendering on the three - dimensional character model of the target virtual character.
[0117] Step 1204, for each initial rendered image, obtain the prompt information of the initial rendered image targeted, and obtain the contour feature map of the initial rendered image targeted.
[0118] Among them, obtaining the prompt information for the targeted initial rendered image includes: obtaining a trained label model, inputting the targeted initial rendered image into the label model to obtain multiple initial labels for the targeted initial rendered image; screening the multiple initial labels to obtain multiple scene element prompt words for the targeted initial rendered image; obtaining multiple positive prompt words for the targeted initial rendered image, where the multiple positive prompt words are used to describe the image quality of the targeted initial rendered image; obtaining multiple negative prompt words for the targeted initial rendered image, where the multiple negative prompt words are used to describe the expected missing features of the targeted initial rendered image; and determining the prompt information for the targeted initial rendered image based on the multiple scene element prompt words, the multiple positive prompt words, and the multiple negative prompt words.
[0119] Step 1206: Obtain multiple pre-trained style constraint models and contour constraint models. The style combinations indicated by the multiple style constraint models fit the reference rendered image of the target virtual character.
[0120] Step 1208: Based on the prompt information and the contour feature map, generate a style rendered image through the multiple style constraint models and the contour constraint models.
[0121] Among them, the contour of the style rendered image matches the targeted initial rendered image, and the style matches the reference rendered image.
[0122] Step 1210: Determine the style description information of the style rendered image and determine the first expression description information of the style rendered image.
[0123] Among them, the first expression description information is used to describe the expression of the target virtual character in the style rendered image.
[0124] Step 1212: Based on the style description information and the expression description information, determine the graphic and text data pair containing the style rendered image.
[0125] Step 1214: Use the graphic and text data pair to train the style constraint model to be trained to obtain the target style constraint model.
[0126] Step 1216: When obtaining the first rendered image obtained by two-dimensional rendering based on the three-dimensional character model, determine the prompt information of the first rendered image and determine the contour feature map of the first rendered image.
[0127] Among them, determining the prompt information of the first rendered image includes: obtaining the initial prompt information preset for the target virtual character; determining the second expression description information of the first rendered image, where the second expression description information is used to describe the expression of the target virtual character in the first rendered image; adding the second expression description information to the initial prompt information to obtain the prompt information of the first rendered image. Among them, the initial prompt information may include multiple picture element prompt words, multiple positive prompt words, and multiple negative prompt words.
[0128] Step 1218, obtain the target style constraint model, the contour constraint model, and multiple pre-trained style constraint models.
[0129] Step 1220, inject the model parameters of the multiple pre-trained style constraint models and the target style constraint model into the pre-trained image generation model with different weights respectively to obtain the target image generation model.
[0130] Among them, the weight corresponding to the target style constraint model is greater than the weights corresponding to the multiple pre-trained style constraint models respectively.
[0131] Step 1222, input the prompt information of the first rendered image into the target image generation model, and input the contour feature map of the first rendered image into the contour constraint model.
[0132] Step 1224, perform image generation through the target image generation model, and perform contour constraint on the image generated by the target image generation model through the output of the contour constraint model to obtain a second rendered image with a style matching the reference rendered image.
[0133] Step 1226, perform face detection on the second rendered image, intercept the detected face area from the second rendered image, and magnify the intercepted face area to obtain a face image.
[0134] Step 1228, redraw the face image to obtain a redrawn face image. After shrinking the redrawn face image to match the size of the face area, post the obtained face image back to the face area to obtain a repaired second rendered image.
[0135] In some embodiments, the present application further provides an application scenario. In this application scenario, the image processing method provided by the embodiments of the present application is applied to perform stylized rendering after the initial 3D to 2D rendering of the animation. Here, since the requirements for the initial 3D to 2D rendering are simple, its Shader configuration is relatively simple and does not require excessive manpower, so it will not be elaborated here.
[0136] In this application scenario, the image processing method of this application is generally described in four parts, namely: 1) generation of style-rendered images; 2) training of character style Lora; 3) secondary rendering of the picture based on ControlNet; 4) repair of character facial expressions. The following is an introduction to each part respectively:
[0137] 1. Generation of style-rendered images assisted by SD
[0138] First, some initial rendered images are obtained through preliminary 3D to 2D rendering of the 3D character model. Then, appropriate text prompts, the Stable Diffusion base model, LoRA for controlling the painting style and details, and ControlNet are set to make the final rendered makeup photo more suitable for the character image. Among them:
[0139] The text prompt is the instruction for controlling Stable Diffusion to generate images. The prompt describes details, features, elements, styles, etc. that the image should have. A good prompt should not only include the elements of the picture but also descriptions of details such as style. In the process of generating style-rendered images, since the number of required style-rendered images is small, the text prompt can be generated by combining the picture tag inference model and manual annotation. Specifically, first, the picture is tagged with a picture label model (open source, such as deepboru, etc.). After the tags are screened, they are retained as the prompts for the picture elements. In addition, some words indicating style and image quality can be added to improve the picture effect. In addition, negative prompts can be added to reduce the lack of picture effects. Such prompts are used to prevent the picture from generating similar descriptions. Thus, a text prompt composed of picture element prompts + positive image quality prompts + negative image quality prompts can be obtained.
[0140] After adding a certain LoRA model to the SD model, the SD model will have the capabilities corresponding to this LoRA model. In the field of image generation, LoRA models are often used to create a certain character, painting style, item, etc. In this application scenario, according to the painting styles indicated by each LoRA model in the LoRA model set and the artist's drawings (i.e., the reference rendered images above), it can be decided which LoRA models to use to correct the generated picture. This part requires continuous debugging of the painting style combinations to confirm which LoRA models to finally select.
[0141] In order to make the generated pictures retain information such as the layout and characters of the original picture, it is far from enough to only use text prompts. It is necessary to introduce the ControlNet model to make the generated characters as consistent as possible. In this application scenario, the ControlNet model mainly uses Canny and SoftHED for control.
[0142] Finally, the artist confirms which set of prompt + LoRA model + ControlNet model produces the effect of the character's style. Then, style-rendered images can be batch-generated from the initial rendered images. After that, the artist helps with the screening. After obtaining 50 style-rendered images, the next step, assisting in the training of the LoRA model, can begin.
[0143] 2. Training of the Character Stylization LoRA Model
[0144] First, label the obtained style-rendered images, including determining the style description information and adding expression description information to avoid all generated character images having the same expression. After obtaining the text-image data pairs, put them into the training script for training.
[0145] It should be noted that although it was mentioned earlier that the painting style and elements of the picture can be controlled by combining multiple LoRA models. However, multiple LoRA models often affect each other, resulting in problems of out-of-control details in the generated pictures. Therefore, we need to first generate a small number of style-rendered images in the first step, and then train the LoRA model through these style-rendered images to solve the problem of out-of-control details in the picture caused by the conflict between multiple LoRA models.
[0146] 3. Secondary Rendering of the Picture Based on the ControlNet Model
[0147] The process is similar to that of generating style-rendered images in the first step, that is, for each picture, a set of positive image quality prompts, negative image quality prompts, and picture element prompts can be preset. In this step, the generation of these three types of prompts is the same as that in the first step. The control of the picture here is also guided by the ControlNet model, which is the same as in the first step. The main differences from the first step are in the selection and use of the LoRA model and the size of the generated images.
[0148] First, regarding the use of the LoRA model: When generating style-rendered images, only multiple selected LoRA models need to be used, while in this step, the trained LoRA model and the multiple previously selected LoRA models are used together for image generation.
[0149] Secondly, regarding the issue of image size. When generating style-rendered images, since the generated images are required for LoRA model training, to ensure the training efficiency, the image size is generally set at 512*512, and there is also 768*768, but the resolution will not exceed 1024. When performing animation shot rendering in this step, images of 1920*1080 need to be generated. Therefore, minor debugging adjustments need to be made to the parameters of the ControlNet model. Specifically, the threshold of the canny operator of the preprocessor corresponding to the ControlNet model used to generate the contour feature map can be adjusted so that all the edges that can maintain control can be detected at a size of 1920. For example, the image size in Figure (a) is the same as that in the training process, and the image size in Figure (b) is adjusted to 1080. It can be seen that in the case of a size of 1080, the generated images have better effects in details (hands, text).
[0150] 4. Repair of the facial expressions of the character
[0151] After the generation in the third step, a series of rendered video frames of the character can be obtained. However, since the human element prompts used in the third step are all the same, it will cause the distortion of the character's expression. For example, an originally angry expression becomes peaceful. Therefore, the text prompt of this video needs to be modified, mainly by adding expression-related prompts.
[0152] In addition, the generated images may also have the problem of face breakdown, especially when the face is small. Therefore, after obtaining the images for secondary rendering, the Adetailer tool can also be used for face repair.
[0153] In the above embodiments, the 3D rendering to 2D style of anime characters has the following advantages compared with the traditional 3D rendering to 2D based on Shader parameters: 1. Cost reduction and efficiency improvement: With the assistance of the generation model, the repetitive work of painters is greatly reduced. 2. Fast style switching: Characters of different styles can be quickly generated through this solution. 3. Strong stability: Compared with ordinary SD, this solution shows better stability in terms of character details, expressions, etc. In actual business use, in the redrawing link, the AI intelligent redrawing is 20 times more efficient than manual work.
[0154] It should be understood that although the steps in the flowcharts involved in the above embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0155] Based on the same inventive concept, the embodiments of the present application further provide an image generation device for implementing the above-mentioned image generation method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following image generation devices can refer to the limitations on the image generation method in the above text, and will not be repeated here.
[0156] In some embodiments, as Figure 13 shown, an image generation device 1300 is provided, including:
[0157] An initial rendering image acquisition module 1302, configured to acquire a plurality of initial rendering images of a target virtual character, and the initial rendering images are obtained by performing two-dimensional rendering on the three-dimensional character model of the target virtual character;
[0158] A prompt information acquisition module 1304, configured to, for each initial rendering image, acquire the prompt information of the initial rendering image targeted, and acquire the contour feature map of the initial rendering image targeted;
[0159] A model acquisition module 1306, configured to acquire a plurality of pre-trained style constraint models and contour constraint models, and the style combinations indicated by the plurality of style constraint models fit the reference rendering image of the target virtual character;
[0160] A style rendering image generation module 1308, configured to generate a style rendering image based on the prompt information and the contour feature map through the plurality of style constraint models and the contour constraint models, and the contour of the style rendering image matches the initial rendering image targeted, and the style matches the reference rendering image;
[0161] The model training module 1310 is used to train a target style constraint model by using stylized rendering images. The target style constraint model is used to re-render a first rendered image obtained by two-dimensionally rendering a three-dimensional character model to obtain a second rendered image with a style matching that of a reference rendered image.
[0162] The above image generation device obtains a plurality of initial rendered images of a target virtual character, which are obtained by two-dimensionally rendering the three-dimensional character model of the target virtual character. For each initial rendered image, the prompt information of the initial rendered image is obtained, and the contour feature map of the initial rendered image is obtained. A plurality of pre-trained style constraint models and contour constraint models are obtained. The style combinations indicated by the plurality of style constraint models fit the reference rendered image of the target virtual character. Based on the prompt information and the contour feature map, through the plurality of style constraint models and contour constraint models, stylized rendering images are generated. The contours of the generated stylized rendering images match the initial rendered images they correspond to, and the styles match the reference rendered image. The target style constraint model is trained by using the stylized rendering images. The target style constraint model is used to re-render a first rendered image obtained by two-dimensionally rendering a three-dimensional character model to obtain a second rendered image with a style matching that of the reference rendered image. Since the style combinations indicated by the plurality of style constraint models fit the reference rendered image of the target virtual character, the target style constraint model of the target virtual character can be further trained by the stylized rendering images generated by the plurality of style constraint models and contour constraint models. Furthermore, the two-and-a-half rendering image of the target virtual character can be re-rendered with a fixed style through the target style constraint model, improving the quality of the images obtained by two-and-a-half rendering.
[0163] In some embodiments, the model training module is further used to determine the style description information of the stylized rendering image; based on the style description information, determine the graphic and text data pair including the stylized rendering image; and use the graphic and text data pair to train the style constraint model to be trained to obtain the target style constraint model.
[0164] In some embodiments, the model training module is further used to determine the first expression description information of the stylized rendering image, where the first expression description information is used to describe the expression of the target virtual character in the stylized rendering image; and based on the style description information and the expression description information, determine the graphic and text data pair including the stylized rendering image.
[0165] In some embodiments, the prompt information acquisition module is further configured to: obtain a trained label model, input the initial rendering image thereto into the label model, and obtain multiple initial labels of the initial rendering image; screen the multiple initial labels to obtain multiple scene element prompt words of the initial rendering image; obtain multiple positive prompt words of the initial rendering image, where the multiple positive prompt words are used to describe the image quality of the initial rendering image; and determine the prompt information of the initial rendering image based on the multiple scene element prompt words and the multiple positive prompt words.
[0166] In some embodiments, the prompt information acquisition module is further configured to: obtain multiple negative prompt words of the initial rendering image, where the multiple negative prompt words are used to describe the expected missing features of the initial rendering image; and determine the prompt information of the initial rendering image based on the multiple scene element prompt words, the multiple positive prompt words, and the multiple negative prompt words.
[0167] In some embodiments, the above device further includes: a secondary rendering module, configured to, when obtaining a first rendering image obtained by performing two-dimensional rendering based on a three-dimensional character model, determine the prompt information of the first rendering image and determine the contour feature map of the first rendering image; obtain a target style constraint model and a contour constraint model; and generate a second rendering image with a style matching that of a reference rendering image through the target style constraint model and the contour constraint model based on the prompt information of the first rendering image and the contour feature map of the first rendering image.
[0168] In some embodiments, the secondary rendering module is further configured to obtain multiple pre-trained style constraint models; and generate a second rendering image with a style matching that of a reference rendering image through the multiple pre-trained style constraint models, the target style constraint model, and the contour constraint model based on the prompt information of the first rendering image and the contour feature map of the first rendering image.
[0169] In some embodiments, the secondary rendering module is further configured to respectively inject the model parameters of the multiple pre-trained style constraint models and the target style constraint model into a pre-trained image generation model with different weights to obtain a target image generation model, where the weight corresponding to the target style constraint model is greater than the weights corresponding to the multiple pre-trained style constraint models respectively; input the prompt information of the first rendering image into the target image generation model, and input the contour feature map of the first rendering image into the contour constraint model; perform image generation through the target image generation model, and perform contour constraint on the image generated by the target image generation model through the output of the contour constraint model to obtain a second rendering image with a style matching that of a reference rendering image.
[0170] In some embodiments, the secondary rendering module is further configured to obtain initial prompt information preset for the target virtual character; determine second expression description information of the first rendered image, where the second expression description information is used to describe the expression of the target virtual character in the first rendered image; and add the second expression description information to the initial prompt information to obtain prompt information of the first rendered image.
[0171] In some embodiments, the above device further includes: a face restoration module, configured to perform face detection on the second rendered image when the second rendered image is obtained; intercept the detected face region from the second rendered image, and enlarge the intercepted face region to obtain a face image; redraw the face image to obtain a redrawn face image; and after shrinking the redrawn face image to match the size of the face region, post the obtained face image back to the face region to obtain a restored second rendered image.
[0172] In some embodiments, the multiple style constraint models are respectively obtained by training the low-rank matrix of the StableDiffusion model, and the contour constraint model is obtained by training a trainable copy cloned from the weights of the StableDiffusion model.
[0173] Each module in the above image generation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of the processor, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0174] In some embodiments, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 14 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the image data involved in this application. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements the image generation method of any embodiment of this application.
[0175] In some embodiments, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in Figure 15 . The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the image generation method of any embodiment of the present application. The display unit of the computer device is used to form a visually visible picture, which may be a display screen, a projection device, or a virtual reality imaging device. The display screen may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0176] Those skilled in the art can understand that Figure 14 , Figure 15 The structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0177] In some embodiments, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the above-mentioned image generation method are implemented.
[0178] In some embodiments, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned image generation method are implemented.
[0179] In some embodiments, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the above-mentioned image generation method are implemented.
[0180] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0181] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.
[0182] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.
[0183] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. An image processing method, characterized in that, The method includes: Obtaining a plurality of initial rendered images of a target virtual character, where the initial rendered images are obtained by performing two-dimensional rendering on a three-dimensional character model of the target virtual character; For each initial rendered image, obtaining the prompt information of the targeted initial rendered image, and obtaining the contour feature map of the targeted initial rendered image; Obtaining a plurality of pre-trained style constraint models and contour constraint models, where the style combinations indicated by the plurality of style constraint models fit a reference rendered image of the target virtual character; Based on the prompt information and the contour feature map, generating a style rendered image through the plurality of style constraint models and the contour constraint models, where the contour of the style rendered image matches that of the targeted initial rendered image and the style matches that of the reference rendered image; Training a target style constraint model using the style rendered image, where the target style constraint model is used to re-render a first rendered image obtained by performing two-dimensional rendering on the three-dimensional character model to obtain a second rendered image with a style matching that of the reference rendered image.
2. The method according to claim 1, characterized in that The training the target style constraint model using the style rendered image includes: Determining the style description information of the style rendered image; Based on the style description information, determining a graphic-text data pair including the style rendered image; Using the graphic-text data pair to train a style constraint model to be trained to obtain a target style constraint model.
3. The method according to claim 2, wherein The determining the graphic-text data pair including the style rendered image based on the style description information includes: Determining first expression description information of the style rendered image, where the first expression description information is used to describe the expression of the target virtual character in the style rendered image; Based on the style description information and the expression description information, determining a graphic-text data pair including the style rendered image.
4. The method according to claim 1, wherein The obtaining the prompt information of the targeted initial rendered image includes: Obtaining a trained label model, inputting the targeted initial rendered image into the label model to obtain a plurality of initial labels of the targeted initial rendered image; Filtering the plurality of initial labels to obtain a plurality of scene element prompt words of the targeted initial rendered image; Obtaining a plurality of positive prompt words of the targeted initial rendered image, where the plurality of positive prompt words are used to describe the image quality of the targeted initial rendered image; Based on the plurality of scene element prompt words and the plurality of positive prompt words, determining the prompt information of the targeted initial rendered image.
5. The method according to claim 4, wherein The determining the prompt information of the targeted initial rendered image based on the plurality of scene element prompt words and the plurality of positive prompt words includes: Obtaining a plurality of negative prompt words of the targeted initial rendered image, where the plurality of negative prompt words are used to describe the expected missing features of the targeted initial rendered image; Based on the plurality of scene element prompt words, the plurality of positive prompt words, and the plurality of negative prompt words, determining the prompt information of the targeted initial rendered image.
6. The method according to claim 1, wherein The method further includes: When the first rendered image obtained by two-dimensional rendering based on the three-dimensional character model is acquired, determine the prompt information of the first rendered image and determine the contour feature map of the first rendered image; Obtain the target style constraint model and the contour constraint model; Based on the prompt information of the first rendered image and the contour feature map of the first rendered image, generate a second rendered image with a style matching that of the reference rendered image through the target style constraint model and the contour constraint model.
7. The method according to claim 6, characterized in that, The generating, based on the prompt information of the first rendered image and the contour feature map of the first rendered image, a second rendered image with a style matching that of the reference rendered image through the target style constraint model and the contour constraint model includes: Obtain the multiple pre-trained style constraint models; Based on the prompt information of the first rendered image and the contour feature map of the first rendered image, generate a second rendered image with a style matching that of the reference rendered image through the multiple pre-trained style constraint models, the target style constraint model and the contour constraint model.
8. The method according to claim 6, wherein The generating, based on the prompt information of the first rendered image and the contour feature map of the first rendered image, a second rendered image with a style matching that of the reference rendered image through the multiple pre-trained style constraint models, the target style constraint model and the contour constraint model includes: Inject the model parameters of the multiple pre-trained style constraint models and the target style constraint model into a pre-trained image generation model with different weights respectively to obtain a target image generation model, where the weight corresponding to the target style constraint model is greater than the weights corresponding to the multiple pre-trained style constraint models respectively; Input the prompt information of the first rendered image into the target image generation model, and input the contour feature map of the first rendered image into the contour constraint model; Perform image generation through the target image generation model, and perform contour constraint on the image generated by the target image generation model through the output of the contour constraint model to obtain a second rendered image with a style matching that of the reference rendered image.
9. The method according to claim 6, wherein The determining the prompt information of the first rendered image includes: Obtain the initial prompt information preset for the target virtual character; Determine the second expression description information of the first rendered image, where the second expression description information is used to describe the expression of the target virtual character in the first rendered image; Add the second expression description information to the initial prompt information to obtain the prompt information of the first rendered image.
10. The method according to claim 1, characterized in that The method further includes: When the second rendered image is obtained, perform face detection on the second rendered image; Crop the detected face region from the second rendered image and magnify the cropped face region to obtain a face image; Redraw the face image to obtain a redrawn face image; After shrinking the redrawn face image to match the size of the face region, the obtained face image is pasted back to the face region to obtain a repaired second rendered image.
11. The method according to any one of claims 1 to 10, characterized in that, The multiple style constraint models are respectively obtained by training the low-rank matrix of the StableDiffusion model, and the contour constraint model is obtained by training a trainable copy cloned from the weights of the StableDiffusion model.
12. An image processing apparatus, characterized in that, The device includes: An initial rendered image acquisition module, configured to acquire multiple initial rendered images of a target virtual character, where the initial rendered images are obtained by performing two-dimensional rendering on a three-dimensional character model of the target virtual character; A prompt information acquisition module, configured to, for each initial rendered image, acquire the prompt information of the initial rendered image targeted thereby, and acquire the contour feature map of the initial rendered image targeted thereby; A model acquisition module, configured to acquire multiple pre-trained style constraint models and contour constraint models, where the style combinations indicated by the multiple style constraint models fit a reference rendered image of the target virtual character; A style rendered image generation module, configured to generate a style rendered image based on the prompt information and the contour feature map, through the multiple style constraint models and the contour constraint model, where the contour of the style rendered image matches the initial rendered image targeted thereby, and the style matches the reference rendered image; A model training module, configured to train and obtain a target style constraint model by using the style rendered image, where the target style constraint model is configured to perform secondary rendering on a first rendered image obtained by performing two-dimensional rendering on the three-dimensional character model, to obtain a second rendered image having a style matching the reference rendered image.
13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 11.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 11.