Image exterior drawing method and device, equipment and storage medium
By calling the multimodal model and image generation model, combined with the control of the stable diffusion model by the internal drawing control network ControlNet++, the problem of single information in the existing technology of Chinese and foreign drawing areas is solved, and vivid and natural image external drawing is realized. The generated external drawing image information is rich and the content is consistent.
Patent Information
- Application Number
- CN202510481801.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The prior art is difficult to achieve vivid and natural expansion of the original image, and the texture and other information of the outer drawing area are relatively single, making it difficult to generate an outer drawing area with rich information.
The pre-trained multimodal model is called to process the original image, generate image generation prompt words, combine the pre-trained image generation model, and use extended models such as the in-draw control network ControlNet++ to control the image generation process of the stable diffusion model to generate vivid and natural out-drawn images.
The information-rich outer drawing area is generated, and the original image content is basically retained, so that the image after the outer drawing can be used as the target image extended from the original image, and ultimately realizes the vivid and natural image outer drawing task.
Smart Images

Figure CN120014119A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image external drawing method, device, equipment and storage medium. Background Art
[0002] The Image Outpainting task refers to the image editing task of drawing new image content outside the given original image boundary to achieve image expansion.
[0003] Currently, image external painting can be performed by copying or splicing similar texture content outside the original image boundary based on the texture information in the original image. For example, when external painting a grassland image, a grassland texture sample can be extracted from the grassland image and directly or after reasonable deformation of the grassland texture sample, it can be spliced outside the image boundary to achieve image expansion; similar image blocks can also be obtained from the image library and the image expansion can be achieved by filling the obtained image blocks into the area that needs external painting.
[0004] However, the texture and other information of the external drawing area obtained based on the above scheme is relatively simple, and the existing image external drawing scheme is difficult to achieve a vivid and natural expansion of the original image. Summary of the invention
[0005] In view of the above problems, the present application provides an image external drawing method, device, equipment and storage medium to generate an information-rich external drawing area and realize a vivid and natural image external drawing task.
[0006] The specific plan is as follows:
[0007] The first aspect of the present application provides an image drawing method, comprising:
[0008] Calling a pre-trained multimodal model to process the original image to obtain an image generation prompt word, wherein the image generation prompt word includes a description text for generating an outer drawing area of the original image;
[0009] A pre-trained image generation model is called to generate an externally drawn image according to the image generation prompt word and the original image; wherein the image generation model is a stable diffusion model configured with an extended model, the extended model is a pre-trained internal drawing control network ControlNet++, and the internal drawing control network ControlNet++ uses the original image to control the image generation process of the stable diffusion model, so that the area corresponding to the original image in the generated image of the stable diffusion model tends to be consistent with the original image.
[0010] A second aspect of the present application provides an image drawing device, comprising:
[0011] A prompt word generation unit, used to call a pre-trained multimodal model to process the original image to obtain an image generation prompt word, wherein the image generation prompt word includes a description text for generating an outer drawing area of the original image;
[0012] An image generation unit is used to call a pre-trained image generation model to generate an externally drawn image according to the image generation prompt word and the original image; wherein the image generation model is a stable diffusion model configured with an extended model, the extended model is a pre-trained internal drawing control network ControlNet++, and the internal drawing control network ControlNet++ uses the original image to control the image generation process of the stable diffusion model so that the area corresponding to the original image in the generated image of the stable diffusion model tends to be consistent with the original image.
[0013] The third aspect of the present application provides an image drawing device, comprising at least one processor and a memory connected to the processor, wherein:
[0014] The memory is used to store computer programs;
[0015] The processor is used to execute the computer program to implement the image drawing method described in the first aspect.
[0016] A fourth aspect of the present application provides a storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the image drawing method described in the first aspect.
[0017] By means of the above technical solution, the present application first calls the pre-trained multimodal model to process the original image, obtains the image generation prompt word, and then calls the pre-trained image generation model to generate the externally drawn image according to the image generation prompt word and the original image. Since the image generation prompt word contains the description text of the externally drawn area used to generate the original image, the externally drawn area information is provided for the subsequent externally drawn image generation, which is helpful to generate the information-rich externally drawn area; on this basis, the image generation model is a stable diffusion model configured with an extended model, and the extended model is the pre-trained internal drawing control network ControlNet++. Since the internal drawing control network ControlNet++ uses the original image to control the image generation process of the stable diffusion model, the area corresponding to the original image in the generated image of the stable diffusion model can be made consistent with the original image, thereby basically retaining the original image content, so that the generated externally drawn image can be used as the target image extended by the external drawing of the original image, and finally a vivid and natural image external drawing task is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present application. Also, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:
[0019] Figure 1 A schematic diagram of a flow chart of an image drawing method provided in this application;
[0020] Figure 2 A structural schematic diagram of an image generation model is shown;
[0021] Figure 3 A schematic diagram of the noise prediction process based on the joint guidance method is shown;
[0022] Figure 4 A schematic diagram of an external drawing process is shown;
[0023] Figure 5 A schematic diagram of the structure of an image drawing device provided in this application;
[0024] Figure 6 A schematic diagram of the structure of an image drawing device provided in this application. DETAILED DESCRIPTION
[0025] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation mode of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0026] Figure 1 FIG. 1 is a flow chart of an image drawing method according to an embodiment of the present application. Figure 1 As shown, the method may include the following steps:
[0027] Step S101: call a pre-trained multimodal model to process the original image to obtain image generation prompt words.
[0028] The image generation prompt includes a description text for generating the outer drawing area of the original image, which can be inferred by a multimodal model based on the original image content, providing an information basis for subsequent image generation, especially image generation corresponding to the outer drawing area. Exemplarily, the multimodal model can be a multimodal large language model, such as GPT-4o, or other image understanding models, such as mini-CPM-v.
[0029] Step S102: calling a pre-trained image generation model to generate an externally drawn image based on the image generation prompt word and the original image.
[0030] Among them, the image generation model is a stable diffusion model configured with an extended model, and the extended model is a pre-trained internal drawing control network ControlNet++. The internal drawing control network ControlNet++ (referred to as internal drawing ControlNet++) uses the original image to control the image generation process of the stable diffusion model, so that the area corresponding to the original image in the generated image of the stable diffusion model tends to be consistent with the original image. Internal drawing ControlNet++ is obtained by optimizing internal drawing ControlNet. Internal drawing ControlNet++ has better performance in maintaining the consistency of the original image. By using the internal drawing ControlNet++ control generation method, there is no need to introduce additional consistency control modules (such as image prompt adapter IP-Adapter), thereby reducing the demand for computing resources and reducing the cost of task implementation to a certain extent. In a possible implementation, the above-mentioned stable diffusion model can refer to the text-to-image generation model SDXL-Lighting. It should be noted that the model performs well in terms of image quality and can generate images with higher resolution and richer details while maintaining good diversity and image-text matching. The model is also flexible in usage and can generate images in 1, 2, 4 or 8 steps. The image quality improves with the increase of inference steps. In addition, the model uses progressive adversarial distillation technology to generate high-quality and high-resolution images in 2 or 4 steps, reducing the computational cost and time by ten times. It can even generate images for time-sensitive applications in 1 step at the expense of some quality. Based on the above, the SDXL-Lighting model can significantly reduce the number of sampling steps while maintaining the generation quality, thereby improving the efficiency of image generation.
[0031] This application first calls a pre-trained multimodal model to process the original image, obtains image generation prompt words, and then calls the pre-trained image generation model to generate an externally painted image based on the image generation prompt words and the original image. Since the image generation prompt words contain the description text of the externally painted area used to generate the original image, the externally painted area information is provided for the subsequent externally painted image generation, which is helpful to generate information-rich externally painted areas; on this basis, the image generation model is a stable diffusion model configured with an extended model, and the extended model is a pre-trained internal painting control network ControlNet++. Since the internal painting ControlNet++ uses the original image to control the image generation process of the stable diffusion model, the area corresponding to the original image in the generated image of the stable diffusion model can be made consistent with the original image, thereby basically retaining the original image content, so that the generated externally painted image can be used as a target image extended by the external painting of the original image, and finally a vivid and natural image external painting task is achieved.
[0032] In one or more embodiments provided in the present application, the above step S101, calling a pre-trained multimodal model to process the original image to obtain an image generation prompt word, may include:
[0033] The original image and pre-configured image processing prompt words are input into the multimodal model so that the multimodal model outputs the image generation prompt words.
[0034] Among them, the image processing prompt words are used to instruct the multimodal model to recognize and understand the input image, to infer the external drawing area of the input image based on the recognized image content, and to generate and output prompt words based on the recognized image content and the inferred external drawing content.
[0035] The above step S101 can be called the prompt word inference stage. Through the above scheme, the present application uses a multimodal model to deeply understand the original image content and imagine the content of the outer drawing area, thereby obtaining rich description information supplement of the outer drawing area, providing rich reference information for applying the inner drawing model to solve the outer drawing problem.
[0036] Compared with the commonly used prompt word inference scheme that only understands the original image through the visual language model, this scheme can supplement the missing external drawing area information based on the method of obtaining the original image description to guide the generation of the diffusion model, thereby solving the problem of blurred content and poor image quality in the external drawing area to a certain extent. Even when facing large-scale external drawing tasks, this scheme can provide certain information support for large-area external drawing areas, so that the image generated prompt words obtained in the prompt word inference stage of this application can greatly improve the external drawing quality at a large external drawing scale and solve the problem of blurred external drawing areas to a great extent.
[0037] In one or more embodiments provided in the present application, the above-mentioned step S102, calling the pre-trained image generation model to generate the externally drawn image according to the image generation prompt word and the original image, may include the following steps:
[0038] Step S201: Fill the surrounding area of the original image with a preset color to generate an extended image that includes the original image and has the same size as the image after external drawing, as a control image.
[0039] Exemplarily, the above-mentioned preset color may be white or black; the control diagram may be centered on the center of the original image. In particular, if the external drawing target specifies the position of the original image in the externally drawn image, the position of the original image in the control diagram is consistent with the position of the original image (or the image area corresponding to the original image) in the externally drawn image.
[0040] Step S202, using the control diagram as a conditional input of the internal drawing control network ControlNet++, and using the image generation prompt word as a prompt input of the internal drawing control network ControlNet++ and the stable diffusion model, calling the configured stable diffusion model to generate the external drawing image based on the input original image.
[0041] For example, Figure 2 A schematic diagram of the structure of an image generation model is shown, which shows the structure of the stable diffusion model and the internal drawing ControlNet++, combined with Figure 2 As shown, the original image can be input as "input" to the stable diffusion model and the internal drawing ControlNet++, the image generation prompt word can be input as "prompt" to the stable diffusion model and the internal drawing ControlNet++, and the control chart can be input as "condition N" to the internal drawing ControlNet++.
[0042] In one or more embodiments provided in the present application, the stable diffusion model adopts a sampling guidance method that combines perturbation attention guidance and classifier-free guidance.
[0043] It should be noted that Perturbed-Attention Guidance (PAG) aims to capture structural information by considering the self-attention mechanism and gradually enhance the structure of synthetic samples during the denoising process. Specifically, PAG generates structurally degraded intermediate samples by replacing the selected self-attention map in the diffuse U-Net with the identity matrix, and guides the denoising process away from these degraded samples, such as by replacing the self-attention mask of the specified layer with the identity matrix, so that the mutual attention between different regions in the image is completely inactivated. It should be noted that self-attention is mainly responsible for the reconstruction of the overall structure of the image, and inactivating self-attention can lose supervision of structural details. PAG can achieve sample quality improvement in both unconditional and conditional settings without further training or integration of external modules. Classifier-Free Guidance (CFG) amplifies the influence of conditional signals (such as text conditions) by mixing the noisy prediction results of the conditional branch and the unconditional branch.
[0044] The joint guidance of PAG and CFG can be applied to the noise predicted by the U-Net of the diffusion model in each sampling step. For example, Figure 3 The schematic diagram of the noise prediction process based on the joint guidance method is shown. Figure 3 As shown, the prediction noise at time t can be expressed as: t =σ φ +λ CFG ×(σ text -σ φ )+λ PAG ×(σ text -σ perturbed ). In the formula, σ φ represents unconditional noise, that is, the noise prediction result obtained by injecting the text prompt word into the model with an empty string, that is, the result generated by the model without any guidance / guidance, and the output is usually very random; σ text Represents text conditional noise, that is, the noise prediction result obtained by inputting the text prompt word into the model, and when obtaining σ text During the process, all attention layers in U-Net are turned on normally, and the output usually follows the prompt word; σ perturbed Represents disturbance noise, that is, the noise prediction result obtained by inputting the text prompt word into the model, and when obtaining σ text In the process of λ, some attention layers in U-Net are inactivated; CFG Represents the weight coefficient for measuring the strength of CFG, through λCFG ×(σ text -σ φ ) can make the output image biased towards σ text This result of following the prompt word avoids σ φ This random result; PAG Represents the weight coefficient for measuring the strength of PAG, using λ PAG ×(σ text -σ perturbed ) can make the output image biased towards σ text This results in a good picture structure, avoiding σ perturbed This is the result of a messy picture structure.
[0045] Based on the above content, this application improves the generation quality of the structure in the external drawing area to a certain extent by introducing PAG in sampling. By combining PAG and CFG, the advantage of CFG in better following text prompt words and the advantage of PAG in generating images with better structures are combined, which helps to improve the overall quality of the image after external drawing.
[0046] In one or more embodiments provided in the present application, after generating the externally drawn image, the following steps may also be included:
[0047] Step S103: post-process the externally drawn image to obtain an externally drawn result of the original image.
[0048] The post-processing of the externally drawn image may include:
[0049] The image area in the externally drawn image corresponding to the original image is replaced with the original image, and image fusion processing is performed on the seam area between the externally drawn image and the original image.
[0050] Exemplarily, the above-mentioned image fusion processing may refer to fusing the joint edges of the original image and the externally drawn image by an alpha blending (alpha blending / composition) method, that is, linearly weighted fusion of the two original images and the externally drawn image at the joint edges to eliminate seams.
[0051] Specifically, a mask image can be generated first, in which the area corresponding to the original image is white (alpha=1) to indicate that the area is completely the original image, and the rest of the area is black (alpha=0) to indicate that the area is completely the image after external painting. At the seam, a part of the area is taken inward and a part of the area is taken outward to form a fusion area (also called a seam area or a transition area), and alpha is gradually reduced from the inside to the outside in this area (i.e., from 1.0 to 0.0) to indicate the gradual fusion of the original image and the image after external painting. Assuming that the original image is a rectangle with a width of W and a height of H, the span of the transition area in the width direction can be W / 8, which is composed of 1 / 16 of the width taken inward and 1 / 16 of the width taken outward; correspondingly, the span of the transition area in the length direction can be H / 8; for the process of alpha gradually attenuating from the inside to the outside in the transition area, it can be achieved by applying a filter to the mask image. The filter can be used to blur the white and black boundary of the mask image, and finally form a transition area. Exemplarily, the filter may be a mean filter box filter, which updates the alpha value of the transition area by calculating the average alpha value of the surrounding pixels, thereby achieving a smooth transition.
[0052] It should be noted that since the image generation model randomly samples from a distribution for image generation, it cannot guarantee that the area corresponding to the original image in the externally painted image is completely consistent with the original image. Based on this, the above scheme ensures the consistency of the original image by overlaying the original image to the corresponding position on the externally painted result, and eliminates the seams through image fusion, so that the seam area of the externally painted image can be naturally fused, thereby improving the image external painting quality.
[0053] In one or more embodiments provided in the present application, post-processing the externally drawn image may further include: before replacing the image area corresponding to the original image in the externally drawn image, performing the following steps:
[0054] Step A: Count the pixel values of the original image (which may be represented as I_original) and the image area corresponding to the original image in the outpainted image (which may be represented as I_outpaint).
[0055] Step B: Calculate the pixel value distribution of the image area corresponding to the original image in the original image and the image after external rendering in the RGB color space based on the kernel density estimation method.
[0056] Step C: establishing a pixel value mapping relationship between the original image and an image region in the externally drawn image corresponding to the original image based on the calculated pixel value distribution.
[0057] Specifically, the distribution P_original and P_outpaint of I_original and I_outpaint in the RGB color space can be estimated by the kernel density estimation method, including the distribution estimation of the three channels of R, G, and B. Then, the cumulative distribution function (CDF) of I_original and I_outpaint in the RGB space is calculated by P_original and P_outpaint, respectively, expressed as CDF_original and CDF_outpaint, and then the pixel value mapping relationship between I_original and I_outpaint is constructed by CDF_original and CDF_outpaint.
[0058] Step D: correcting the pixel values of the externally drawn image according to the pixel value mapping relationship, and continuing to execute the step of replacing the image area corresponding to the original image in the externally drawn image with the original image based on the corrected externally drawn image.
[0059] The above scheme establishes a pixel value mapping relationship between the original image and the corresponding area in the external drawing result through kernel density estimation, and corrects the external drawing image accordingly, eliminating the color difference between the external drawing image and the original image, and further ensuring the consistency between the external drawing image and the original image.
[0060] In one or more embodiments provided in the present application, post-processing the externally drawn image may further include:
[0061] Before replacing the image area in the externally painted image corresponding to the original image, a pre-trained image restoration model is called to perform restoration processing on the externally painted image, and based on the restored externally painted image, the step of replacing the image area in the externally painted image corresponding to the original image with the original image is continued.
[0062] Among them, the image restoration model includes a face restoration model and an image super-resolution model. The above-mentioned face restoration model may refer to a GFPGAN (Generative Facial Prior - Guided Facial Attribute Editing Network, GFPGAN) model, which utilizes the generative adversarial network GAN architecture. Specifically, the generator network of the model can perform restoration and enhancement operations on the input face image based on pre-trained face prior knowledge. In the model training stage, a large amount of high-quality face image data can be used to enable the generator to learn key features such as the structure and texture of the face. On this basis, the discriminator network can distinguish between the generated face and the real face, and optimize the generator performance through adversarial training between the generator and the discriminator. In addition, GFPGAN also utilizes the semantic information of facial components, such as the position and shape of the facial features, so that facial details can be restored more accurately, such as repairing facial blemishes, improving clarity, reconstructing facial details and textures, etc., which helps to generate high-quality, natural and realistic face images. In addition, the above-mentioned image super-resolution model may be the basic image restoration model Real-ESRGAN. For example, the GFPGAN model can be used to detect and repair faces to obtain a clearer face area image, and the Real-ESRGAN model can be used to perform whole-image repair to obtain a clearer external drawing whole image. Finally, the repaired face area image is used to replace the corresponding area in the repaired external drawing whole image to obtain a repaired external drawing image.
[0063] Based on the above content, this application uses a face restoration model to repair the faces in the external drawing results, and integrates an image super-resolution model on this basis to improve the overall picture clarity and ultimately improve the image external drawing quality.
[0064] Optionally, during post-processing, restoration processing may be performed first, followed by chromatic aberration correction processing, and finally replacement fusion processing.
[0065] Figure 4 The schematic diagram of the external drawing process of the present application scheme is illustrated, the original image, the control image C_image generated based on the original image, and the obtained external drawing result are shown in FIG. Figure 4As shown, the original image is taken from the open source anime dataset qkrwnstj / anime_dataset, and the image generation prompt word C_text obtained by inputting the original image into the multimodal model can include: "A high-resolution image of an anime-style short-haired woman standing in a lush green park. The woman (i.e., the aforementioned short-haired woman) is wearing a light and flowing dress and raising one hand as if checking something. In the background, tall city buildings can be seen under the clear sky, and the sun shines through the trees, creating a peaceful and peaceful atmosphere. The lighting is soft and natural, highlighting the calm and peaceful environment." The original image and C_text are input into the stable diffusion model, and the original image, C_image and C_text are input into the internal drawing control network ControlNet++, and finally the external drawing result of the original image is obtained after post-processing.
[0066] The image external drawing device provided in the embodiment of the present application is described below. The image external drawing device described below and the image external drawing method described above can refer to each other.
[0067] Figure 5 Schematic diagram of the structure of an image drawing device disclosed in an embodiment of the present application. Figure 5 As shown, the device may include:
[0068] A prompt word generation unit 11 is used to call a pre-trained multimodal model to process the original image to obtain an image generation prompt word, wherein the image generation prompt word includes a description text for generating an outer drawing area of the original image;
[0069] The image generation unit 12 is used to call a pre-trained image generation model to generate an externally drawn image according to the image generation prompt word and the original image; wherein the image generation model is a stable diffusion model configured with an extended model, and the extended model is a pre-trained internal drawing control network ControlNet++, and the internal drawing control network ControlNet++ uses the original image to control the image generation process of the stable diffusion model, so that the area corresponding to the original image in the generated image of the stable diffusion model tends to be consistent with the original image.
[0070] In one or more embodiments provided in the present application, the prompt word generation unit 11 calls the pre-trained multimodal model to process the original image to obtain the process of generating prompt words from the image, which may include:
[0071] The original image and pre-configured image processing prompt words are input into the multimodal model so that the multimodal model outputs the image generation prompt words; the image processing prompt words are used to instruct the multimodal model to recognize and understand the input image, infer the external drawing area of the input image based on the recognized image content, and generate and output prompt words based on the recognized image content and the inferred external drawing content.
[0072] In one or more embodiments provided in the present application, the process in which the image generation unit 12 calls the pre-trained image generation model to generate the rendered image according to the image generation prompt word and the original image may include:
[0073] Filling the surrounding area of the original image with a preset color to generate an extended image containing the original image and having the same size as the image after external drawing as a control image;
[0074] The control diagram is used as a conditional input of the internal drawing control network ControlNet++, and the image generation prompt word is used as a prompt input of the internal drawing control network ControlNet++ and the stable diffusion model, and the configured stable diffusion model is called to generate the external drawing image based on the input original image.
[0075] In one or more embodiments provided in the present application, the stable diffusion model adopts a sampling guidance method that combines perturbation attention guidance and classifier-free guidance.
[0076] In one or more embodiments improved in the present application, the device may further include: a post-processing unit, the post-processing unit being configured to post-process the externally drawn image after generating the externally drawn image to obtain an externally drawn result of the original image.
[0077] Based on the above, the process of the post-processing unit post-processing the externally drawn image may include:
[0078] The image area in the externally drawn image corresponding to the original image is replaced with the original image, and image fusion processing is performed on the seam area between the externally drawn image and the original image.
[0079] In one or more embodiments of the present application, the process of the post-processing unit post-processing the externally drawn image may further include:
[0080] Before replacing the image area corresponding to the original image in the externally drawn image, the following steps are performed:
[0081] Counting pixel values of image areas corresponding to the original image in the original image and the externally drawn image respectively;
[0082] Calculate the pixel value distribution of the image area corresponding to the original image in the original image and the externally drawn image in the RGB color space based on the kernel density estimation method;
[0083] Establishing a pixel value mapping relationship between the original image and an image region in the externally drawn image corresponding to the original image according to the calculated pixel value distribution;
[0084] The pixel values of the externally drawn image are corrected according to the pixel value mapping relationship, and based on the corrected externally drawn image, the step of replacing the image area in the externally drawn image corresponding to the original image with the original image is continued.
[0085] In one or more embodiments of the present application, the process of the post-processing unit post-processing the externally drawn image may further include:
[0086] Before replacing the image area corresponding to the original image in the externally painted image, calling a pre-trained image restoration model to perform restoration processing on the externally painted image, and continuing to perform the step of replacing the image area corresponding to the original image in the externally painted image with the original image based on the restored externally painted image;
[0087] Wherein, the image restoration model includes a face restoration model and an image super-resolution model.
[0088] The image external drawing device provided in the embodiment of the present application can be applied to image external drawing equipment, such as terminals with data processing capabilities: mobile phones, computers, etc. Optionally, Figure 6 The hardware structure diagram of the image drawing device is shown in FIG. Figure 6 ,The hardware structure of the image external drawing device may include: at least one processor 1, at least one communication interface 2, at least one memory 3 and at least one communication bus 4;
[0089] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 communicate with each other through the communication bus 4;
[0090] The processor 1 may be a central processing unit CPU, or an application-specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0091] The memory 3 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), etc., such as at least one disk memory;
[0092] The memory is used to store computer programs, and the processor is used to execute the computer programs, so that the image external drawing device can implement any of the above-mentioned image external drawing methods.
[0093] A storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any image drawing method provided in the embodiment of the present application.
[0094] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any one of the image drawing methods provided in the embodiments of the present application.
[0095] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0096] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can refer to each other.
[0097] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image drawing method, characterized in that: include: Calling a pre-trained multimodal model to process the original image to obtain an image generation prompt word, wherein the image generation prompt word includes a description text for generating an outer drawing area of the original image; A pre-trained image generation model is called to generate an externally drawn image according to the image generation prompt word and the original image; wherein the image generation model is a stable diffusion model configured with an extended model, the extended model is a pre-trained internal drawing control network ControlNet++, and the internal drawing control network ControlNet++ uses the original image to control the image generation process of the stable diffusion model, so that the area corresponding to the original image in the generated image of the stable diffusion model tends to be consistent with the original image.
2. The image drawing method according to claim 1, characterized in that: The calling of the pre-trained multimodal model to process the original image to obtain the image generation prompt words includes: The original image and pre-configured image processing prompt words are input into the multimodal model so that the multimodal model outputs the image generation prompt words; the image processing prompt words are used to instruct the multimodal model to recognize and understand the input image, infer the external drawing area of the input image based on the recognized image content, and generate and output prompt words based on the recognized image content and the inferred external drawing content.
3. The image drawing method according to claim 1 or 2, characterized in that: The calling of the pre-trained image generation model to generate the externally drawn image according to the image generation prompt word and the original image includes: Filling the surrounding area of the original image with a preset color to generate an extended image that includes the original image and has the same size as the image after external drawing, as a control image; The control diagram is used as a conditional input of the internal drawing control network ControlNet++, and the image generation prompt word is used as a prompt input of the internal drawing control network ControlNet++ and the stable diffusion model, and the configured stable diffusion model is called to generate the external drawing image based on the input original image.
4. The image drawing method according to claim 3, characterized in that: The stable diffusion model adopts a sampling guidance method that combines perturbation attention guidance with classifier-free guidance.
5. The image drawing method according to claim 1 or 2, characterized in that: After generating the externally drawn image, the method further includes: Post-processing the image after external drawing to obtain the external drawing result of the original image; wherein post-processing the image after external drawing includes: The image area in the externally drawn image corresponding to the original image is replaced with the original image, and image fusion processing is performed on the seam area between the externally drawn image and the original image.
6. The image drawing method according to claim 5, characterized in that: Post-processing the externally drawn image also includes: Before replacing the image area corresponding to the original image in the externally drawn image, the following steps are performed: Counting pixel values of image areas corresponding to the original image in the original image and the externally drawn image respectively; Calculate the pixel value distribution of the image area corresponding to the original image in the original image and the externally drawn image in the RGB color space based on the kernel density estimation method; Establishing a pixel value mapping relationship between the original image and an image region in the externally drawn image corresponding to the original image according to the calculated pixel value distribution; The pixel values of the externally drawn image are corrected according to the pixel value mapping relationship, and based on the corrected externally drawn image, the step of replacing the image area in the externally drawn image corresponding to the original image with the original image is continued.
7. The image drawing method according to claim 5, characterized in that: Post-processing the externally drawn image also includes: Before replacing the image area corresponding to the original image in the externally painted image, calling a pre-trained image restoration model to perform restoration processing on the externally painted image, and continuing to perform the step of replacing the image area corresponding to the original image in the externally painted image with the original image based on the restored externally painted image; Wherein, the image restoration model includes a face restoration model and an image super-resolution model.
8. An image drawing device, characterized in that: include: A prompt word generation unit, used to call a pre-trained multimodal model to process the original image to obtain an image generation prompt word, wherein the image generation prompt word includes a description text for generating an outer drawing area of the original image; An image generation unit is used to call a pre-trained image generation model to generate an externally drawn image according to the image generation prompt word and the original image; wherein the image generation model is a stable diffusion model configured with an extended model, the extended model is a pre-trained internal drawing control network ControlNet++, and the internal drawing control network ControlNet++ uses the original image to control the image generation process of the stable diffusion model so that the area corresponding to the original image in the generated image of the stable diffusion model tends to be consistent with the original image.
9. An image drawing device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the image drawing device can implement the image drawing method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the image drawing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image redrawing method and device, computer equipment and storage medium
CN117808917A
Method, device and equipment for generating image more biased to hobbies of people and medium
CN118115630A
Image expansion method and device, readable storage medium and program product
CN119625137A
Image restoration method, and image restoration model training method and device
CN119831901A
Generative ai-based video aspect ratio enhancement
US20250063136A1