Image generation method and device, equipment and storage medium
During the image generation process, the image control network is used to extract global features and image information, and the image to be processed is optimized and redrawn multiple times, which solves the problem of image edge inconsistency and improves the effect of image generation.
Patent Information
- Application Number
- CN202311757386.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
In the process of image generation, the generated image edges are incoherent, resulting in poor drawing effects.
By acquiring the image and prompt words to be processed, the global features are extracted using the pre-constructed image control network, the pending area is determined, and redrawn according to the prompt words and global features to generate the first image. Subsequently, the image information is extracted to optimize and redraw the first image, and a second image is generated.
The coherence of image edges is achieved, the picture difference is reduced, and the image generation effect is improved.
Smart Images

Figure CN120182401A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image generation method, apparatus, device, and storage medium. Background Art
[0002] With the rapid development of digital media, digital image processing technologies have become increasingly mature. Currently, in order to draw personalized content on an original image, artificial intelligence technologies are mostly used to assist in the generation work, and the image-to-image generation technology has received more and more attention and research.
[0003] Currently, after selecting an adjustment area of the original image, the image-to-image generation technology is directly used to redraw or add content to the adjustment area. However, there is a color difference between the generated adjustment image and the original image on the screen, and the edges of the original image and the redrawn content in the generated adjustment image are not coherent, resulting in a relatively poor drawing effect of the generated adjustment image. Summary of the Invention
[0004] In order to solve the above technical problems, the present disclosure provides a method, apparatus, electronic device, and storage medium, effectively solving the problem of relatively poor image drawing effect caused by non-coherent edges of the generated image.
[0005] In a first aspect, an embodiment of the present disclosure provides an image generation method, including:
[0006] Obtaining a to-be-processed image and a first prompt word, where the first prompt word is used to determine the redrawn content;
[0007] Extracting global features of the to-be-processed image through a pre-constructed image control network, and determining a to-be-processed area of the to-be-processed image, where the to-be-processed area includes a redrawn area and a non-redrawn area, the redrawn area refers to an extended area of the to-be-processed image and / or an internal editing area of the to-be-processed image, and the non-redrawn area refers to at least a part of the to-be-processed image that does not need to be redrawn;
[0008] Redrawing the to-be-processed area according to the first prompt word and the global features to generate a first image;
[0009] Extracting image information of the to-be-processed image, and redrawing the first image according to the image information to generate a second image, where the image information includes edge information and / or segmentation information.
[0010] Optionally, the obtaining the to-be-processed image and the first prompt word includes:
[0011] Obtain the image to be processed, and extract the description text of the picture in the image to be processed; use the description text and / or the input text as the first prompt, where the input text refers to the new content to be redrawn on the image to be processed; or,
[0012] Obtain the prompt text, and use the text-to-image method to generate the image to be processed based on the prompt text; use one or more of the prompt text, the input text, and the blank text as the first prompt.
[0013] Optionally, redrawing the area to be processed according to the first prompt and the global feature to generate the first image includes:
[0014] Set the redrawing amplitude of the redrawing area to the first amplitude value, and at the same time set the redrawing amplitude of the remaining areas in the area to be processed except the redrawing area to the second amplitude value, where the first amplitude value is greater than the second amplitude value, and the amplitude values of the areas in the remaining areas gradually increase as they approach the redrawing area, and the amplitude values of the areas in the redrawing area gradually decrease as they approach the remaining areas;
[0015] Based on the first prompt and the global feature, redraw the redrawing area at the first amplitude value, and at the same time redraw the remaining areas at the second amplitude value to generate the first image.
[0016] Optionally, extracting the image information of the image to be processed and redrawing the first image according to the image information to generate the second image includes:
[0017] Extract the image information of the image to be processed through a pre-constructed image control network;
[0018] Perform image enhancement processing on the first image to obtain an enhanced image;
[0019] Redraw the enhanced image according to the image information and perform sampling a preset number of times during the redrawing process to generate the second image.
[0020] Optionally, the preset number includes a first value and a second value, where the first value is less than the second value,
[0021] Redrawing the enhanced image according to the image information and performing sampling a preset number of times during the redrawing process to generate the second image includes:
[0022] Obtain a second prompt, where the second prompt refers to blank text or the prompt text of the image to be processed;
[0023] Redraw the enhanced image based on the image information and the second prompt word until the current sampling count obtained in real time during the redrawing process reaches the first value, generating redrawing data;
[0024] Continue to redraw based on the second prompt word on the basis of the redrawing data until the current sampling count obtained in real time during the redrawing process reaches the second value, generating a second image.
[0025] Optionally, the redrawing the first image according to the image information to generate a second image includes:
[0026] Calculate the proportion of the image to be processed in the first image;
[0027] If the proportion is less than a preset threshold, redraw the first image according to the image information and the prompt text of the image to be processed, generating a second image.
[0028] Optionally, the redrawing the first image according to the image information to generate a second image includes:
[0029] Redraw the first image according to the image information to generate a third image of a first size;
[0030] Redraw the third image according to the extracted image information and redrawing parameters of the third image to generate a second image of a second size;
[0031] Wherein, the first size is smaller than the second size and larger than the size of the image to be processed.
[0032] In a second aspect, an embodiment of the present disclosure provides an image generation device, including:
[0033] An acquisition unit, configured to acquire an image to be processed and a first prompt word, wherein the first prompt word is used to determine the redrawing content;
[0034] A determination unit, configured to extract the global features of the image to be processed through a pre-constructed image control network and determine the area to be processed of the image to be processed, wherein the area to be processed includes a redrawing area and a non-redrawing area, the redrawing area refers to the expanded area of the image to be processed and / or the internal editing area of the image to be processed, and the non-redrawing area refers to at least part of the area of the image to be processed that does not need to be redrawn;
[0035] A first generation unit, configured to perform a first redrawing on the image to be processed according to the first prompt word, generating a first image;
[0036] A second generation unit, configured to extract image information of the image to be processed, and perform a second redrawing on the first image according to the image information to generate a second image, where the image information includes edge information and / or segmentation information.
[0037] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:
[0038] A memory;
[0039] A processor; and
[0040] A computer program;
[0041] wherein the computer program is stored in the memory and is configured to be executed by the processor to implement the image generation method as described above.
[0042] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the image generation method as described above are implemented.
[0043] An embodiment of the present disclosure provides an image generation method, including: obtaining an image to be processed and a first prompt word, where the first prompt word is used to determine the redrawing content; extracting global features of the image to be processed through a pre-constructed image control network, and determining a region to be processed of the image to be processed, where the region to be processed includes a redrawing region and a non-redrawing region, the redrawing region refers to an extended region of the image to be processed and / or an internal editing region of the image to be processed, and the non-redrawing region refers to at least a part of the region of the image to be processed that does not need to be redrawn; performing a first redrawing on the region to be processed according to the first prompt word and the global features to generate a first image; extracting image information of the image to be processed, and performing a second redrawing on the first image according to the image information to generate a second image, where the image information includes edge information and / or segmentation information. In the method provided in this application, after the first redrawing of the image to be processed by the prompt word to generate the first image, the second redrawing is performed on the first image according to the image information of the image to be processed. The second redrawing is a process of optimizing the first image generated by the first redrawing. While realizing redrawing-related functions such as image expansion and image element elimination through two redrawings, it also optimizes the redrawing region with discontinuous edge connections in the picture, effectively reducing the picture difference and obtaining a second image with a better generation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0045] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0046] Figure 1 Schematic flowchart of an image generation method provided by an embodiment of the present disclosure;
[0047] Figure 2 Schematic structural diagram of an image generation method provided by an embodiment of the present disclosure;
[0048] Figure 3 Schematic structural diagram of another image generation method provided by an embodiment of the present disclosure;
[0049] Figure 4 Schematic detailed flowchart of S103 in an image generation method provided by an embodiment of the present disclosure;
[0050] Figure 5 Schematic detailed flowchart of S104 in an image generation method provided by an embodiment of the present disclosure;
[0051] Figure 6 Schematic structural diagram of an image generation device provided by an embodiment of the present disclosure;
[0052] Figure 7 Schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners
[0053] In order to be able to more clearly understand the above-mentioned objects, features, and advantages of the present disclosure, the following will further describe the solutions of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0054] In the following description, many specific details are set forth in order to fully understand the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.
[0055] In view of the above technical problems, the embodiments of the present disclosure provide an image generation method, which is applied to the scene of image-to-image, can realize redrawing requirements such as image expansion and internal image editing of the image to be processed, and to a certain extent ensures the edge coherence of the image to be processed and the redrawing area, reduces the picture difference of the generated image, and effectively improves the generation effect. Specifically, it can be described in detail through at least one of the following embodiments.
[0056] Figure 1 The flowchart of an image generation method provided by an embodiment of the present disclosure is applied to a server or a terminal. In a possible application scenario, the server performs a first redrawing on the image to be processed and sends the generated first image to the terminal, and the terminal performs a second redrawing on the received first image to generate a second image with better redrawing effect. In another possible application scenario, the terminal or the server independently performs the first redrawing of the image to be processed and the second redrawing of the first image. Other possible application scenarios are not limited herein. The following embodiments will take the server executing the image generation method as an example for detailed description, specifically including the following steps S101 to S104 as shown in Figure 1 follows:
[0057] S101. Obtain the image to be processed and the first prompt.
[0058] Wherein, the first prompt is used to determine the redrawing content.
[0059] It can be understood that when obtaining the image to be processed and the first prompt, the image to be processed can be understood as the original image, and the first prompt refers to the content to be redrawn on the original image. There are specifically two requirements for generating a redrawn image based on the original image, namely image expansion and internal image editing. Among them, image expansion is an operation of generating an image for the surrounding expansion area according to the content of the picture itself in the original image on the basis of the original image, that is, expanding the picture edge; internal image editing is an operation of selecting the redrawing area in the picture of the original image according to the user's drawing idea and redrawing the redrawing area, which can realize functions such as watermark removal and eraser.
[0060] Optionally, for S101 to obtain the image to be processed and the first prompt above, it can be specifically implemented through the following steps:
[0061] Obtain the image to be processed and extract the description text of the picture in the image to be processed; use the description text and / or the input text as the first prompt, wherein the input text refers to the new content to be redrawn on the image to be processed; or, obtain the prompt text, and generate the image to be processed based on the prompt text by the text-to-image method; use one or more of the prompt text, the input text, and the blank text as the first prompt.
[0062] Understandably, if the image to be processed is directly obtained, the text description of the picture in the image to be processed is extracted through a neural network model to obtain the description text. Among them, the neural network model can be a CLIP (Contrastive Language-Image Pre-Training) model or a multimodal model (Bootstrapping Language-Image Pre-training, BLIP). In this case, the first prompt is the description text and / or the input text, where the input text specifically refers to the new content to be redrawn on the image to be processed. For example, content such as cats and dogs is to be redrawn on the original image. Or, if the obtained is a prompt text (Prompt, or the original prompt), the text-to-image generation model is used to generate the image to be processed according to the prompt text. In this case, the first prompt can be at least one of the prompt text, the input text, and the blank text. The blank text means that the first prompt can be empty, and the surrounding expansion area can be filled according to the content of the picture itself.
[0063] S102. Extract the global features of the image to be processed through a pre-constructed image control network, and determine the area to be processed of the image to be processed.
[0064] Among them, the area to be processed includes a redrawing area and a non-redrawing area. The redrawing area refers to the expansion area of the image to be processed and / or the internal editing area of the image to be processed. The non-redrawing area refers to at least part of the area of the image to be processed that does not need to be redrawn.
[0065] It is understandable that, based on the above S101, the global features of the image to be processed are extracted through a pre-constructed Image Control Network (ControlNet). The global features can also be understood as global information. ControlNet is a simple transfer learning method, and the extracted global features can be used in the subsequent image-to-image generation process, that is, the information of the original image, such as depth maps, segmentation maps, key points and other data, is used to control the newly generated image. Among them, the redrawing area refers to the expanded area of the image to be processed and / or the internal editing area of the image to be processed, and the non-redrawing area refers to at least part of the area of the image to be processed that does not need to be redrawn. Determining the redrawing area of the area to be processed, for the image expansion requirement, the redrawing area can be the expanded area after the expansion of the image to be processed. For the internal editing area of the image, the redrawing area can be a certain editing area in the image to be processed. For both requirements, the redrawing area includes the expanded area and a certain editing area. The editing area refers to the area where new content is generated, and at the same time, the other areas of the image to be processed except a certain editing area can be understood as the non-redrawing area. After determining the redrawing area, the area to be processed is constructed according to the redrawing area and the non-redrawing area of the image to be processed. The non-redrawing area refers to at least part of the area of the image to be processed that does not need to be redrawn. The size of the non-redrawing area can be the same as the size of the image to be processed, that is, the entire image to be processed is the non-redrawing area, or it can be smaller than the size of the image to be processed, that is, part of the area of the image to be processed is the non-redrawing area. The area to be processed can be understood as a mask area. For example, the edge of the non-redrawing area is expanded, and in this case, the content of the area of the image to be processed except the non-redrawing area may be changed.
[0066] S103. Redraw the area to be processed according to the first prompt and the global features to generate a first image.
[0067] It is understandable that, based on the above S102, optionally, after obtaining the global features, an image-to-image generation method is adopted to redraw the area to be processed according to the first prompt and the global features to generate a first image, which specifically includes the following content.
[0068] It is understandable that, based on the above S401, a large model of Wenshengtu (Stable Diffusion 1.5, SD1.5 or Stable Diffusion XL, SDXL) is used to control the redrawing content according to the first prompt word and global features / global information, and the first redrawing is performed to generate the first image. If the first prompt word is blank text, the image expansion can be achieved based on the global information. If the first prompt word is input text, the internal editing of the image can be achieved based on the first prompt word, and the image expansion can be achieved based on the global information. That is to say, in a redrawing process, multiple redrawing functions can be achieved, including internal image editing, image expansion, and single redrawing. Other achievable redrawing functions and function types are not limited here. It is understandable that the first redrawing of the area to be processed according to the first prompt word and global features generates a first image with a relatively coherent picture and insignificant difference in picture content. Specifically, the Refiner Inpaint module in the SDXL model can be used for the first redrawing. If the first redrawing requirement is image expansion, the first prompt word can be blank text or input text without the original prompt word, the original prompt word refers to the description word related to the original screen content of the image to be processed, and the input text refers to the description word related to the newly added content of the image to be processed.
[0069] S104: extracting image information of the image to be processed, and redrawing the first image according to the image information to generate a second image.
[0070] The image information includes edge information and / or segmentation information.
[0071] It is understandable that, based on the above S103, the image information of the image to be processed is extracted through the image control network (ControlNet), and the image information may be edge information and / or segmentation information. Subsequently, the image to be processed is redrawn for the second time according to the image information using the image-to-image method to generate a second image. The second redrawing can be understood as a process of optimizing the first image generated by the first redrawing. It is understandable that the picture in the generated first image may still have a certain color difference and a clear sense of boundary at the edge. Therefore, the first image can be redrawn for the second time based on the edge information and / or segmentation information, that is, the first image is optimized to eliminate the picture difference and the area of incoherent edges.
[0072] Optionally, the redrawing of the first image according to the image information to generate the second image in S104 may be specifically implemented through the following steps:
[0073] Calculate the proportion of the image to be processed in the first image; if the proportion is less than a preset threshold, redraw the first image according to the image information and the prompt text of the image to be processed to generate a second image.
[0074] It is understandable that calculating the ratio of the image to be processed in the first image, if this ratio is less than the preset threshold, that is, if the main subject of the picture occupies a small part of the picture frame, then add the original prompt word (prompt text) during the second redrawing process to avoid more picture differences caused by the second redrawing.
[0075] Optionally, the redrawing of the first image according to the image information in S104 to generate a second image can be specifically implemented through the following steps:
[0076] Redraw the first image according to the image information to generate a third image of the first size; redraw the third image according to the extracted image information and redrawing parameters of the third image to generate a second image of the second size; wherein, the first size is smaller than the second size and larger than the size of the image to be processed.
[0077] It is understandable that during the process of redrawing the image to be processed to generate a second image of the second size, a third image of the first size can be generated first to further improve the generation effect. The third image can be understood as an intermediate image. The first size is smaller than the second size and larger than the size of the image to be processed. The number of third images is not limited. Specifically: redraw the first image according to the image information to generate a third image; redraw the third image according to the image information and redrawing parameters to generate a second image. Among them, the redrawing parameters can include information such as a third prompt word and the generated image size. The third prompt word can be at least one of input text, blank text, prompt text, and description text. The image information can be the information of the image to be processed or the information of the third image. In one possible case, the first redrawing realizes the image expansion function, the first prompt word is blank text, the second redrawing realizes the internal image editing function, the third prompt word is input text, such as a cat, and the third redrawing is the redrawing optimization process, and the second prompt word is blank text. In another possible case, the first redrawing realizes the image expansion function, and both the second redrawing and the third redrawing are optimization processes. Other possible generation processes are not limited here and can be determined according to user needs. That is, the drawing process of adjusting the image screen such as image expansion and / or image editing is realized based on global information, and the optimization process of the redrawn image is realized based on image information.
[0078] Exemplarily, refer to Figure 2 , Figure 2Schematic structural diagram of an image generation method provided by an embodiment of the present disclosure. Based on the image expansion requirements implemented by SD1.5, specifically, an original image of 512*512 is obtained, and the original image is preliminarily expanded. For example, the original image is expanded vertically to 512*1024. Subsequently, the expanded image of 512*1024 is border-expanded to determine the redrawing area. As Figure 2 shown, the redrawing area is the area where 512*1024 is border-expanded to 1024*1024. The target area where the redrawing area and the original image are located is combined to obtain a 1024*1024 area to be processed. The 1024*1024 area to be processed is used as the input for the first redrawing to generate a 1024*1024 first image. Subsequently, the 1024*1024 first image is redrawn (optimized) for the second time to generate a 1024*1024 second image with coherent pictures.
[0079] Exemplarily, refer to Figure 3 , Figure 3 Schematic structural diagram of another image generation method provided by an embodiment of the present disclosure. Based on the image expansion requirements implemented by SDXL, after obtaining the original prompt Prompt (prompt text), use the text-to-image function of SDXL to generate an original image with a size of 1024*1024 according to Prompt, and use Prompt as the first prompt word for the first redrawing. When the main body occupies a relatively small part of the picture, Prompt can also be added as the second prompt word for the second redrawing to avoid differences caused by redrawing. Or, after obtaining an original image with a size of 1024*1024, obtain the description text of the picture in the original image through CLIP or BLIP, and use this description text as the first prompt word for the first redrawing. Subsequently, determine the redrawing area of the image to be processed, and splice the redrawing area and the target area where the original image is located to obtain the original area. As Figure 3 shown, two redrawing areas of 512*1024 and a target area of 1024*1024 are spliced to obtain a 2024*1024 area to be processed. Subsequently, the 2024*1024 area to be processed is redrawn for the first time to generate a first image, and then the first image is optimized to generate a second image.
[0080] For the image generation method provided by the embodiment of the present disclosure, after the image to be processed is redrawn for the first time through a prompt word to generate a first image, the first image is redrawn for the second time according to the image information of the image to be processed. The second redrawing is a process of continuously optimizing based on edge information and / or segmentation information on the basis of the first image. While realizing redrawing-related functions such as image expansion and image element elimination, the redrawing area with incoherent edge connection in the picture is optimized. Through the process of multiple redrawings, the redrawing accuracy is effectively improved, the picture difference is reduced, and a second image with better generation effect is obtained.
[0081] Based on the above embodiments, Figure 4 FIG. S103 is a detailed flowchart provided by an embodiment of the present disclosure. Optionally, the step of redrawing the area to be processed according to the first prompt word and the global feature to generate a first image specifically includes the following steps S401 to S402 as shown in Figure 4 follows:
[0082] S401. Set the redrawing amplitude of the redrawing area to a first amplitude value, and at the same time set the redrawing amplitude of the remaining areas in the area to be processed except the redrawing area to a second amplitude value.
[0083] Wherein, the first amplitude value is greater than the second amplitude value, and the amplitude values of the areas in the remaining areas gradually increase as they approach the redrawing area, and the amplitude values of the areas in the redrawing area gradually decrease as they approach the remaining areas.
[0084] S402. Based on the first prompt word and the global feature, redraw the redrawing area at the first amplitude value, and at the same time redraw the remaining areas at the second amplitude value to generate a first image.
[0085] It can be understood that after obtaining the area to be processed, set the redrawing amplitude of the redrawing area to the first amplitude value, and set the redrawing amplitude of the remaining areas in the area to be processed except the redrawing area to the second amplitude value. When the redrawing area is an extended area of the image to be processed, the remaining area is the target area. When the redrawing area is a certain editing area within the image to be processed, the target area is the complete area of the image to be processed, that is, the redrawing area is within the target area, so that the combined size of the area to be processed is the same as the size of the target area / image to be processed. In this case, the remaining area can be understood as the remaining area in the target area except the redrawing area. Among them, the first amplitude value is greater than the second amplitude value. Preferably, control the redrawing amplitude of the remaining area of the Mask to 0 to keep the image to be processed unchanged. The remaining area is the area where the image to be processed is located. The closer to the remaining area, the smaller the redrawing amplitude, that is, the content of the remaining area image hardly changes. The closer to the redrawing area, the greater the redrawing amplitude.
[0086] The image generation method provided by the embodiment of the present disclosure changes the redrawing position corresponding to the image to be processed and the redrawing amplitude corresponding to different positions through Mask, and can use different gray values of Mask to blur at the edge for the fusion of the image edge to perform precise redrawing and the coherence of the image, ensuring the redrawing effect.
[0087] Based on the above embodiments, Figure 5It is a schematic diagram of the refined process of S104 provided by the embodiments of the present disclosure. Optionally, image information of the to-be-processed image is extracted, and the first image is redrawn according to the image information to generate a second image, which specifically includes the following steps S501 to S503 as shown in Figure 5 follows:
[0088] S501. Extract the image information of the to-be-processed image through a pre-constructed image control network.
[0089] Among them, the image information includes edge information and / or segmentation information.
[0090] It is understandable that ControlNet is used to obtain the image information of the to-be-processed image, and the image information includes edge information (Canny) or segmentation information (Segmentation).
[0091] S502. Perform image enhancement processing on the first image to obtain an enhanced image.
[0092] It is understandable that before optimizing the first image, to reduce the sense of style of the picture, the first image can be preprocessed. For example, the first image is processed through image sharpening and other enhancement methods to enhance the edge sense of the picture, so that more details can be retained after subsequent image generation from image. The specific enhancement method is not limited.
[0093] S503. Redraw the enhanced image according to the image information, and perform sampling for a preset number of times during the redrawing process to generate a second image.
[0094] It is understandable that on the basis of the above S501 and S502, using the image generation from image function of SDXL or SD1.5, setting the second prompt word to be empty, setting the redrawing amplitude of the redrawing area to 0.2, optimizing the enhanced image according to the image information, and setting the number of sampler sampling times to be the preset number of times during the redrawing process to generate a second image with coherent pictures.
[0095] Among them, the preset number of times includes a first value and a second value, where the first value is less than the second value.
[0096] Optionally, in the above S503, redrawing the enhanced image according to the image information and the second prompt word, and performing sampling for a preset number of times during the redrawing process to generate a second image can be specifically implemented through the following steps:
[0097] Obtain a second prompt word, where the second prompt word refers to blank text or a prompt text of the image to be processed; redraw the enhanced image based on the image information and the second prompt word until the currently obtained sampling times during the redrawing process reach the first value, generating redrawing data; continue to redraw based on the second prompt word on the basis of the redrawing data until the currently obtained sampling times during the redrawing process reach the second value, generating a second image.
[0098] It can be understood that taking 20 sampling steps as an example of the preset number of times, starting from step 0, use the image information to control the generation of the picture and generate redrawing data. Stop using the image information for control when the sampling step reaches the first value (50%), that is, step 10. Starting from step 11, do not use the image information for control, and continue to automatically redraw on the basis of the redrawing data until the sampling step is step 20 to obtain a second image with a coherent picture.
[0099] Preferably, the parameter settings for redrawing based on the image-to-image function of SDXL are as follows: Set the Mask edge blur (MaskBlur) to 15 to obtain a better edge blur fusion effect; when generating the first image for the first redrawing, select the description text or the prompt text of the image to be processed as the first prompt word in the masked area (MaskedContent). The prompt text can be at least part of the picture content in the image to be processed, that is, the description text and the prompt text can be the same or different. Use the DPM++2MKarras sampler for sampling. 2M refers to the second-order sampling scheme, and Karras is a noise schedule. The main performance is that the noise step size will be smaller near the end, which helps to improve the quality of the redrawn image. In addition, the images generated during the redrawing process can only reach the latent space and do not need to generate images visible to the human eye. This method helps to improve the redrawing speed. Subsequently, perform global optimization on the first image that has been redrawn in the latent space. During the second redrawing process, the sampler remains unchanged, use 20 sampling steps, set the CFGScale parameter to 7, and set the redrawing amplitude to be between 0.3 and 0.4. Use the SDXL image-to-image function to decode the generated image into the original space to obtain a second image with a unified picture.
[0100] The image processing method provided by the embodiments of the present disclosure reduces the sense of picture segmentation and retains more image details by enhancing the first image. Subsequently, when optimizing the enhanced image, through the method of staged sampling, control the generation of the picture through image information in the first sampling stage and adopt the automatic generation method in the second sampling stage to obtain a second image with a coherent picture.
[0101] Figure 6The structural schematic diagram of the image generation device provided by an embodiment of the present disclosure. The image generation device provided by an embodiment of the present disclosure may execute the processing flow provided by the above-mentioned image generation method embodiment, such as Figure 6 As shown, the image generation device 600 includes an acquisition unit 601, a determination unit 602, a first generation unit 603, and a second generation unit 604, where:
[0102] The acquisition unit 601 is configured to acquire a to-be-processed image and a first prompt word, where the first prompt word is used to determine the content to be redrawn;
[0103] The determination unit 602 is configured to extract the global feature of the to-be-processed image through a pre-constructed image control network, and determine the to-be-processed area of the to-be-processed image, where the to-be-processed area includes a redrawing area and a non-redrawing area, the redrawing area refers to the extended area of the to-be-processed image and / or the internal editing area of the to-be-processed image, and the non-redrawing area refers to at least part of the area of the to-be-processed image that does not need to be redrawn;
[0104] The first generation unit 603 is configured to redraw the to-be-processed area according to the first prompt word and the global feature to generate a first image;
[0105] The second generation unit 604 is configured to extract the image information of the to-be-processed image, and redraw the first image according to the image information to generate a second image, where the image information includes edge information and / or segmentation information.
[0106] Optionally, the acquisition unit 601 is configured to:
[0107] Acquire a to-be-processed image, and extract the description text of the picture in the to-be-processed image; use the description text and / or the input text as the first prompt word, where the input text refers to the new content to be redrawn on the to-be-processed image; or,
[0108] Acquire a prompt text, and generate the to-be-processed image based on the prompt text by using the text-to-image method; use one or more of the prompt text, the input text, and the blank text as the first prompt word.
[0109] Optionally, the first generation unit 603 is configured to:
[0110] Set the redrawing amplitude of the redrawing area to a first amplitude value, and at the same time set the redrawing amplitude of the remaining areas in the area to be processed except the redrawing area to a second amplitude value, where the first amplitude value is greater than the second amplitude value, and the amplitude values of the areas in the remaining areas gradually increase as they approach the redrawing area, and the amplitude values of the areas in the redrawing area gradually decrease as they approach the remaining areas;
[0111] Based on the first prompt word and the global feature, redraw the redrawing area at the first amplitude value, and at the same time redraw the remaining areas at the second amplitude value to generate a first image.
[0112] Optionally, the second generation unit 604 is configured to:
[0113] Extract the image information of the image to be processed through a pre-constructed image control network;
[0114] Perform image enhancement processing on the first image to obtain an enhanced image;
[0115] Redraw the enhanced image according to the image information and the second prompt word, and perform sampling a preset number of times during the redrawing process to generate a second image.
[0116] Optionally, the preset number of times in the device 600 includes a first value and a second value, where the first value is less than the second value.
[0117] Optionally, the second generation unit 604 is configured to:
[0118] Obtain a second prompt word, where the second prompt word refers to blank text or the prompt text of the image to be processed;
[0119] Redraw the enhanced image based on the image information and the second prompt word until the currently obtained sampling times during the redrawing process reach the first value to generate redrawing data;
[0120] Continue to redraw based on the second prompt word on the basis of the redrawing data until the currently obtained sampling times during the redrawing process reach the second value to generate a second image.
[0121] Optionally, the second generation unit 604 is configured to:
[0122] Calculate the proportion of the image to be processed occupying the first image;
[0123] If the proportion is less than a preset threshold, redraw the first image according to the image information, the obtained second prompt word, and the prompt text of the image to be processed to generate a second image.
[0124] Optionally, the second generation unit 604 is configured to:
[0125] Redraw the first image according to the image information to generate a third image with a first size;
[0126] Redraw the third image according to the extracted image information and redrawing parameters of the third image to generate a second image with a second size;
[0127] Wherein, the first size is smaller than the second size and larger than the size of the image to be processed.
[0128] Figure 6 The image generation device in the illustrated embodiment can be used to execute the technical solutions of the above method embodiments. The implementation principles and technical effects are similar and will not be elaborated here.
[0129] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Specifically refer to Figure 7 , which shows a schematic structural diagram of an electronic device 700 suitable for implementing the present disclosure. The electronic device 700 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 7 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0130] As Figure 7 shown, the electronic device 700 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 701, which can execute various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage device 708 into the random access memory (RAM) 703 to implement the image generation method of the embodiments as described in the present disclosure. In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored. The processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.
[0131] Typically, the following devices can be connected to the I / O interface 705: input devices 706 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 707 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and a communication device 709. The communication device 709 can allow the electronic device 700 to communicate with other devices wirelessly or wireline to exchange data. Although Figure 7 the electronic device 700 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.
[0132] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart, so as to implement the image generation method as described above. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 709, or installed from the storage device 708, or installed from the ROM 702. When the computer program is executed by the processing device 701, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0133] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0134] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0135] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.
[0136] Optionally, when the above one or more programs are executed by the electronic device, the electronic device can also perform the other steps described in the above embodiments.
[0137] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0139] The units involved in the embodiments described in this disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases.
[0140] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0141] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0142] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or gateway that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or gateway. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or gateway that includes the elements.
[0143] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image generation method, characterized in that, Including: Obtain a to-be-processed image and a first prompt word, where the first prompt word is used to determine the content to be redrawn; Extract the global features of the to-be-processed image through a pre-constructed image control network, and determine the to-be-processed area of the to-be-processed image, where the to-be-processed area includes a redrawing area and a non-redrawing area, the redrawing area refers to the expanded area of the to-be-processed image and / or the internal editing area of the to-be-processed image, and the non-redrawing area refers to at least part of the area of the to-be-processed image that does not need to be redrawn; Redraw the to-be-processed area according to the first prompt word and the global features to generate a first image; Extract the image information of the to-be-processed image, and redraw the first image according to the image information to generate a second image, where the image information includes edge information and / or segmentation information.
2. The method according to claim 1, characterized in that, The obtaining of the to-be-processed image and the first prompt word includes: Obtain a to-be-processed image, and extract the description text of the picture in the to-be-processed image; use the description text and / or the input text as the first prompt word, where the input text refers to the new content to be redrawn on the to-be-processed image; or, Obtain a prompt text, and generate the to-be-processed image based on the prompt text by using the text-to-image method; use one or more of the prompt text, the input text, and the blank text as the first prompt word.
3. The method according to claim 1, characterized in that, The redrawing of the to-be-processed area according to the first prompt word and the global features to generate a first image includes: Set the redrawing amplitude of the redrawing area to a first amplitude value, and at the same time set the redrawing amplitude of the remaining areas in the to-be-processed area except the redrawing area to a second amplitude value, where the first amplitude value is greater than the second amplitude value, and the amplitude values of the areas in the remaining areas gradually increase as they get closer to the redrawing area, and the amplitude values of the areas in the redrawing area gradually decrease as they get closer to the remaining areas; Based on the first prompt word and the global features, redraw the redrawing area at the first amplitude value, and at the same time redraw the remaining areas at the second amplitude value to generate a first image.
4. The method according to claim 1, characterized in that, The extracting of the image information of the to-be-processed image and the redrawing of the first image according to the image information to generate a second image includes: Extract the image information of the to-be-processed image through a pre-constructed image control network; Perform image enhancement processing on the first image to obtain an enhanced image; Redraw the enhanced image according to the image information, and perform sampling a preset number of times during the redrawing process to generate a second image.
5. The method according to claim 4, characterized in that, The preset number includes a first value and a second value, where the first value is less than the second value, The redrawing of the enhanced image according to the image information and the sampling a preset number of times during the redrawing process to generate a second image includes: Obtain a second prompt word, where the second prompt word refers to the blank text or the prompt text of the to-be-processed image; Redraw the enhanced image based on the image information and the second prompt word until the currently sampled number obtained in real time during the redrawing process reaches the first value, generating redrawing data; Based on the second prompt word, continue to redraw on the basis of the redrawing data until the currently sampled number obtained in real time during the redrawing process reaches the second value, generating a second image.
6. The method according to claim 1, characterized in that, The redrawing the first image according to the image information to generate a second image includes: Calculate the proportion of the image to be processed occupying the first image; If the proportion is less than a preset threshold, redraw the first image according to the image information and the prompt text of the image to be processed, generating a second image.
7. The method according to claim 1, characterized in that, The redrawing the first image according to the image information to generate a second image includes: Redraw the first image according to the image information to generate a third image of a first size; Redraw the third image according to the extracted image information of the third image and the obtained redrawing parameters to generate a second image of a second size; Wherein, the first size is smaller than the second size and larger than the size of the image to be processed.
8. An image generation device, characterized in that, The apparatus includes: An acquisition unit, configured to acquire an image to be processed and a first prompt word, wherein the first prompt word is used to determine the redrawing content; A determination unit, configured to extract the global features of the image to be processed through a pre-constructed image control network and determine the area to be processed of the image to be processed, wherein the area to be processed includes a redrawing area and a non-redrawing area, the redrawing area refers to the expanded area of the image to be processed and / or the internal editing area of the image to be processed, and the non-redrawing area refers to at least part of the area of the image to be processed that does not need to be redrawn; A first generation unit, configured to redraw the area to be processed according to the first prompt word and the global features, generating a first image; A second generation unit, configured to extract the image information of the image to be processed and redraw the first image according to the image information, generating a second image, wherein the image information includes edge information and / or segmentation information.
9. An electronic device, characterized in that, Includes: A memory; A processor; And A computer program; Wherein, the computer program is stored in the memory and is configured to be executed by the processor to implement the image generation method according to any one of claims 1 to 7.
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the steps of the image generation method according to any one of claims 1 to 7.
Citation Information
Cited By
Generative large model-oriented dynamic adaptive interaction system and method
CN121541952A