Image preview method and device

Through the combination of deep learning image generation model and preview model, the preview model is used to iterate the intermediate image, which solves the problem of long preview image generation time in the prior art, and realizes the preview effect close to the final image can be generated in the early stage of the image generation process, improving the preview efficiency.

CN120151591APending Publication Date: 2025-06-13SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311713731.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing image preview method requires a long wait for a preview image that is closer to the final result, which is less efficient and affects the user experience.

Method used

By combining the deep learning image generation model and the preview model, the preview model is used to iterate the intermediate image to generate a preview image that is closer to the final image.

Benefits of technology

In the early stage of the image generation process, preview images that are closer to the final image can be output, which improves the efficiency of image preview and saves users' waiting time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151591A_ABST
    Figure CN120151591A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of image previewing, and provides an image previewing method, which comprises the following steps: according to an initial noise image input by a deep learning image generation model, under the constraint of at least one cue word or cue statement, outputting an intermediate image generated by the ith iteration, wherein the deep learning image generation model comprises N deep learning network layers, and the deep learning network layers are used for executing a single iteration generation step of an image; the intermediate image is input to a preset preview model, the preview model inputs the intermediate image and iterates M times to generate a preview image, the preview model comprises M deep learning network layers, i is smaller than N, and M is smaller than N-i. Compared with the prior art, the image previewing method provided by the invention can be used for previewing the previewing result close to the final image in the early stage of the image generation process, so that the image previewing efficiency is improved, and the time is saved for a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image preview, and particularly relates to an image preview method and device. Background Art

[0002] In the field of image generation, it usually takes several seconds to dozens of seconds to generate a high-quality image. During the process of waiting for the final image to be generated, it is very important to provide a preview image for the user. If the preview image does not meet the user's requirements, the user can terminate the generation process in time and adjust parameters such as the prompt words to regenerate the image. This approach can save computing power resources compared to the method of regenerating the image when the result is unsatisfactory after completing the entire generation process, and also saves time for the user.

[0003] However, the current image preview methods usually directly decode the intermediate results of the image generation process. This method has a fast decoding speed, but the generated preview image has a poor effect. Generally, only when the image generation process reaches 60% - 80%, can a preview image closer to the final result be generated. Therefore, users need to wait for a long time to preview a preview image closer to the final result, with low efficiency and affecting the user experience. Summary of the Invention

[0004] The embodiments of this application provide an image preview method and device, which can solve the problem that the current preview method requires a long waiting time to generate a preview image closer to the final result.

[0005] In a first aspect, the embodiments of this application provide an image preview method, including:

[0006] According to an initial noise image input by a deep learning image generation model, under the constraint of at least one prompt word or prompt statement, an intermediate image generated in the i-th iteration is output, where the deep learning image generation model includes N deep learning network layers, and the deep learning network layers are used to execute a single iteration generation step of the image;

[0007] The intermediate image is input into a preset preview model, and the preview model inputs the intermediate image and iterates M times to generate a preview image. The preview model includes M deep learning network layers, where i is less than N, M is less than N - i, and i, N, and M are all positive integers.

[0008] In a possible implementation, the image preview method further includes:

[0009] Input at least one constraint image into the deep learning image generation model to make the N deep learning network layers iterate under the constraint of the constraint image.

[0010] In a possible implementation, the constraint image includes at least one of a line constraint map, a depth constraint map, a pose constraint map, an item type constraint map, and a style constraint map corresponding to the at least one prompt or prompt statement.

[0011] In a possible implementation, the image preview method further includes:

[0012] Select a diffusion model;

[0013] Using the historical final image marked with the at least one prompt or prompt statement as the output of the diffusion model and a random initial noise as the input, train to obtain the deep learning image generation model and the preview model.

[0014] In a possible implementation, using the historical final image marked with the at least one prompt or prompt statement as the output of the diffusion model and a random initial noise as the input, training to obtain the deep learning image generation model and the preview model includes:

[0015] Execute multiple training steps, each training step including inputting the historical final image marked with the at least one prompt or prompt statement into the deep learning image generation model and the preview model, and solidifying the outputs of the deep learning image generation model and the preview model as the historical final image to obtain an updated deep learning image generation model and preview model.

[0016] In a second aspect, an embodiment of the present application provides an image generation method, including:

[0017] According to an initial noise image input by a deep learning image generation model, under the constraint of at least one prompt or prompt statement, output an intermediate image generated in the i-th iteration, where the deep learning image generation model includes N deep learning network layers, and the deep learning network layer is used to execute a single iteration generation step of the image;

[0018] Input the intermediate image into a preset preview model, the preview model inputs the intermediate image and iterates M times to generate a preview image, the preview model includes M layers of the deep learning network layer, where i is less than N, M is less than N - i, and i, N, and M are all positive integers;

[0019] Judge whether the preview image meets a preset standard, if it meets, continue to generate the final image.

[0020] In a third aspect, an embodiment of the present application provides an image preview device, including:

[0021] An intermediate image acquisition module, configured to output an intermediate image generated in the i-th iteration under the constraint of at least one prompt or prompt statement, where the deep learning image generation model includes N deep learning network layers, and the deep learning network layers are used to execute a single iteration generation step of an image;

[0022] A preview image generation module, configured to input the intermediate image into a preset preview model, the preview model inputs the intermediate image and iterates M times to generate a preview image, the preview model includes M of the deep learning network layers, where i is less than N, M is less than N - i, and i, N, and M are all positive integers.

[0023] In a possible implementation, i / N is less than 90%, and / or M is equal to 2.

[0024] In a fourth aspect, an embodiment of the present application provides a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the method described above is implemented.

[0025] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.

[0026] Advantages of the present application

[0027] The present application is applicable to the technical field of image preview, and provides an image preview method, including: according to an initial noise image input by a deep learning image generation model, outputting an intermediate image generated in the i-th iteration under the constraint of at least one prompt or prompt statement, where the deep learning image generation model includes N deep learning network layers, and the deep learning network layers are used to execute a single iteration generation step of an image; inputting the intermediate image into a preset preview model, the preview model inputs the intermediate image and iterates M times to generate a preview image, the preview model includes M of the deep learning network layers, where i is less than N, M is less than N - i, and i, N, and M are all positive integers.

[0028] The image preview method provided by the present application uses a preview model to process the intermediate image generated by a deep learning image generation model. By setting the number of iterations of the image generation model and the preview model, the preview model can output a preview image closer to the final image in the early stage of the image generation process, improving the efficiency of image preview and saving preview time for users. Description of the Drawings

[0029] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a schematic flowchart of an image preview method provided by an embodiment of the present application;

[0031] Figure 2 It is a schematic diagram of a deep learning image generation model and a preview model provided by an embodiment of the present application;

[0032] Figure 3 It is one of the comparison diagrams of the preview effect of an embodiment of the present application and the preview effect of the prior art;

[0033] Figure 4 It is a schematic flowchart of an image generation method provided by an embodiment of the present application;

[0034] Figure 5 It is a schematic structural diagram of an image preview device provided by an embodiment of the present application;

[0035] Figure 6 It is the second comparison diagram of the preview effect of an embodiment of the present application and the preview effect of the prior art;

[0036] Figure 7 It is the third comparison diagram of the preview effect of an embodiment of the present application and the preview effect of the prior art;

[0037] Figure 8 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0038] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0039] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0040] It should also be understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0041] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "once" or "in response to determining" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]" depending on the context.

[0042] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0043] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.

[0044] In the field of image generation, it usually takes several seconds to dozens of seconds to generate a high-quality image. During the waiting process for the final image to be generated, it is very important to provide a preview image for the user. If the preview image does not meet the user's requirements, the user can terminate the generation process in time and adjust parameters such as the prompt words to regenerate the image. This approach can save computing power resources compared to the way of regenerating when the result is unsatisfactory after completing the entire generation process, and also saves time for the user.

[0045] However, the current image preview methods usually directly decode the intermediate results of the image generation process. This method has a very fast decoding speed, but the generated preview image has a poor effect. Generally, only when the image generation process reaches 60% - 80%, can a preview image closer to the final result be generated. Therefore, the user needs to wait for a relatively long time to preview a preview image closer to the final result, with low efficiency and affecting the user experience.

[0046] This application is applicable to the field of image preview technology and provides an image preview method, including: according to an initial noise image input by a deep learning image generation model, under the constraint of at least one prompt word or prompt statement, output an intermediate image generated in the i-th iteration, where the deep learning image generation model includes N deep learning network layers, and the deep learning network layers are used to execute a single iteration generation step of the image; input the intermediate image into a preset preview model, the preview model inputs the intermediate image and generates a preview image after M iterations, and the preview model includes M deep learning network layers, where i is less than N, M is less than N - i, and i, N, and M are all positive integers.

[0047] The image preview method provided by this application uses the preview model to process the intermediate image generated by the deep learning image generation model. By setting the number of iterations of the image generation model and the preview model, it is possible to preview a preview image that is relatively close to the final image in the early stage of the image generation process, saving time for users to preview.

[0048] To illustrate the technical solution of this application, it will be described below through specific embodiments.

[0049] Refer to Figure 1 The flow of an embodiment of the shown image preview method, by way of example and not limitation, includes the following steps:

[0050] Step S100: According to an initial noise image input by a deep learning image generation model, under the constraint of at least one prompt word or prompt statement, output an intermediate image generated in the i-th iteration, where the deep learning image generation model includes N deep learning network layers, and the deep learning network layers are used to execute a single iteration generation step of the image;

[0051] Step S200: Input the intermediate image into a preset preview model, the preview model inputs the intermediate image and generates a preview image after M iterations, and the preview model includes M deep learning network layers, where i is less than N, M is less than N - i, and i, N, and M are all positive integers.

[0052] To better understand this solution, a brief introduction to artificial intelligence-generated content is provided first. Artificial Intelligence Generated Content (AIGC) is a new type of content creation method. It is a method based on artificial intelligence technologies such as Generative Adversarial Networks (GAN) and large pre-trained models. By learning from existing data and performing pattern recognition, it is a technology that can generate relevant content with appropriate generalization ability. The core idea of AIGC technology is to use artificial intelligence algorithms to generate content with a certain degree of creativity and quality. Through training models and learning from a large amount of data, AIGC can generate content related to the input conditions or instructions. For example, by inputting keywords, descriptions, or samples, AIGC can generate matching articles, images, audio, etc. In the embodiments of this application, it mainly involves the text-image generation process.

[0053] It should be noted that in step S100, the deep learning image generation model includes N deep learning network layers. Therefore, generating the final image requires N iterations. Here, the deep learning network layer is used to perform a single iteration generation step of the image, that is, through the calculation and optimization of the network layer, the initial noise image is gradually transformed into an intermediate image closer to the target image.

[0054] For the intermediate image generated in each iteration, it can be directly decoded to generate the corresponding preview image. However, for the intermediate images generated in the first few iterations, due to the small number of iterations, their clarity is very low, and the preview images directly decoded cannot well reflect some features in the final image. Therefore, the preview effect is poor.

[0055] Therefore, in the embodiments of this application, for the intermediate image generated in the i-th iteration obtained in step S100, further processing is required, that is, the intermediate image generated in the i-th iteration is input into a preset preview model. In the preview model, the intermediate image generated in the i-th iteration will be further iterated. After iterating M times, its iteration result is decoded to obtain the corresponding preview image. It should be noted that the number of iterations M is less than N - i. Therefore, the step length of each iteration in the preview model is longer than the step length of one iteration of the deep learning image generation model. Therefore, the speed of obtaining the preview image through the preview model is faster than the speed of obtaining the final image by the deep learning image generation model.

[0056] Exemplarily, when N = 10, i = 1, and M = 2, the number of iterations of the deep learning image generation model is 10. The intermediate image generated in the first iteration of the deep learning image generation model, that is, the intermediate result when the iteration reaches 10%, is input into the preview model. In the preview model, the intermediate image needs to be iterated twice to obtain the preview image. It should be noted that the two iterations in the preview model need to complete 90% of the iteration process that is not completed in the deep learning image generation model. Therefore, in the preview model, the step size of each iteration is longer than that of the deep learning image generation model. Assuming that the time required for both the deep learning image generation model and the preview model to iterate once is t, for the intermediate image generated in the first iteration, the time required to generate the final image by iteration in the deep learning image generation model is 9t, while the time required to generate the preview image in the preview model is only 2t. Therefore, the speed of obtaining the preview image using the preview model is faster than the speed of obtaining the final image using the deep learning image generation model.

[0057] In a possible implementation manner, the image preview method further includes:

[0058] Step A1: Input at least one constraint image into the deep learning image generation model, so that the N-layer deep learning network layer iterates under the constraint of the constraint image.

[0059] It should be noted that the ultimate goal of image generation is to make the finally generated image closer to the picture that the user wants to obtain. Therefore, more constraint conditions can more effectively reduce the diversity and randomness during image generation. Therefore, in this embodiment, the constraint image corresponding to at least one prompt word or prompt statement is used as a constraint condition and input into each layer, so that the N-layer deep learning network layer iterates under the constraint of the constraint image.

[0060] Exemplarily, a constraint image can be added to the deep learning image generation model to achieve more refined control over the final image, including control over elements such as the main body, background, style, and form of the final image. For example, a photo can be selected, and the skeleton of the person in the picture can be collected, so as to generate a person with the same pose in the new picture; the line drawing of the picture in the picture can also be selected, so as to generate a picture with the same line drawing in the new picture; the existing style in the picture can also be selected, so as to generate a picture with the same style in the new picture.

[0061] In a possible implementation manner, the constraint image includes at least one of a line constraint diagram, a depth constraint diagram, a pose constraint diagram, an item type constraint diagram, and a style constraint diagram corresponding to the at least one prompt word or prompt statement.

[0062] It can be understood that image generation based on constraint text means that the generated image is expected to conform to the description of the text, while image generation based on constraint images means that the final generated image is expected to conform to the description of the constraint image.

[0063] Exemplarily, the constraint image can be a constraint on the content of the picture. For example, in the case of an item type constraint diagram, inputting it into a deep learning image generation model enables each deep learning network layer to iterate under the constraint of the item type constraint diagram, and the content of the final picture obtained will be similar to the constraint image.

[0064] Similarly, the constraint image can also be a constraint on the structure of the picture, that is, the similarity in spatial structure, such as a line constraint diagram and a pose constraint diagram, etc. Additionally, the constraint image can also be a constraint on the style of the picture, that is, the style of the final generated image is expected to conform to the style of the constraint image. For example, it is expected that the picture is in a Chinese style, classical style, minimalist style, or cyberpunk style, etc.

[0065] It should be noted that under the joint constraints of the prompt words or prompt statements and the constraint image, the final generated image will better meet the user's needs.

[0066] In one possible implementation, the image preview method further includes:

[0067] Step B1: Select a diffusion model;

[0068] Step B2: Use the historical final images that have been marked with the at least one prompt word or prompt statement as the output of the diffusion model, and use a random initial noise as the input to train the deep learning image generation model and the preview model.

[0069] To better understand this solution, the diffusion model will be briefly introduced first.

[0070] The diffusion model (Diffusion) consists of two processes: the forward diffusion process (Forward Diffusion Process) and the reverse generation process (Reverse Diffusion Process). The forward diffusion process gradually adds Gaussian noise to an image until it becomes random noise, while the reverse generation process is a denoising process. We will start from a random noise and gradually denoise it until an image is generated. In each iteration of the forward diffusion process, noise is sampled from a Gaussian distribution, and this noise is recorded during the forward diffusion process. In the reverse diffusion process, only the noise recorded correspondingly during the forward diffusion process is used as a label to train the diffusion model.

[0071] When using Diffusion, images are generated from text. In fact, Diffusion uses a TextEncoder to generate the embedding corresponding to the text (text embedding), which is then used as the input to Diffusion together with a random noise embedding and a timestep embedding, and finally the final image is generated. The text embedding can be generated using the CLIP model. The full English name of CLIP is Contrastive Language-Image Pre-training, which is a pre-training method or model based on contrasting text-image pairs. The training data for CLIP is text-image pairs: an image and its corresponding text description. Through contrastive learning, the model can learn the matching relationship between text-image pairs.

[0072] It should be noted that after selecting a diffusion model and training it, an initial noise image is required, and under the constraint of at least one prompt word or prompt statement, the intermediate image generated in the i-th iteration is output; then, the intermediate image is input into a preset preview model, and a preview image is generated after iterating M times in the preview model.

[0073] In a possible implementation manner, step B2 includes:

[0074] Execute multiple training steps. Each training step includes inputting the historical final image marked with the at least one prompt word or prompt statement into the deep learning image generation model and the preview model, and solidifying the outputs of the deep learning image generation model and the preview model as the historical final image to obtain an updated deep learning image generation model and preview model.

[0075] It should be noted that the deep learning image generation model and the preview model are trained in the same way. The difference is that the preview model has a smaller layer structure. Therefore, when using it, the number of iterations required to generate a preview image using the preview model is less than the number of iterations required to generate the final image using the deep learning image generation model. Therefore, the time taken to generate a preview image using the preview model is less than the time taken to generate the final image using the deep learning image generation model. Using the preview model can obtain a preview image closer to the final image faster.

[0076] The following describes this solution in combination with specific field embodiments.

[0077] Refer to Figure 2 , in one embodiment, a diffusion model is selected as the deep learning image generation model to generate the final image. It should be noted that the diffusion model and the preview model are trained using the same training method.

[0078] The diffusion model can be understood as a process of image denoising. In each iteration, part of the noise in the image is removed. Taking the intermediate image generated in the first iteration as an example, in the diffusion model, to remove all the noise, it still needs to go through 9 deep learning network layers and complete 9 iterations. However, in the preview model, only 2 deep learning network layers need to be passed through and 2 iterations need to be completed to remove all the noise.

[0079] It should be noted that for the intermediate image generated in the first iteration of the diffusion model, the iteration completion degree is 10%. After inputting it into the preview model, the preview model needs to go through two iterations to complete the 90% of the iteration process that was not completed in the diffusion model. Therefore, in the preview model, the step size of each iteration is longer than that of the diffusion model.

[0080] Assume that the time required for both the diffusion model and the preview model to iterate once is t. Then, for the intermediate image generated in the first iteration, the time required to generate the final image through iteration in the deep learning image generation model is 9t, while the time required to generate the preview image in the preview model is only 2t. Therefore, the speed of obtaining the preview image using the preview model is faster than the speed of obtaining the final image using the deep learning image generation model.

[0081] Exemplarily, set N = 10, i = 1, M = 2, that is, the diffusion model includes 10 deep learning network layers, and the preview model includes 2 deep learning network layers. The constraint conditions include five prompt words, a depth constraint map, and a line constraint map. Among them, the prompt words include: masterpiece, best quality, Santorini, Zeus statue, and colorful sky. The diffusion model iterates continuously under the constraints of the above prompt words and constraint images until the final image is output. Further, the intermediate image generated in the first iteration output by the diffusion model is input into a preset preview model. The preview model includes 2 of the above-mentioned deep learning network layers. After the preview model inputs the intermediate image, it iterates 2 times to generate a preview image, as shown in the left picture in Figure 3 In the prior art, the commonly used method is to directly decode the intermediate image generated in the first iteration output by the diffusion model, and the obtained preview image is as shown in the right picture in Figure 3 It can be seen that the preview effect is very unclear, approaching the state of a noise map, and cannot provide an effective preview function. However, in this embodiment, since the intermediate image has been processed by the preview model, a preview image closer to the final image can be obtained. Therefore, compared with the prior art, the image preview method provided by this application can preview a preview result closer to the final image in the early stage of the image generation process.

[0082] In a second aspect, referring to Figure 4 , the embodiments of this application provide an image generation method, including:

[0083] Step P1: According to the initial noise image input by a deep learning image generation model, under the constraint of at least one prompt word or prompt statement, output the intermediate image generated in the i-th iteration, where the deep learning image generation model includes N deep learning network layers, and the deep learning network layers are used to execute the single-iteration generation step of the image;

[0084] Step P2: Input the intermediate image into a preset preview model. The preview model inputs the intermediate image and iterates M times to generate a preview image. The preview model includes M of the deep learning network layers, where i is less than N, M is less than N - i, and i, N, and M are all positive integers;

[0085] Step P3: Determine whether the preview image meets the preset standard. If it meets, continue to generate the final image.

[0086] In the embodiment of the present application, the deep learning image generation model outputs the intermediate image generated in the i-th iteration. Further, the intermediate image generated in the i-th iteration is input into a preset preview model. In the preview model, the intermediate image generated in the i-th iteration will be further iterated. After iterating M times, its iteration result is decoded to obtain the corresponding preview image. It should be noted that the number of iterations M is less than N - i. Therefore, the step length of each iteration in the preview model is longer than the step length of one iteration of the deep learning image generation model. Therefore, the speed of obtaining the preview image through the preview model is faster than the speed of obtaining the final image by the deep learning image generation model. After obtaining the preview image, determine whether it meets the preset standard. If it meets, continue to generate the final image. If it does not meet, the image generation process can be terminated in time, and the constraint conditions can be adjusted to generate the final image that meets the requirements. Therefore, by using the above image generation method, users can better control the image generation process and make adjustments to the image generation process more timely, improving the efficiency of image generation.

[0087] In a third aspect, the embodiment of the present application provides an image preview device. Refer to Figure 5 , including:

[0088] An intermediate image acquisition module 301, configured to output the intermediate image generated in the i-th iteration under the constraint of at least one prompt word or prompt statement, where the deep learning image generation model includes N deep learning network layers, and the deep learning network layers are used to execute the single-iteration generation step of the image;

[0089] A preview image generation module 302, configured to input the intermediate image into a preset preview model. The preview model inputs the intermediate image and iterates M times to generate a preview image. The preview model includes M of the deep learning network layers, where i is less than N, M is less than N - i, and i, N, and M are all positive integers.

[0090] It should be noted that the intermediate image acquisition module 301 is used to output the intermediate images generated during the iterative process of the deep learning image generation model. The deep learning image generation model includes N deep learning network layers. Therefore, generating the final image requires N iterations, and for each intermediate image generated in each iteration, it can be output. The traditional image preview method directly decodes the intermediate image to generate the corresponding preview image. However, for the intermediate images generated in the previous several iterations, due to the small number of iterations, their clarity is very low, and the preview images directly decoded cannot well reflect some features in the final image. Therefore, the preview effect is poor.

[0091] Therefore, in the embodiment of the present application, the intermediate image output by the intermediate image acquisition module 301 needs to be further processed, that is, the intermediate image generated in the i-th iteration is input into a preset preview model. In the preview model, the intermediate image generated in the i-th iteration will be further iterated. After iterating M times, the iteration result is decoded to obtain the corresponding preview image. It should be noted that the number of iterations M is less than N - i. Therefore, the step length of each iteration in the preview model is longer than the step length of one iteration of the deep learning image generation model. Therefore, the speed of obtaining the preview image through the preview model is faster than the speed of obtaining the final image by the deep learning image generation model. Therefore, using the image preview device provided by the present application, a preview image closer to the final image can be obtained more quickly.

[0092] In a possible implementation manner, i / N is less than 90%, and / or M is equal to 2.

[0093] Refer to Figure 6 , when i / N = 90%, since the picture generation process in the deep learning image generation model is almost over, the intermediate image at this time is already relatively close to the final image. Therefore, directly decoding the intermediate image can also obtain a relatively clear preview image. And the image preview method and device provided by the present application have a good preview effect on the intermediate images in the early stage of the iterative process. Refer to Figure 7 , when i / N is less than 90%, for example, i / N is equal to 60%, it can be clearly seen at this time that the preview effect of using the image preview method and device provided by the present application is better than the preview image generated by direct decoding.

[0094] It should be noted that if the number of deep learning network layers of the preview model is set to be relatively large, the time required to iteratively generate the preview image will be longer. Therefore, in this embodiment, the number of deep learning network layers of the preview model is set to two layers. The intermediate image only needs to be iterated twice in the preview model to obtain the preview image, and the generation speed is relatively fast, reducing the waiting time of the user.

[0095] Figure 8A schematic structural diagram of a terminal device provided by an embodiment of the present application. The terminal device 400 includes: at least one processor 401( Figure 8 Only one processor is shown), a memory 402, and a computer program 403 stored in the memory 402 and executable on the at least one processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned method embodiment 8 are implemented.

[0096] The terminal device 400 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art can understand that Figure 8 This is only an example of the terminal device 400 and does not constitute a limitation on the terminal device 400. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0097] The so-called processor 401 may be a central processing unit (CPU). The processor 401 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0098] In some embodiments, the memory 402 may be an internal storage unit of the terminal device 400, such as the hard disk or memory of the terminal device 400. In other embodiments, the memory 402 may also be an external storage device of the terminal device 400, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 400. Further, the memory 402 may also include both the internal storage unit and the external storage device of the terminal device 400. The memory 402 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory 402 may also be used to temporarily store data that has been output or will be output.

[0099] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example for illustration. In actual applications, the above-mentioned functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0100] The embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented.

[0101] The embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the foregoing method embodiments when executed.

[0102] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps in the foregoing method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0103] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0104] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0105] In the embodiments provided in this application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0106] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0107] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. An image preview method, characterized in that, comprising: According to an initial noise image input by a deep learning image generation model, under the constraint of at least one prompt word or prompt statement, output an intermediate image generated in the i-th iteration, wherein the deep learning image generation model includes N layers of deep learning network layers, and the deep learning network layers are used to execute a single iteration generation step of the image; Input the intermediate image into a preset preview model, the preview model inputs the intermediate image, and iterates M times to generate a preview image, the preview model includes M layers of the deep learning network layers, wherein i is less than N, M is less than N - i, and i, N, and M are all positive integers.

2. The image preview method according to claim 1, characterized in that, further comprising: Input at least one constraint image into the deep learning image generation model, so that the N layers of deep learning network layers iterate under the constraint of the constraint image.

3. The image preview method according to claim 2, characterized in that, The constraint image includes at least one of a line constraint map, a depth constraint map, a pose constraint map, an item type constraint map, and a style constraint map corresponding to the at least one prompt word or prompt statement.

4. The image preview method according to claim 1, characterized in that, further comprising: Select a diffusion model; Using the historical final image marked with the at least one prompt word or prompt statement as the output of the diffusion model, and using a random initial noise as the input, train to obtain the deep learning image generation model and the preview model.

5. The image preview method according to claim 4, characterized in that, Using the historical final image marked with the at least one prompt word or prompt statement as the output of the diffusion model, and using a random initial noise as the input, training to obtain the deep learning image generation model and the preview model includes: Execute multiple training steps, each training step includes inputting the historical final image marked with the at least one prompt word or prompt statement into the deep learning image generation model and the preview model, and solidifying the output of the deep learning image generation model and the preview model as the historical final image, to obtain an updated deep learning image generation model and preview model.

6. An image generation method, characterized in that, Executed by the entire system, including: According to an initial noise image input by a deep learning image generation model, under the constraint of at least one prompt word or prompt statement, output an intermediate image generated in the i-th iteration, wherein the deep learning image generation model includes N layers of deep learning network layers, and the deep learning network layers are used to execute a single iteration generation step of the image; Input the intermediate image into a preset preview model, the preview model inputs the intermediate image, and iterates M times to generate a preview image, the preview model includes M layers of the deep learning network layers, wherein i is less than N, M is less than N - i, and i, N, and M are all positive integers; Judge whether the preview image meets a preset standard, if so, continue to generate the final image.

7. An image preview device, It is characterized in that including an intermediate image acquisition module, configured to output an intermediate image generated in the i-th iteration under the constraint of at least one prompt word or prompt statement, wherein the deep learning image generation model includes N deep learning network layers, and the deep learning network layers are configured to execute a single iteration generation step of an image a preview image generation module, configured to input the intermediate image into a preset preview model, the preview model inputs the intermediate image and generates a preview image after M iterations, the preview model includes M deep learning network layers, where i is less than N, M is less than N - i, and i, N, and M are all positive integers 8. The image preview device according to claim 7 It is characterized in that i / N is less than 90%, and / or M is equal to 2 9. A terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor It is characterized in that when the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented 10. A computer-readable storage medium storing a computer program It is characterized in that when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented