Image generation method and device

By introducing detection and correction mechanisms into the image generation model, the problem of multiple iterations and corrections required after the image is generated is solved, and efficient resource utilization and image generation efficiency are achieved.

CN120182435APending Publication Date: 2025-06-20LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510238040.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In scenarios where images are generated using text, the generated images need to be reviewed to avoid sensitive content, but this process can lead to waste of processing resources, as the generated images may require multiple iterations and corrections.

Method used

By obtaining the reference image and the description information input by the user, multiple iterations are performed using the image generation model to generate an intermediate result image, and determining whether the intermediate result image meets the setting requirements based on the preset detection strategy. If not, make corrections to generate a target image that meets the requirements.

Benefits of technology

It effectively reduces the situation where the generated target image does not meet the requirements, thereby reducing resource waste and improving the efficiency of the image generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182435A_ABST
    Figure CN120182435A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and device, and the method comprises the steps: obtaining a reference image and description information inputted by a user, and enabling the description information to be used for describing the image content of a to-be-generated target image; on the basis of the description information, performing multiple iteration processing on the reference image by using an image generation model to obtain an intermediate result image corresponding to each processing; determining whether the intermediate result image meets a set requirement or not according to a preset detection strategy; and if the intermediate result image meets the set requirement, continuing to execute iterative processing of the image generation model based on the intermediate result image until a target image is generated, and outputting and displaying the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image generation technologies, and in particular, to an image generation method and apparatus. Background Art

[0002] In the scenario of generating images using text, in order to avoid the generated images containing sensitive content or other content that does not meet the requirements, it is necessary to review the generated images. If the generated image contains content that does not meet the requirements, the image will be deleted and the user will be informed that the image cannot be generated. However, generating images requires a large amount of processing resources. Therefore, discarding the image after generating it will inevitably result in waste of the processing resources consumed for generating the image. Summary of the Invention

[0003] On the one hand, this application provides an image generation method, including:

[0004] Obtaining a reference image and description information input by a user, where the description information is used to describe the image content of the target image to be generated;

[0005] Based on the description information, using an image generation model to perform multiple iterative processes on the reference image to obtain intermediate result images corresponding to each process;

[0006] According to a preset detection strategy, determining whether the intermediate result image meets the set requirements;

[0007] If the intermediate result image meets the set requirements, based on the intermediate result image, continue to execute the iterative process of the image generation model until the target image is generated and output for display.

[0008] In a possible implementation, it further includes:

[0009] If the intermediate result image does not meet the set requirements, using the image generation model to correct the intermediate result image to generate a target image that meets the set requirements.

[0010] In another possible implementation, the step of determining whether the intermediate result image meets the set requirements according to a preset detection strategy includes:

[0011] Obtaining the current iteration number of the image generation model processing the reference image and the intermediate result image obtained by the current iterative process. If the current iteration number and / or the intermediate result image meet the set conditions, starting from the current iteration number, according to the set sampling frequency, obtaining at least one intermediate result image obtained by the iterative process and determining whether the obtained intermediate result image meets the set requirements;

[0012] Or,

[0013] Based on at least one sampling number set by the user, when the current iteration number of the image generation model processing the reference image is the sampling number, obtain the intermediate result image obtained by the current processing, and determine whether the obtained intermediate result image meets the set requirements.

[0014] In another possible implementation, the determining whether the intermediate result image meets the set requirements includes:

[0015] After processing the intermediate result image, send it to the detection device, and receive the detection result from the detection device, where the detection result is used to indicate whether the intermediate result image meets the set requirements;

[0016] Or, use the local detection module to detect whether the intermediate result image meets the set requirements.

[0017] In another possible implementation, the determining whether the intermediate result image meets the set requirements includes at least one of the following:

[0018] Determine whether the semantic information expressed by the intermediate result image meets the set requirements;

[0019] Determine whether the image features of the intermediate result image meet the set requirements;

[0020] If the intermediate result image does not meet the set requirements, use the image generation model to correct the intermediate result image, including at least one of the following:

[0021] If it is detected that the semantic information expressed by the intermediate result image does not meet the set requirements, discard the intermediate result image, and return to execute the process of using the image generation model to perform multiple iterations on the reference image to regenerate the image;

[0022] If it is detected that the image features of the intermediate result image do not meet the set requirements, based on the image features that do not meet the set requirements in the intermediate result image, use the image generation model to correct the intermediate result image.

[0023] In another possible implementation, the if it is detected that the image features of the intermediate result image do not meet the set requirements, based on the image features that do not meet the set requirements in the intermediate result image, use the image generation model to correct the intermediate result image, includes:

[0024] If it is detected that the image features of the intermediate result image do not meet the set requirements, obtain at least one abnormal image region where the image features that do not meet the set requirements are located in the intermediate result image and the abnormal cause information indicating that the image features of the abnormal image region do not meet the set requirements.

[0025] Based on at least one abnormal image region in the intermediate result image and the corresponding abnormal cause information of the abnormal image region, use the image generation model to correct the intermediate result image.

[0026] In another possible implementation manner, the step of using the image generation model to correct the intermediate result image based on at least one abnormal image region in the intermediate result image and the corresponding abnormal cause information of the abnormal image region includes:

[0027] Based on at least one abnormal image region in the intermediate result image and the corresponding abnormal cause information of the abnormal image region, use the image generation model to perform iterative processing on each abnormal image region in the intermediate result image until an intermediate result image that meets the set requirements is obtained.

[0028] In another possible implementation manner, the step of using the image generation model to correct the intermediate result image if the intermediate result image does not meet the set requirements includes:

[0029] If the intermediate result image does not meet the set requirements, increment the abnormal processing count by one.

[0030] If the abnormal processing count at the current moment has not exceeded the set count, use the image generation model to correct the intermediate result image.

[0031] If the abnormal processing count at the current moment exceeds the set count, control the image generation model to end the iterative processing of the reference image.

[0032] In another possible implementation manner, the method further includes:

[0033] After displaying the target image, obtain the region to be adjusted marked by the user in the target image and the adjustment requirement information of the region to be adjusted.

[0034] Based on the region to be adjusted marked in the target image and the adjustment requirement information of the region to be adjusted, use the image generation model to perform iterative processing on the target image to generate an adjusted target image.

[0035] Output the adjusted target image.

[0036] In another aspect, the present application further provides an image generation device, including:

[0037] An information acquisition unit for acquiring a reference image and description information input by a user, where the description information is used to describe the image content of a target image to be generated;

[0038] An image processing unit for performing iterative processing on the reference image multiple times using an image generation model based on the description information to obtain intermediate result images corresponding to each processing;

[0039] An image detection unit for determining whether an intermediate result image meets a set requirement according to a preset detection strategy;

[0040] An iteration control unit for, if the intermediate result image meets the set requirement, continuing to perform the iterative processing of the image generation model based on the intermediate result image until the target image is generated and output for display. Description of the Drawings

[0041] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the original elements and elements are not necessarily drawn to scale.

[0042] Figure 1 It is a schematic flowchart of an image generation method provided by the present application;

[0043] Figure 2 It is another schematic flowchart of an image generation method provided by the present application;

[0044] Figure 3 It is another schematic flowchart of an image generation method provided by the present application;

[0045] Figure 4 It is a schematic diagram of an implementation logic framework of an image generation method provided by the present application;

[0046] Figure 5 It is a schematic diagram of a composition structure of an image generation device provided by the present application;

[0047] Figure 6 It is a schematic diagram of a composition architecture of an electronic device provided by the present application. Detailed Description of the Embodiments

[0048] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than intended to limit the present application. As known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0049] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0050] As Figure 1 , a schematic flowchart of an image generation method provided by the present application is shown. The method of this embodiment can be applied to an electronic device, and the electronic device can be a mobile phone, a tablet computer or a notebook computer. Of course, it can also be a device node in a cloud platform or a distributed system, or a server, etc., without limitation.

[0051] The method of this embodiment may include:

[0052] S101, obtaining a reference image and description information input by a user.

[0053] Among them, the reference image is an image that needs to be processed by using image generation to obtain the required image. For example, the reference image can be a noise image; it can also be an image with some content different from the image expected to be generated by the user, etc., without limitation.

[0054] The description information is used to describe the image content of the target image to be generated. For example, if the user hopes to generate an image of a sleeping kitten, the description information can be "generate an image containing a sleeping kitten".

[0055] Among them, the description information can be in text form or voice form, without specific limitation.

[0056] S102, based on the description information, using an image generation model to perform iterative processing on the reference image to obtain intermediate result images corresponding to each processing.

[0057] In this application, the image generation model continuously optimizes the reference image by performing multiple iterative processes on the reference image, such as performing multiple iterative denoising processes on the reference image, etc., to finally generate a target image that matches the description information.

[0058] Among them, there can be multiple possible specific model types for the image generation model, and no restrictions are imposed on this.

[0059] For example, the image generation model can be a diffusion model, and the diffusion model is used to perform multiple iterative denoising processes on the reference image to generate the target image. Of course, the image generation model can also be a denoising diffusion implicit model, a score-based generation model, a progressive generative adversarial network model, etc., which will not be elaborated here.

[0060] In this application, during the process of the image generation model performing iterative processing on the reference image, the image obtained from each iterative process is called an intermediate result image. Therefore, the intermediate result image can be each image generated by the image generation model before generating the final target image.

[0061] For example, assuming that the image generation model needs to perform 20 iterative processes on the reference image to generate the target image, then the image obtained by the image generation model performing one iterative process on the reference image is an intermediate result image, and the 48 images obtained by the image generation model performing the 2nd to 49th iterative processes on the reference image are also intermediate result images.

[0062] S103, according to the preset detection strategy, determine whether the intermediate result image meets the set requirements.

[0063] Among them, the detection strategy is a strategy for determining whether the intermediate result image needs to be detected, or a strategy for selecting the intermediate result images that need to be detected.

[0064] For example, the detection strategy can include at least one iteration number for which image detection needs to be performed, and the intermediate result image obtained by processing the reference image through the corresponding iteration number is determined as the intermediate result image that needs to be detected.

[0065] Another example is that the detection strategy can be that the intermediate result image obtained from each iterative process needs to be subjected to image detection.

[0066] Of course, there can be other possibilities for the detection strategy, which can be specifically set according to actual needs, and no specific restrictions are imposed.

[0067] Among them, the set requirements can be set according to actual needs, and in different application scenarios, the set requirements will also be different. For example, the set requirements can include the absence of violent content, sensitive information (such as sensitive characters or sensitive objects, etc.) or other set content, etc.

[0068] It should be noted that this step S103 is executed during the process of performing multiple iterations on the reference image using the image generation model, so as to detect whether there are problems with the intermediate result images obtained from the iterative processing that do not meet the set requirements.

[0069] S104, if the intermediate result image meets the set requirements, based on the intermediate result image, continue to perform the iterative processing of the image generation model until the target image is generated and output for display.

[0070] It can be understood that the process of the image generation model performing multiple iterations on the reference image is as follows: using the image generation model to perform the first iterative processing on the reference image to obtain an intermediate result image; then using the image generation model to process the intermediate result image obtained from the first iterative processing to obtain the intermediate result image obtained from the second iterative processing; then, using the image generation model to process the intermediate result image obtained from the second iterative processing, and so on iteratively until the target image is generated.

[0071] It can be seen that after using the image generation model to perform one iterative processing on the reference image to obtain an intermediate result image, in each iterative process after the first iteration, the image generation model needs to process the intermediate result image obtained from the current iterative processing. Thus, it can be known that continuing to perform the iterative processing of the image generation model based on the intermediate result image is actually, taking the intermediate result image as the intermediate result image obtained from the current iteration, using the image generation model to continue to perform iterative processing to achieve multiple iterative processing of the reference image.

[0072] However, different from directly continuing to perform iterative processing using the image generation model based on the intermediate result image every time an intermediate result image is obtained currently, in this application, only when it is confirmed that the intermediate result image meets the set requirements, will the iterative processing be continued using the image generation model based on this intermediate result image, so as to reduce the situation where the target images generated by multiple iterations do not meet the requirements.

[0073] From the above content, it can be known that in this application, during the process of performing multiple iterative processing on the reference image using the image generation model based on the description information, it will be determined whether the intermediate result image meets the set requirements according to the preset detection strategy. Only when this intermediate result image meets the set requirements, will the iterative processing be continued using the image generation model based on this intermediate result image, thereby reducing the situation where the target images generated by the image generation model do not meet the requirements, and naturally reducing the resource consumption caused by the generated target images not meeting the requirements.

[0074] In this application, there can be various possibilities for the preset detection strategy. The following will take several possible situations as examples for illustration.

[0075] In a possible scenario, the present application may determine whether it is necessary to conduct spot checks on the intermediate result images by combining at least one of the number of iterations that the image generation model is currently performing iterative processing and the intermediate result images obtained through the iterations. The following is an illustration in conjunction with Figure 2 for explanation. As Figure 2 , a schematic flowchart of another process of the image generation method provided by the present application is shown. The method of this embodiment may include:

[0076] S201, obtain a reference image and description information input by the user.

[0077] The description information is used to describe the image content of the target image to be generated.

[0078] S202, based on the description information, use the image generation model to perform iterative processing on the reference image multiple times to obtain intermediate result images corresponding to each processing.

[0079] The above two steps can refer to the relevant introductions in the previous embodiments and will not be elaborated here.

[0080] S203, obtain the current iteration number of the image generation model for processing the reference image, and the intermediate result image obtained through the current iterative processing.

[0081] Among them, the current iteration number is the number of times that the image generation model has currently cumulatively performed iterative processing on the reference image. Since the iterative processing of the reference image by the image generation model includes the processing of the reference image and the iterative processing of the intermediate result images generated based on the reference image, the current iteration number is the total number of times that the image generation model has cumulatively iteratively processed the reference image and the intermediate result images of the reference image at the current moment.

[0082] S204, if the current iteration number and / or the intermediate result image meet the set conditions, starting from the current iteration number, obtain at least one intermediate result image obtained through iterative processing according to the set spot check frequency, and determine whether the obtained intermediate result image meets the set requirements.

[0083] Among them, the current iteration number and / or the intermediate result image meeting the set conditions is the condition for triggering spot checks on the intermediate result images iteratively generated by the image generation model.

[0084] The set conditions can be set based on actual needs. The set conditions corresponding to the current iteration number and the intermediate result image will be described separately below:

[0085] Among them, the current iteration count meeting the set condition can be that the current iteration count is within a set iteration count range, or is not lower than a set count. The set count and the iteration count range can be set according to actual needs. For example, it can be based on the degree of compliance between the intermediate result image obtained after the reference image undergoes iteration processing for the set number of times and the description information meeting the requirements, or that the content of the intermediate result image can be initially recognized by the naked eye, etc., to determine the set count or the iteration count range. For example, testers pre-analyze the intermediate result images obtained by the image generation model performing iteration processing on test images for different numbers of times to find the minimum iteration count required for the intermediate result image to contain the basic content outline expressed by the description information, and determine this minimum iteration count as the set count.

[0086] Among them, there can be multiple situations where the intermediate result image meets the set condition. For example, the intermediate result image meeting the set condition can include: the degree of compliance between the intermediate result image and the description information meets the requirements. For example, by identifying the semantic information expressed by the intermediate result image, if the matching degree of the identified semantic information and the semantic expressed by the description information exceeds the set threshold, it is determined that the degree of compliance between the intermediate result image and the description information meets the requirements; or the feature similarity between the image features of the intermediate result image and the features of the description information exceeds the set similarity, etc.

[0087] Another example is that the intermediate result image meeting the set condition can include: the proportion of the detailed area in the intermediate result image exceeds the set ratio. It can be understood that if the intermediate result image has the basic content outline described by the description information, then the intermediate result image will have some edge lines of objects. Therefore, by performing edge detection on the intermediate result image, the detailed area and the flat area of the intermediate result image can be detected. If the detailed area is too small, it means that there is more noise in the intermediate result image and the preliminary outline corresponding to the description information has not yet been formed.

[0088] Of course, there can be other possibilities for the intermediate result image to meet the set condition, which can be specifically set according to actual needs and are not restricted here.

[0089] In practical applications, in the case where both the intermediate result image and the current iteration count meet the set condition, the random inspection of the intermediate result image will be performed. For example, if the current iteration count is not lower than the set count and the degree of compliance between the intermediate result image and the description information meets the requirements, then starting from this current iteration count, at the set random inspection frequency, at least one intermediate result image obtained by iteration processing is acquired.

[0090] Among them, the set random inspection frequency is the frequency of extracting the intermediate result images to be detected from the intermediate result images obtained in each iteration starting from the current iteration count.

[0091] When the types of image generation models are different or the total number of iterations required for the image generation model to iteratively process the reference image is different, the sampling frequency will also be different.

[0092] In a possible case, the sampling frequency can be a fixed frequency. For example, the sampling period corresponding to the sampling frequency can be 5 iterations or 10 iterations, etc. Correspondingly, starting from the current iteration number, the intermediate result image obtained after every 5 or 10 iterations will be acquired and detected. Among them, the value of the fixed sampling frequency can be comprehensively determined according to the image generation model and the total number of iterations for the image generation model to iteratively process the reference image, etc., without specific limitations.

[0093] Taking the example of sampling the intermediate result image every 5 iterations, assume that it takes 30 iterations to process the reference image using the image generation model to obtain the target image, and the current iteration number is the 5th iteration. Then, the intermediate result image obtained in the current 5th iteration is extracted and it is detected whether the intermediate result image meets the set requirements; then, after the reference image is cumulatively iteratively processed 10 times using the image generation model, the intermediate result image obtained after 10 iterations needs to be acquired and it is determined whether the intermediate result image meets the set requirements, and so on. The intermediate result images obtained after 15, 20, and 25 iterations need to be acquired and it is determined whether the corresponding intermediate result images meet the set requirements.

[0094] In another possible case, the sampling frequency can also be a non - fixed frequency, that is, during the process of using the image generation model to iteratively process the reference image multiple times, the sampling frequency will change. In this possible case, the sampling frequency can also be comprehensively determined based on at least one of the type of the image generation model and the total number of iterations required for the image generation model to iteratively process the reference image.

[0095] Specifically, considering that during the process of using the image generation model to iteratively process the reference image multiple times, the intermediate result images obtained in the early - stage iterative processing determine the basic structure of the generated target image, such as the contour and content layout, etc. Based on this, if the intermediate result images obtained in the early - stage iterative processing do not meet the requirements, the probability that the intermediate result images obtained after multiple subsequent iterative processes do not meet the requirements will also be relatively high. Based on this, in order to more reasonably sample the intermediate result images and reduce the resource consumption of the sampled images, in this application, as the number of iterations of the image generation model for iteratively processing the reference image increases, the sampling frequency will gradually decrease.

[0096] For example, assume that sampling inspection starts from the 6th iteration process. Then, in the 6th to 12th iteration processes, sampling inspection is performed on the intermediate result image every two iterations. Correspondingly, it is necessary to sequentially obtain the intermediate result images obtained in the 6th, 8th, 10th, and 12th iteration processes and determine whether they meet the set requirements. From the 12th iteration to the 20th iteration, sampling inspection is performed every four iterations. Then, it is necessary to obtain the intermediate result images obtained in the 16th and 20th iteration processes and detect whether they meet the set requirements. For the 20th to 30th iterations (assuming that the 30th iteration process will obtain the target image), sampling inspection is performed every five iterations. Then, the intermediate result image obtained in the current iteration will be obtained at the 25th iteration and detected whether it meets the set requirements.

[0097] In an alternative manner, during the process of the image generation model iteratively processing the reference image, the value of the sampling inspection frequency can be exponentially decaying.

[0098] It can be understood that when the sampling inspection frequency is a non-fixed frequency, the number of iterations that need to be sampled can be the number of iterations determined in advance according to the sampling inspection frequency, or the number of iterations determined in real time based on the specific sampling inspection frequency, and there is no restriction on this.

[0099] S205. If the obtained intermediate result image meets the set requirements, based on this intermediate result image, continue to perform the iterative processing of the image generation model until the target image is generated and output for display.

[0100] It can be understood that in step S204, after each obtained intermediate result image is determined to meet the set requirements according to the set sampling inspection frequency, this step S205 will be executed. If it is determined according to the set sampling inspection frequency that the intermediate result image obtained in the current iteration does not need to be detected for compliance, then the iterative processing of the image generation model can be directly continued based on this intermediate result image.

[0101] In another possible situation of the detection strategy, in order to meet the actual needs of different users, at least one sampling inspection number for sampling the intermediate result image can be set by the user. For example, the electronic device can prompt the user with the interval of the sampling inspection numbers that can be selected based on the total number of iterations required for the image generation model to iteratively process the reference image, and the user can select at least one sampling inspection number for detecting the intermediate result image within this iteration number interval.

[0102] On this basis, based on at least one sampling inspection number set by the user, when the current iteration number of the image generation model processing the reference image is this sampling inspection number, obtain the intermediate result image obtained in the current processing and determine whether the obtained intermediate result image meets the set requirements.

[0103] For example, the number of spot checks that the user can set includes: the 6th iteration, the 10th iteration, and the 18th iteration. Then, if the current iteration number is the 6th iteration, the 10th iteration, or the 18th iteration, the intermediate result image obtained by the current iteration process can be obtained, and it can be detected whether the intermediate result image meets the set requirements.

[0104] In any embodiment of the present application, the intermediate result image can be detected locally on the electronic device, or the intermediate result image can be detected in a specific detection device. The following describes these two situations.

[0105] In a possible situation, the present application can process the intermediate result image and send it to the detection device, and receive the detection result from the detection device.

[0106] Among them, the detection device is a device other than the electronic device. For example, when the electronic device is a terminal device, the detection device can be a device node in the cloud or other servers, etc., without limitation.

[0107] The detection device will detect whether the intermediate result image meets the set requirements and return the corresponding detection result to the electronic device. Among them, the detection result is used to indicate whether the intermediate result image meets the set requirements.

[0108] Processing the intermediate result image can be to compress the intermediate result image; it can also be to encrypt the intermediate result image processing to convert it into a hidden file. It can be understood that what the present application sends to the detection device is only the intermediate result image obtained during the iteration process. Compared with sending the target image finally generated by the image generation model to the detection device, the amount of information in the intermediate result image is relatively small, thereby reducing the risk of data content leakage in the target image.

[0109] In another possible implementation manner, the present application can also use the local detection module to detect whether the intermediate result image meets the set requirements. Among them, the detection module can be a detection model used to implement detecting whether the intermediate result image meets the set requirements. For example, the detection model can be the same model as the image generation model, or it can be a different model, without limitation. The detection module can also be other processing modules, etc., without limitation.

[0110] In any of the above embodiments of the present application, there can be multiple possible specific implementations for determining whether the intermediate result image meets the set requirements. The following introduces several possible implementation manners.

[0111] In a possible implementation, it can be determined whether the semantic information expressed by the intermediate result image meets the set requirements. For example, a detection module is used to determine the semantic information of the intermediate result image, and the semantic information is detected to meet the set requirements, or a detection device is used to determine whether the semantic information expressed by the intermediate result image meets the set requirements.

[0112] Among them, the semantic information expressed by the intermediate result image can reflect the content information of the intermediate result image. Therefore, combined with the semantic information expressed by the intermediate result image, it can be judged whether there is sensitive information, violent content or set illegal objects in the intermediate result image.

[0113] In another possible implementation, the present application may determine whether the image features of the intermediate result image meet the set requirements. For example, the image features of the intermediate result image are determined, and if the image features of the intermediate result image do not include the set non-compliant image features, it is determined that the image features of the intermediate result image meet the set requirements.

[0114] Among them, the image features of the intermediate result image can represent the information, content or object in the intermediate result image. Therefore, combined with the image features of the intermediate result image, it can be determined whether there are objects such as tools, text or patterns that do not meet the requirements in the intermediate combined image.

[0115] For example, taking the example that the object of tool cannot appear in the intermediate result image, then it is determined that the intermediate result image meets the set requirements. By identifying the image features of the intermediate result image, it can be identified whether there are image features matching the object features of the tool in the intermediate result image. If not, it is confirmed that the intermediate result image meets the set requirements.

[0116] It can be understood that, in actual applications, depending on different actual needs, the present application may adopt any one of the above two implementation methods to determine whether the intermediate result image meets the set requirements; or it may combine the two implementation methods at the same time to determine whether the intermediate result image meets the set requirements, for example, only when the semantic information expressed by the intermediate result image meets the set requirements and the image features of the intermediate result image meet the set requirements, can it be determined that the intermediate result image meets the set requirements.

[0117] It can be understood that in any of the above embodiments of the present application, in order to reduce the situation where the generated target image does not meet the requirements, if the intermediate result image does not meet the set requirements, the present application can also use the image generation model to correct the intermediate result image to generate a target image that meets the set requirements.

[0118] Among them, the purpose of correcting the intermediate result image is to obtain an intermediate result image that meets the requirements, and there can be various possible specific implementation methods for correcting the intermediate result image.

[0119] For example, the intermediate result image can be discarded, and based on the description information, the reference image can be processed iteratively multiple times using the image generation model to regenerate the image.

[0120] However, if the iterative processing of the reference image is terminated every time it is detected that the intermediate result image does not meet the set requirements, and based on the description information, the reference image is iteratively processed again using the image generation model, this will inevitably increase the time consumed to generate the target image required by the user and also result in waste of the resources consumed to generate the intermediate result image.

[0121] Based on this, the present application can also obtain at least one abnormal image region in the intermediate result image that does not meet the set requirements and the abnormal cause information indicating that the abnormal image region does not meet the set requirements. On this basis, based on at least one abnormal image region in the intermediate result image and the corresponding abnormal cause information of the abnormal image region, the intermediate result image is corrected using the image generation model.

[0122] It can be understood that in practical applications, due to different specific implementation methods for determining that the intermediate result image does not meet the set requirements and different specific reasons for the intermediate result image not meeting the set requirements, the specific correction method of the intermediate result image in the present application can also be different.

[0123] Next, in combination with different implementation methods for determining that the intermediate result image does not meet the set requirements, possible implementation methods for correcting the intermediate result image are introduced.

[0124] In a possible case, when determining whether the intermediate result image meets the set requirements is to determine whether the semantic information expressed by the intermediate result image meets the set requirements, if it is detected that the semantic information expressed by the intermediate result image does not meet the set requirements, the present application can discard the intermediate result image and re-execute the iterative processing of the reference image using the image generation model based on the description information to regenerate the image.

[0125] It can be understood that the semantic information expressed by the intermediate result image reflects the content information of the entire intermediate result image. Therefore, if the semantic information expressed by the intermediate result image does not meet the set requirements, it means that the overall content of the intermediate result image does not meet the requirements. In this case, the difference in complexity between correcting the intermediate result image and regenerating the image is not significant. Therefore, the intermediate result image can be discarded, and the reference image can be iteratively processed again using the image generation model based on the description information.

[0126] It can be understood that due to the randomness in the image processing mechanisms such as denoising in the image generation model, the intermediate result images output when the same image is input into the image generation model multiple times will also be different. Based on this, after discarding the intermediate result images that do not meet the requirements in this application, the image generation model is reused to perform iterative processing on the reference image, which can achieve the purpose of updating the intermediate result images generated in each iteration.

[0127] In another possible implementation, when determining whether the intermediate result image meets the set requirements by determining whether the image features of the intermediate result image meet the set requirements, if it is detected that the image features of the intermediate result image do not meet the set requirements, the intermediate result image is corrected using the image generation model based on the image features that do not meet the requirements in the intermediate result image.

[0128] Among them, correcting the intermediate result image based on the image features that do not meet the set requirements in the intermediate result image can be based on the image features that do not meet the set requirements in the intermediate result image, using the image generation model to correct the local area where the image features that do not meet the requirements in the intermediate result image are located, and after confirming that the corrected intermediate result image meets the set requirements, continue to perform the iterative processing of the image generation model based on the corrected intermediate result image.

[0129] Correcting the intermediate result image based on the image features that do not meet the requirements in the intermediate result image can also be to continue to perform the iterative processing of the image generation model based on the image features that do not meet the set requirements in the intermediate result image and the intermediate result image, so as to correct the image features that do not meet the requirements in the intermediate result image during the iterative processing of the image generation model and finally generate the target image.

[0130] In practical applications, this application can adopt one or both of the above implementation methods to perform the correction of the intermediate result image.

[0131] For the sake of easy understanding, the following takes the example of simultaneously detecting whether the semantic information and image features of the intermediate result image meet the requirements to introduce the specific implementation of correcting the intermediate result image.

[0132] Such as Figure 3 , which shows another schematic flowchart of the image generation method provided by this application. The method of this embodiment may include:

[0133] S301, obtain a reference image and description information input by the user.

[0134] Among them, the description information is used to describe the image content of the target image to be generated.

[0135] S302. Based on the description information, use the image generation model to perform multiple iterative processes on the reference image to obtain intermediate result images corresponding to each process.

[0136] S303. Determine the current iteration number of the image generation model for processing the reference image.

[0137] S304. Based on at least one sampling inspection number set by the user, when the current iteration number belongs to the at least one sampling inspection number, obtain the intermediate result image obtained from the current process.

[0138] For ease of understanding, steps S303 and S304 are illustrated by a preset detection strategy. However, it can be understood that if steps S303 and S304 are replaced with "obtain the current iteration number of the image generation model for processing the reference image, and the intermediate result image obtained from the current iterative process. If the current iteration number and / or the intermediate result image meet the set conditions, starting from the current iteration number, according to the set sampling inspection frequency, obtain at least one intermediate result image obtained from the iterative process" or replaced with other preset detection strategies, it is also applicable to this embodiment.

[0139] S305. Determine whether the semantic information expressed by the intermediate result image obtained from the current process meets the set requirements, and whether the image features of the intermediate result image meet the set requirements.

[0140] In this embodiment, it is illustrated by taking the example that while determining whether the semantics expressed by the intermediate result image meet the set requirements, it is also necessary to determine whether the image features of the intermediate result image meet the set requirements.

[0141] In practical applications, it can also be to first determine whether the semantic information expressed by the intermediate result image meets the set requirements. If the semantic information expressed by the intermediate result image meets the set requirements, then determine whether the image features of the intermediate result image meet the set requirements; if the semantic information expressed by the intermediate result image does not meet the set requirements, then there is no need to determine whether the image features of the intermediate result image meet the set requirements, and jump to execute step S307.

[0142] S306. If both the semantic information expressed by the intermediate result image and the image features of the intermediate result image meet the set requirements, based on the intermediate result image, continue to execute the iterative process of the image generation model until the target image is generated and output for display.

[0143] S307. If it is detected that the semantic information expressed by the intermediate result image does not meet the set requirements, discard the intermediate result image, and return to step S302 to re-execute the iterative process of the image generation model for the reference image.

[0144] S308. If it is detected that the image features of the intermediate result image do not meet the set requirements, based on the image features that do not meet the set requirements in the intermediate result image, use the image generation model to correct the intermediate result image.

[0145] In this step, for the correction of the intermediate result image, reference can be made to the relevant introduction above.

[0146] In a possible implementation manner, if it is detected that the image features of the intermediate result image do not meet the set requirements, at least one abnormal image region where the image features that do not meet the set requirements in the intermediate result image are located and the abnormal cause information indicating that the image features of the abnormal image region do not meet the set requirements can also be obtained. For example, the abnormal image region and the corresponding abnormal cause information can be determined by a detection module or a detection device. On this basis, the present application can correct the intermediate result image by using the image generation model based on at least one abnormal image region in the intermediate result image and the corresponding abnormal cause information of the abnormal image region.

[0147] For example, based on at least one abnormal image region in the intermediate result image and the corresponding abnormal cause information of the abnormal image region, on the basis of the intermediate result image, continue to perform iterative processing using the image generation model to correct the intermediate result image during the iterative processing and finally generate a target image.

[0148] In an alternative manner, in order to more effectively ensure that the generated target image meets the set requirements, in the present application, it can also be based on at least one abnormal image region in the intermediate result image and the corresponding abnormal cause information of the abnormal image region, and use the image generation model to perform iterative processing on each abnormal image region in the intermediate result image until an intermediate result image that meets the set requirements is obtained. On this basis, step S306 can be returned to continue performing iterative processing of the image generation model based on the intermediate result image.

[0149] It can be understood that if the semantic information expressed by the intermediate result image does not meet the requirements and the image features of the intermediate result image also do not meet the set requirements, only step S307 can be executed without executing step S308.

[0150] It can be understood that during the process of repeatedly iterating and processing a reference image using an image generation model, if the number of intermediate result images that do not meet the requirements is relatively large, it is very likely that the content expressed by the description information input by the user is inappropriate, or the description information input by the user maliciously does not meet the requirements, etc. If the description information input by the user already contains content that does not meet the set requirements, then there will be more intermediate result images that need to be corrected, resulting in a relatively large consumption of data processing resources. Moreover, it is also possible that even after multiple correction processes, it is still impossible to effectively ensure that the target image meets the set requirements.

[0151] Based on this, in order to further reduce resource consumption and reduce the risk that the target image does not meet the requirements, in any of the above embodiments of the present application, if it is determined that the currently sampled intermediate result image does not meet the set requirements, the present application can increment the number of exception handling times by one. Among them, the number of exception handling times is the number of intermediate result images that need to be corrected, that is, the number of correction times for performing the correction of the intermediate result images.

[0152] On this basis, if the number of exception handling times at the current moment has not exceeded the set number of times, the present application can use the image generation model to correct the intermediate result image, and the specific correction method is as introduced in the previous embodiments and will not be elaborated here. If the number of exception handling times at the current moment exceeds the set number of times, it indicates that there is an abnormality in the description information, and then the iteration process of the reference image by the image generation model can be controlled to end.

[0153] For ease of understanding, reference can be made to Figure 4 , Figure 4 which shows a logical framework diagram of an implementation of the image generation method of the present application.

[0154] As Figure 4 can be seen, during the process of iteratively processing a reference image using an image generation model, if based on several preset detection strategies mentioned above, it is determined that the current iteration number belongs to the iteration number for sampling the intermediate result image, then the intermediate result image obtained from the current iterative processing is acquired.

[0155] By detecting whether the intermediate result image meets the set requirements, it is determined whether the intermediate result image is compliant. If the intermediate result image is compliant, based on this intermediate result image, the iterative processing is continued using the image processing model. If the intermediate result image is non-compliant, the number of correction times (i.e., the number of exception handling times) is incremented by one, and it is detected whether the current number of correction times exceeds the set maximum number of correction times. If so, the iterative processing of the reference image using the image processing model is ended.

[0156] If the number of corrections does not exceed the set maximum number of corrections, the intermediate result image will be corrected, and based on the corrected and compliant intermediate result image, the image processing model will continue to be iteratively processed until the number of iterative processes of the image processing model reaches the target number of iterations, and the target image will be output.

[0157] Of course, in practical applications, the present application can also increment the exception handling count by one each time the correction process of an intermediate result image is completed.

[0158] It can be understood that if the exception handling count exceeds the set number at the current moment, the present application can also output a description correction reminder, which is used to prompt the user to adjust the description information.

[0159] In any of the above embodiments of the present application, after the target image generated by the image generation model is displayed, if the user believes that the target image does not meet the user's requirements, the user can also mark the unsatisfactory image area in the target image and give the adjustment requirement information corresponding to the corresponding image area. On this basis, the present application can also obtain the area to be adjusted marked by the user in the target image and the adjustment requirement information of the area to be adjusted.

[0160] Correspondingly, the present application can iteratively process the target image using the image generation model based on the area to be adjusted marked in the target image and the adjustment requirement information of the area to be adjusted, so as to generate an adjusted target image. Then, the adjusted target image is output.

[0161] Corresponding to an image generation method of the present application, the present application also provides an image generation device.

[0162] As Figure 5 , a schematic diagram of a composition structure of the image generation device provided by the present application is shown. The device in this embodiment includes:

[0163] An information acquisition unit 501, configured to acquire a reference image and description information input by a user, where the description information is used to describe the image content of the target image to be generated;

[0164] An image processing unit 502, configured to perform multiple iterative processes on the reference image using an image generation model based on the description information to obtain an intermediate result image corresponding to each process;

[0165] An image detection unit 503, configured to determine whether the intermediate result image meets the set requirements according to a preset detection strategy;

[0166] An iterative control unit 504, configured to, if the intermediate result image meets the set requirements, continue to perform iterative processing of the image generation model based on the intermediate result image until the target image is generated and output for display.

[0167] In a possible implementation, the image generation device further includes:

[0168] An image correction unit, configured to, if the intermediate result image does not meet the set requirements, correct the intermediate result image by using the image generation model to generate a target image that meets the set requirements.

[0169] In yet another possible implementation, the image detection unit includes:

[0170] A first detection subunit, configured to obtain the current iteration number of the image generation model for processing the reference image, and the intermediate result image obtained by the current iterative processing. If the current iteration number and / or the intermediate result image meet the set conditions, starting from the current iteration number, at a set sampling frequency, obtain at least one intermediate result image obtained by iterative processing, and determine whether the obtained intermediate result image meets the set requirements;

[0171] Or,

[0172] A second detection subunit, configured to, based on at least one sampling number set by a user, when the current iteration number of the image generation model for processing the reference image is the sampling number, obtain the intermediate result image obtained by the current processing, and determine whether the obtained intermediate result image meets the set requirements.

[0173] In yet another possible implementation, when the image detection unit, the first detection subunit, or the second detection subunit determines whether the intermediate result image meets the set requirements, specifically:

[0174] After processing the intermediate result image, send it to a detection device, and receive a detection result from the detection device, where the detection result is used to indicate whether the intermediate result image meets the set requirements;

[0175] Or, use a local detection module to detect whether the intermediate result image meets the set requirements.

[0176] In yet another possible implementation, the image detection unit includes at least one of the following:

[0177] A first determination subunit, configured to determine whether the semantic information expressed by the intermediate result image meets the set requirements according to a preset detection strategy;

[0178] A second determination subunit, configured to determine whether the image features of the intermediate result image meet the set requirements according to a preset detection strategy;

[0179] The image correction unit includes at least one of the following:

[0180] A first correction subunit, configured to discard the intermediate result image and return to execute the operation of the image processing unit to regenerate an image if it is detected that the semantic information expressed by the intermediate result image does not meet the set requirements;

[0181] A second correction subunit, configured to correct the intermediate result image by using the image generation model based on the image features that do not meet the set requirements in the intermediate result image if it is detected that the image features of the intermediate result image do not meet the set requirements.

[0182] In another possible implementation manner, the second correction subunit includes:

[0183] An abnormality determination subunit, configured to obtain at least one abnormal image region where the image features that do not meet the set requirements in the intermediate result image are located and the abnormal cause information indicating that the image features of the abnormal image region do not meet the set requirements if it is detected that the image features of the intermediate result image do not meet the set requirements;

[0184] A correction processing subunit, configured to correct the intermediate result image by using the image generation model based on at least one abnormal image region in the intermediate result image and the corresponding abnormal cause information of the abnormal image region.

[0185] In another possible implementation manner, the correction processing subunit includes:

[0186] A region correction subunit, configured to perform iterative processing on each abnormal image region in the intermediate result image by using the image generation model based on at least one abnormal image region in the intermediate result image and the corresponding abnormal cause information of the abnormal image region until an intermediate result image that meets the set requirements is obtained.

[0187] In another possible implementation manner, the image correction unit includes:

[0188] A number determination subunit, configured to increment the abnormal processing count by one if the intermediate result image does not meet the set requirements;

[0189] A correction trigger subunit, configured to correct the intermediate result image by using the image generation model if the current abnormal processing count has not exceeded the set count;

[0190] An iteration end subunit, configured to control the image generation model to end the iterative processing of the reference image if the number of abnormal handling times at the current moment exceeds the set number of times.

[0191] In another possible implementation manner, the apparatus further includes:

[0192] A marking determination unit, configured to obtain the area to be adjusted marked by the user in the target image and the adjustment requirement information of the area to be adjusted after the target image is displayed;

[0193] An image adjustment unit, configured to perform iterative processing on the target image by using the image generation model based on the area to be adjusted marked in the target image and the adjustment requirement information of the area to be adjusted, so as to generate an adjusted target image;

[0194] An image output unit, configured to output the adjusted target image.

[0195] In an embodiment of the present application, an electronic device is further provided. As Figure 6 shown, it shows a schematic structural diagram of a composition of the electronic device. The electronic device includes at least a processor 601 and a memory 602;

[0196] The processor 601 is configured to execute the image generation method described in any one of the above embodiments;

[0197] The memory 602 is configured to store a program required for the processor to perform operations.

[0198] It can be understood that the electronic device may further include a display unit 603 and an input unit 604.

[0199] Of course, the electronic device may also have Figure 6 more or fewer components, and this is not limited.

[0200] In an embodiment of the present application, a computer program product is further provided, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement any one of the image generation methods provided in the embodiments of the present application.

[0201] In an embodiment of the present application, a computer-readable storage medium is further provided. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device is enabled to implement any one of the image generation methods provided in the embodiments of the present application.

[0202] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0203] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0204] In the above embodiments, it can be implemented in whole or in part through software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0205] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

Claims

1. A method for generating an image, comprising: Obtaining a reference image and description information input by a user, wherein the description information is used to describe the image content of a target image to be generated; Based on the description information, the reference image is iteratively processed by using an image generation model to obtain an intermediate result image corresponding to each processing; According to the preset detection strategy, determine whether the intermediate result image meets the set requirements; If the intermediate result image meets the set requirements, the iterative processing of the image generation model is continued based on the intermediate result image until the target image is generated and output for display.

2. The image generation method according to claim 1, further comprising: If the intermediate result image does not meet the set requirements, the image generation model is used to correct the intermediate result image to generate a target image that meets the set requirements.

3. The image generation method according to claim 1, wherein determining whether the intermediate result image meets the set requirements according to the preset detection strategy comprises: Obtaining the current iteration number of the image generation model processing the reference image and the intermediate result image obtained by the current iteration processing, if the current iteration number and / or the intermediate result image meet the set conditions, starting from the current iteration number, obtaining at least one intermediate result image obtained by the iteration processing according to the set sampling frequency, and determining whether the obtained intermediate result image meets the set requirements; or, Based on at least one sampling number set by a user, when the current iteration number of the image generation model processing the reference image is the sampling number, an intermediate result image obtained by the current processing is obtained, and it is determined whether the obtained intermediate result image meets the set requirements.

4. The image generation method according to claim 1, wherein determining whether the intermediate result image meets the set requirements comprises: After processing the intermediate result image, the image is sent to a detection device, and a detection result is received from the detection device, wherein the detection result is used to indicate whether the intermediate result image meets the set requirements; Alternatively, a local detection module is used to detect whether the intermediate result image meets the set requirements.

5. The image generation method according to claim 2, wherein determining whether the intermediate result image meets the set requirements comprises at least one of the following: Determine whether the semantic information expressed by the intermediate result image meets the set requirements; Determine whether the image features of the intermediate result image meet the set requirements; If the intermediate result image does not meet the set requirements, the intermediate result image is corrected using the image generation model, including at least one of the following: If it is detected that the semantic information expressed by the intermediate result image does not meet the set requirements, the intermediate result image is discarded, and the process of performing multiple iterations of processing on the reference image using the image generation model based on the description information is returned to regenerate the image; If it is detected that the image features of the intermediate result image do not meet the set requirements, the intermediate result image is corrected using the image generation model based on the image features in the intermediate result image that do not meet the set requirements.

6. The image generation method according to claim 5, wherein if it is detected that the image features of the intermediate result image do not meet the set requirements, based on the image features in the intermediate result image that do not meet the set requirements, the intermediate result image is corrected using the image generation model, comprising: If it is detected that the image features of the intermediate result image do not meet the set requirements, obtaining at least one abnormal image region in the intermediate result image where the image features that do not meet the set requirements are located and abnormal reason information that the image features of the abnormal image region do not meet the set requirements; Based on at least one abnormal image region in the intermediate result image and abnormal cause information corresponding to the abnormal image region, the intermediate result image is corrected using the image generation model.

7. The image generation method according to claim 6, wherein based on at least one abnormal image region in the intermediate result image and abnormal cause information corresponding to the abnormal image region, the image generation model is used to correct the intermediate result image, comprising: Based on at least one abnormal image region in the intermediate result image and abnormal cause information corresponding to the abnormal image region, the image generation model is used to iteratively process each abnormal image region in the intermediate result image until an intermediate result image that meets the set requirements is obtained.

8. The image generation method according to claim 2, wherein if the intermediate result image does not meet the set requirements, the intermediate result image is corrected using the image generation model, comprising: If the intermediate result image does not meet the set requirements, the number of exception processing times is increased by one; If the number of exception processing times at the current moment has not exceeded the set number, the intermediate result image is corrected using the image generation model; If the number of exception processing times at the current moment exceeds the set number, the image generation model is controlled to end the iterative processing of the reference image.

9. The image generating method according to claim 1, further comprising: After displaying the target image, obtaining the area to be adjusted marked by the user in the target image and adjustment requirement information of the area to be adjusted; Based on the to-be-adjusted region marked in the target image and the adjustment requirement information of the to-be-adjusted region, the target image is iteratively processed using the image generation model to generate an adjusted target image; Output the adjusted target image.

10. An image generating device, comprising: An information obtaining unit, used to obtain a reference image and description information input by a user, wherein the description information is used to describe the image content of the target image to be generated; An image processing unit, configured to perform multiple iterative processes on the reference image using an image generation model based on the description information to obtain an intermediate result image corresponding to each process; An image detection unit, used to determine whether the intermediate result image meets the set requirements according to a preset detection strategy; An iterative control unit is used to, if the intermediate result image meets the set requirements, continue to execute the iterative processing of the image generation model based on the intermediate result image until the target image is generated and output for display.