Image generation method, system and device and storage medium

By performing multi-stage training and loss function optimization on the image generation model, the problem of low accuracy of the image generation model is solved, and high-precision image matching and aesthetic enhancement in multiple dimensions are achieved.

CN120707655APending Publication Date: 2025-09-26BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410346530.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-25
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing image generation models are not trained sufficiently, resulting in low accuracy and distortion of generated images.

Method used

By conducting multi-stage training on the image generation model, including noise processing, quality dimension comparison and aesthetic scoring, a loss function is constructed to improve the accuracy of the model in various quality and aesthetic dimensions.

Benefits of technology

The accuracy of the image generation model and the matching degree of the generated images are improved, ensuring that the images are consistent with the original images in multiple quality and aesthetic dimensions, and improving the overall quality of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707655A_ABST
    Figure CN120707655A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and discloses an image generation method, system and device and a storage medium, and the method comprises the steps: obtaining a first image, and carrying out the noise processing of the first image, and obtaining a second image; the second image is input into a trained image generation model, the second image is converted into a third image through the image generation model, and the first image and the third image are matched in multiple quality dimensions representing image quality; and taking the third image as an image generated based on the first image. The image precision can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to an image generation method, system, device, and storage medium. Background Art

[0002] With technological advancements, numerous image generation models have emerged. These models can generate images using various methods. For example, a text-to-image model can generate an image corresponding to a text prompt based on the acquired text prompt. Another example is a picture-to-image model that can perform style transfer on the acquired image, transforming the input image into an image of a specified style.

[0003] Currently, in some technologies, the training of image generation models is not sufficient, resulting in low accuracy of images generated by the image generation models and many defects, such as image distortion.

[0004] Therefore, a method that can improve image accuracy is urgently needed. Summary of the Invention

[0005] In view of this, embodiments of the present disclosure provide an image generation method, a data processing system, an electronic device, and a computer-readable storage medium, which can improve image accuracy.

[0006] In one aspect, the present disclosure provides an image generation method, the method comprising:

[0007] Acquire a first image, and perform noise reduction processing on the first image to obtain a second image;

[0008] inputting the second image into a trained image generation model to convert the second image into a third image through the image generation model, wherein the first image and the third image match in multiple quality dimensions representing image quality;

[0009] The third image is used as an image generated based on the first image.

[0010] In one aspect, the present disclosure provides an image generation system, the system comprising:

[0011] a noise processing module, configured to acquire a first image and perform noise processing on the first image to obtain a second image;

[0012] an image conversion module, configured to input the second image into a trained image generation model to convert the second image into a third image through the image generation model, wherein the first image and the third image match in a plurality of quality dimensions representing image quality;

[0013] An image determination module is configured to use the third image as an image generated based on the first image.

[0014] On the other hand, the present disclosure further provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method described above is implemented.

[0015] On the other hand, the present disclosure further provides an electronic device, which includes a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method described above is implemented.

[0016] Based on the second image generated by noise-processing the first image, the image generation model can match the first image with the third image according to multiple quality dimensions that characterize image quality when converting the second image to the third image. This quality-based approach to matching the first and third images results in a more precise matching process, ensuring a high degree of match between the third image and the first image, thereby improving image accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The features and advantages of the present disclosure will be more clearly understood by referring to the accompanying drawings, which are schematic and should not be construed as limiting the present disclosure in any way. In the accompanying drawings:

[0018] Figure 1 Schematic diagrams showing some image generation models generating images;

[0019] Figure 2 A schematic diagram of a flow chart of a model training method provided by one embodiment of the present application is shown;

[0020] Figure 3 Shown based on Figure 2 Schematic diagram of module interaction for the first stage of training using the model training method in [1].

[0021] Figure 4 A schematic diagram of module interaction of the first stage training provided by another embodiment of the present application is shown;

[0022] Figure 5 A schematic diagram of the process of the second stage training provided by an embodiment of the present application is shown;

[0023] Figure 6 Shown Figure 5 Schematic diagram of module interaction in the second stage of training;

[0024] Figure 7 A schematic diagram of module interaction of the second stage training provided by another embodiment of the present application is shown;

[0025] Figure 8 A schematic diagram showing a flow chart of an image generation method provided by an embodiment of the present application is shown;

[0026] Figure 9 A schematic diagram of a module of an image generation system provided by an embodiment of the present application is shown;

[0027] Figure 10 A schematic diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0028] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0029] See also Figure 1 , schematic diagrams of images generated by some image generation models. Figure 1 In [1], the image generation model can be a diffusion model. Prompts are used to guide the image generation model in generating the desired image. For example, if the image generation model is required to generate an image of three children sitting on a sofa, the prompt might be "Three children on a sofa, full shot." Another example is if the image generation model is required to generate a night scene of the moon and a park, the prompt might be "The moon is very bright in the night sky, casting a gentle light over the abandoned dam usement park."

[0030] The principle of the image generation model is to use the prompt word as a guide, and generate an image by multi-stepping the initial image including random noise. Figure X t Perform noise reduction to obtain an image that matches the prompt word. For example Figure 1 In, X t-1 Indicates the initial Figure X t The image obtained after the first denoising, X t-2 Indicates that in X t-1 Based on the initial Figure X t The image obtained after the second denoising is performed. The image X0 obtained after multiple denoising can be used as the image matching the prompt word.

[0031] It is understandable that based on different prompt words, the initial Figure X t The denoising methods used for denoising can be different, so that the image generation model can generate different images for different prompt words. The training process of the image generation model is to let the image generation model learn the denoising method that matches the prompt word, so that after the image generation model is trained, the image generation model can denoise the initial image according to the denoising method that matches the prompt word based on the input prompt word and the learned denoising knowledge. Figure X t Perform noise reduction to generate an image that matches the prompt word.

[0032] At present, the training of image generation models is not sufficient, the accuracy of the trained image generation models is low, and the denoising of the initial images is not accurate enough, which makes the final generated image not match the prompt word, or the image is distorted, resulting in low accuracy of the generated image.

[0033] In view of this, the present application provides a model training method that can improve the accuracy of the image generation model, and thus improve the accuracy of the generated image. The model training method can be applied to electronic devices. Among them, electronic devices include but are not limited to desktop computers, laptops, tablet computers, servers, etc. It should be noted that the image generation model trained using the model training method of the present application can be an image generation model that has completed initial training. The so-called initial training is to train the image generation model based on the prompt words and the initial image including random noise. The image generation model after completing the initial training has the basic ability to denoise images including noise. However, if Figure 1 As described above, the image generation model after the initial training may have the problem of low accuracy. Therefore, the image generation model after the initial training can be further trained based on the model training method of the present application to fine-tune the image generation model, thereby achieving the purpose of improving the model accuracy and then improving the accuracy of the images generated by the image generation model.

[0034] See also Figure 2 and Figure 3 . Figure 2 A flowchart of a model training method provided for one embodiment of the present application. Figure 3 Based on Figure 2 Schematic diagram of the module interaction for the first stage of training using the model training method in . Figure 2 In [1], the model training method includes the following steps:

[0035] Step S21 : obtaining a first sample image and performing noise reduction processing on the first sample image to obtain a second sample image.

[0036] Specifically, the first sample image may be a standard image that needs to be generated by the image generation model. After the first sample image is acquired, random noise may be added to the first sample image to obtain the second sample image.

[0037] Step S22: input the second sample image into the image generation model, so as to convert the second sample image into a third sample image through the image generation model.

[0038] Specifically, the third sample image is an image obtained by denoising the second sample image based on the denoising knowledge learned by the image generation model during the initial training process.

[0039] It is understandable that if the image generation model is highly accurate and the noise reduction of the second sample image is accurate, the third sample image should be consistent with the first sample image, or the difference between the third sample image and the first sample image should be within an acceptable threshold range. However, if the image generation model is not highly accurate and the noise reduction of the second sample is inaccurate, the third sample image may differ significantly from the first sample image. Therefore, the difference between the third sample image and the first sample image can be used to evaluate the accuracy of the image generation model. Simply put, the smaller the difference between the third sample image and the first sample image, the higher the accuracy of the image generation model; and the greater the difference between the third sample image and the first sample image, the lower the accuracy of the image generation model.

[0040] Step S23 : comparing the first sample image and the third sample image to obtain quality difference values ​​of the first sample image and the third sample image in each quality dimension.

[0041] Specifically, the quality dimension is used to determine the image quality of the third sample image. The image quality of the third sample image represents the degree of similarity between the third sample image and the first sample image. A higher degree of similarity between the third sample image and the first sample image indicates better image quality, while a lower degree of similarity between the third sample image and the first sample image indicates worse image quality. When comparing the third sample image and the first sample image for similarity, the third sample image and the first sample image can be compared along multiple dimensions. The multiple dimensions used herein can be referred to as quality dimensions.

[0042] In this embodiment, the quality dimension may include an instance dimension and an image style dimension. Comparing the similarity between the third sample image and the first sample image from the instance dimension may involve comparing the similarity between the instance segmentation results of the third sample image and the first sample image. Comparing the similarity between the third sample image and the first sample image from the image style dimension may involve comparing the similarity between the image styles of the third sample image and the first sample image. It is understood that the selection can be based on actual needs. In this embodiment, dividing the quality dimension into the instance dimension and the image style dimension does not constitute a limitation to this application.

[0043] The difference between the third sample image and the first sample image in each quality dimension can be represented by a quality difference value. Specifically, if the difference between the third sample image and the first sample image in a quality dimension is large, then the quality difference value between the third sample image and the first sample image in that quality dimension can be large. Conversely, if the difference between the third sample image and the first sample image in a quality dimension is small, then the quality difference value between the third sample image and the first sample image in that quality dimension can be correspondingly small.

[0044] In this embodiment, the quality difference value includes an instance difference value between the third sample image and the first sample image in the instance dimension, and a style difference value between the third sample image and the first sample image in the image style dimension.

[0045] The following takes the instance dimension and the image style dimension as an example to illustrate how to compare the third sample image with the first sample image to obtain the quality difference value between the third sample image and the first sample image in each quality dimension.

[0046] Specifically, in some embodiments, after obtaining the first sample image in step S21, the first sample image may be instance-labeled so that the first sample image has instance segmentation labels. Based on this, from the instance dimension, the above-mentioned comparison of the first sample image and the third sample image may include:

[0047] Performing instance segmentation on the third sample image to obtain an instance segmentation result of the third sample image;

[0048] The difference value between the instance segmentation annotation and the instance segmentation result is used as the instance difference value between the third sample image and the first sample image in the instance dimension.

[0049] See also Figure 3 , performing instance segmentation on the third sample image, and determining the difference between the instance segmentation result of the third sample image and the instance segmentation annotation of the first sample image, can be achieved by using a trained instance segmentation model. Performing instance segmentation on an image is a technique that should be known to those skilled in the art and will not be described in detail here. In this application:

[0050] Use m I (x ′ 0) represents the instance segmentation result of the third sample image, where m I represents the instance segmentation network for instance segmentation of the third sample image, x ′ 0 represents the third sample image;

[0051] Use GT(x0) to represent the instance segmentation annotation of the first sample image, where x0 represents the first sample image;

[0052] Use L instance (m I (x ′ 0), GT(x0)) represents the quality difference value (i.e., instance difference value) between the third sample image and the first sample image in the instance dimension. If the instance segmentation result of the third sample image differs greatly from the instance segmentation annotation of the first sample image, L instance (m I (x ′ 0), the value of GT(x0)) can be larger; if the difference between the instance segmentation result of the third sample image and the instance segmentation annotation of the first sample image is small, L instance (m I (x ′ 0), the value of GT(x0)) can be smaller.

[0053] Furthermore, in some embodiments, from the perspective of image style, the comparison of the first sample image and the third sample image may include:

[0054] extracting a first style value representing a style of the first sample image from the first sample image, and extracting a second style value representing a style of the third sample image from the third sample image;

[0055] The difference between the first style value and the second style value is used as the style difference between the first sample image and the third sample image in the image style dimension.

[0056] Specifically, image styles can be divided according to actual needs, for example, image styles can be divided into retro style and modern style. Different image styles can have different style values. Figure 3 Extracting the style values ​​of the first and third sample images, and calculating the style difference between the first and third sample images, can be achieved using a trained image style extraction model. Image style extraction is a technique well known to those skilled in the art and will not be described in detail here.

[0057] In this application:

[0058] Gram(V(x0)) is used to represent the first style value of the first sample image style, where x0 represents the first sample image, V represents the image style extraction network for extracting image style features, and Gram represents the use of the Gram matrix to calculate the extracted image style features. The calculated result can be used as the second style value;

[0059] Using Gram(V(x ′ 0)) represents the second style value of the third sample image style, where x ′ 0 represents the third sample image. For V and Gram, please refer to the relevant description of Gram(V(x0)), which will not be repeated here.

[0060] Use ||Gram(V(x ′ 0))-Gram(V(x0))||2 represents the quality difference value (i.e., style difference value) between the third sample image and the first sample image in the image style dimension. If the image style of the third sample image is significantly different from that of the first sample image, ||Gram(V(x ′ 0))-Gram(V(x0))||2 can be larger; if the image style of the third sample image is slightly different from that of the first sample image, ||Gram(V(x ′ 0))-Gram(V(x0))||2 can take a smaller value.

[0061] In step S24 , if the quality difference value indicates that the first sample image and the third sample image do not match, the image generation model is trained in the first stage according to the quality difference value.

[0062] Specifically, in this embodiment, whether the first sample image and the third sample image match can be determined separately according to the quality difference value of each quality dimension. For example, for the instance dimension, if the instance difference value of the first sample image and the third sample image in the instance dimension is less than a first threshold value (such as 0.5), it can be indicated that the first sample image and the third sample image match in the instance dimension; if the instance difference value of the first sample image and the third sample image in the instance dimension is not less than the first threshold value, it can be indicated that the first sample image and the third sample image do not match in the instance dimension. For the image style dimension, if the style difference value of the first sample image and the third sample image in the image style dimension is less than a second threshold value (such as 0.3), it can be indicated that the first sample image and the third sample image match in the image style dimension; if the instance difference value of the first sample image and the third sample image in the image style dimension is not less than the second threshold value, it can be indicated that the first sample image and the third sample image do not match in the image style dimension.

[0063] The image generation model can be trained in the first phase based on the quality difference values ​​for each quality dimension. For example, if the instance difference value for the instance dimension determines that the first sample image and the second sample image do not match, the image generation model can be trained in the first phase based on the instance difference value to ensure that the first sample image and the third sample image match in the instance dimension. Similarly, if the style difference value for the image style dimension determines that the first sample image and the second sample image do not match, the image generation model can be trained in the first phase based on the style difference value to ensure that the first sample image and the third sample image match in the image style dimension.

[0064] Furthermore, if the first sample image and the second sample image match in one of the quality dimensions, then the image generation model does not need to be trained in the first stage based on the quality difference value of that quality dimension. For example, if the first sample image and the second sample image are determined to be mismatched based on the instance difference value of the instance dimension, but are matched based on the style difference value of the image style dimension, then the image generation model can be trained in the first stage based solely on the instance difference value, without having to train the image generation model in the first stage based on the style difference value.

[0065] In this embodiment, a loss function as shown in Expression (1) can be constructed based on the instance difference value to train the image generation model from the instance dimension:

[0066]

[0067] Where E represents the expectation, D represents the first sample image set, represents the third sample image set.

[0068] Furthermore, based on the style difference value, a loss function as shown in Expression (2) can be constructed to train the image generation model from the image style dimension:

[0069]

[0070] It is understandable that after the image generation model is trained from different quality dimensions, the trained image generation model can generate a third sample image that matches the first sample image according to each quality dimension. Compared to some technologies that compare the first sample image and the third sample image as a whole and train the image generation model based on the results of the overall comparison, the present application splits and compares the first sample image and the third sample image according to each quality dimension, and based on the results of the split comparison, trains the image generation model from each quality dimension. In this way, the training of the image generation model is more refined, so that the trained image generation model can have higher accuracy, which in turn can improve the accuracy of the generated image.

[0071] See also Figure 4 , is a schematic diagram of module interaction for the first phase of training provided in another embodiment of the present application. In some embodiments, the first phase of training may also include:

[0072] Inputting the third sample image into the sub-scoring model, so as to score the third sample image according to multiple aesthetic dimensions representing the beauty of the image through the sub-scoring model, thereby obtaining an aesthetic score of the third sample image in each aesthetic dimension;

[0073] If the aesthetic score does not exceed the score threshold, the image generation model is trained based on the aesthetic score.

[0074] Specifically, the aesthetic dimension includes, but is not limited to, image lighting effects, image color, image layout, etc. This application does not impose any restrictions on the division of the aesthetic dimension. It should be noted that the aesthetic dimension is a dimensional division of the third sample image from the perspective of its aesthetics, while the above-mentioned quality dimension is a dimensional division of the third sample image from the perspective of evaluating the similarity between the third sample image and the first sample image. Therefore, the aesthetic dimension and the quality dimension may be the same or different.

[0075] See also Figure 4 In this embodiment, each aesthetic dimension may correspond to a trained sub-scoring model, which is used to score the aesthetics of the third sample image in the corresponding aesthetic dimension. For example, Figure 4 In the example, the illumination scoring model may be used to score the illumination effect of the third sample image, and the color scoring model may be used to score the color of the third sample image.

[0076] For any sub-scoring model, the higher the aesthetic score output by that sub-scoring model, the better the aesthetic quality of the third sample image in the corresponding aesthetic dimension. Based on the aesthetic scores output by each sub-scoring model, the image generation model can be trained separately so that the aesthetic scores of the third sample images generated by the trained image generation model meet the scoring threshold in each aesthetic dimension. This improves the aesthetic quality of the third sample images.

[0077] Specifically, based on the aesthetic scores output by each sub-scoring model, a loss function as shown in Expression (3) can be constructed to train the image generation model from the aesthetic dimension:

[0078]

[0079] Among them, c represents the prompt word used to guide the image generation model to generate the third sample image, x ′ 0 represents the third sample image, r d represents the sub-scoring model, r d (x ′ 0,c) represents the beauty score for the third sample image, α d Represents the scoring threshold, and the summation brackets represent scoring of multiple aesthetic dimensions of the third sample image. In expression (3), for any aesthetic dimension, the aesthetic score r of the third sample image in the aesthetic dimension is d (x ′ 0,c) reaches the score threshold α d When , the aesthetic rating r based on the aesthetic dimension can be stopped d (x ′ 0,c) Train the image generation model.

[0080] After completing the first phase of training, the image generation model has already achieved relatively good performance in terms of image quality and image aesthetics. After the first phase of training, the image generation model can be further trained using the following methods to further improve its inference speed.

[0081] For details, please refer to Figure 5 and Figure 6 . Figure 5 A flowchart of the second stage training is provided for one embodiment of the present application. Figure 6 for Figure 5 Schematic diagram of module interaction in the second stage of training.

[0082] Figure 5 The second stage of training may specifically include the following steps:

[0083] Step S51 : inputting the first sample image and the third sample image into a trained total score model, so as to output a first image score of the first sample image and a second image score of the third sample image through the total score model.

[0084] Specifically, the image score can represent the quality of the sample image. The overall score model can be a reward model. A higher image score output by the overall score model indicates a better image, and a lower image score output by the overall score model indicates a worse image.

[0085] In this embodiment, the image score can be a score obtained using the first sample image as a benchmark. Simply put, the first sample image is considered the best image, meaning that the first sample image's first image score can be the highest. The higher the similarity between the third sample image and the first sample image, the higher the second image score can be. Conversely, the lower the similarity between the third sample image and the first sample image, the lower the second image score can be.

[0086] In this application, the use of r a (x0) represents the first image score output by the total score model for the first sample image, and the use of r a (x ′ 0, c) represents the second image score output by the total score model for the third sample image, where r a represents the total score model, x0 represents the first sample image, x ′ 0 represents the third sample image, and c represents the prompt word that guides the image generation model to generate the third sample image.

[0087] Step S52: Based on the second image score, the image generation model is trained in the second stage to increase the second image score of the third sample image output by the image generation model.

[0088] Specifically, based on the second image score, a loss function as shown in Expression (4) can be constructed to train the image generation model:

[0089]

[0090] Among them, in expression (4), in r d1 (x ′ 0,c) before adding a negative sign means: in the second image score r a (x ′ 0,c), the higher the loss of the image generation model.

[0091] Step S53: Based on the first image score and the second image score, the total score model is trained so that the first image score output by the total score model is increased and the second image score is decreased.

[0092] It can be understood that after the image generation model is trained based on expression (4), the third sample image generated by the image generation model will become better and better, that is, the second image score r d (x ′ 0, c) will become higher and higher. In this case, in order to make the third sample image generated by the image generation model better and better, the total scoring model can be trained based on the first image score and the second image score, so that the total scoring model has more and more stringent requirements for the third sample image. Correspondingly, under the stricter standards, the second image score of the third sample image by the total scoring model should be lower. In this way, the image generation model can continue to be trained under the lower second image score. This training process can also be called an adversarial training process. Specifically, in the adversarial training process, the total scoring model can regard the first sample image as a good image, output a high first image score for the first sample image, and regard the third sample image as a bad image, output a low second image score for the third sample image. In this way, the accuracy of the image generation model can be improved through such adversarial training.

[0093] Specifically, in this embodiment, a loss function as shown in Expression (5) can be constructed based on the first image score and the second image score to train the total score model:

[0094]

[0095] Among them, σ is the adjustment parameter.

[0096] In the second stage of training, the third sample image with a low number of denoising steps can also receive sufficient feedback optimization, so that after completing the second stage of training, the third sample image obtained by the image generation model using fewer denoising steps can also have better image quality, which is equivalent to speeding up the inference speed of the image generation model.

[0097] See also Figure 7 , which is a schematic diagram of module interaction for the second stage training provided in another embodiment of the present application. Figure 7 In [1], while training the image generation model based on the overall scoring model, the image generation model can also be trained based on multiple sub-scoring models. This can further improve the accuracy of the image generation model. The principle of the sub-scoring model can be found in [1]. Figure 4 The description is not repeated here.

[0098] So far, the entire description of the first and second stage training has been completed. Based on the image generation model after the first and second stage training, this application provides an image generation method that can improve image accuracy. The image generation method can be applied to electronic devices. Electronic devices include but are not limited to tablet computers, laptop computers, tablet computers, servers, etc. Figure 8 , which is a flow chart of an image generation method provided in one embodiment of the present application. Figure 8 In the method for generating an image, the image comprises the following steps:

[0099] Step S81: Acquire a first image and perform noise reduction processing on the first image to obtain a second image.

[0100] Specifically, step S81 is similar to step S21 and will not be described in detail here.

[0101] Step S82: Input the second image into the trained image generation model to convert the second image into a third image through the image generation model, wherein the first image and the third image match in multiple quality dimensions representing image quality.

[0102] Specifically, the first image and the third image match in multiple quality dimensions representing image quality, which may mean that the similarity between the first image and the third image in each quality dimension is higher than a similarity threshold.

[0103] It can be understood that since the image generation model of the present application is trained according to various quality dimensions, the third image and the first image generated by it can match each quality dimension with high accuracy.

[0104] Step S83: Use the third image as an image generated based on the first image.

[0105] Furthermore, in some embodiments, when the first image and the third image match in multiple quality dimensions, the first image and the third image may also satisfy the following conditions:

[0106] The first image and the third image match in multiple aesthetic dimensions representing the aesthetics of the images.

[0107] The relevant principles can be found in the training process of the image generation model, which will not be repeated here.

[0108] The image generation model of this application can be applied in multiple image generation fields, such as image generation and text generation, and can generate images with high precision. In the text generation field, the image generation model can use the input prompt word as a guide to denoise the initial image containing random noise, thereby obtaining an image that matches the prompt word. Because the image generation model of this application performs well in the denoising process, the generated images are highly accurate.

[0109] In summary, in the technical solutions of some embodiments of the present application, based on the second image obtained by noise-processing the first image, the image generation model can match the first image with the third image according to multiple quality dimensions that characterize image quality when converting the second image into the third image. This method of matching the first and third images according to quality dimensions provides a more precise matching process, allowing the third image to have a high degree of match with the first image, thereby improving image accuracy.

[0110] See also Figure 9 , which is a module diagram of an image generation system provided in one embodiment of the present application. Figure 9 In [1], the image generation system includes:

[0111] a noise processing module, configured to acquire a first image and perform noise processing on the first image to obtain a second image;

[0112] an image conversion module, configured to input the second image into a trained image generation model to convert the second image into a third image through the image generation model, wherein the first image and the third image match in a plurality of quality dimensions representing image quality;

[0113] The image determination module is configured to use the third image as an image generated based on the first image.

[0114] See also Figure 10 , is a schematic diagram of an electronic device provided by an embodiment of the present application. The electronic device includes a processor and a memory, the memory being used to store a computer program, and when the computer program is executed by the processor, the above method is implemented.

[0115] The processor may be a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination of the above chips.

[0116] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer executable programs, and modules, such as the program instructions / modules corresponding to the methods in the embodiments of the present disclosure. The processor executes the non-transitory software programs, instructions, and modules stored in the memory to perform various processor functions and data processing, thereby implementing the methods in the aforementioned method embodiments.

[0117] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include a memory remotely located relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0118] One embodiment of the present application further provides a computer-readable storage medium, which is used to store a computer program. When the computer program is executed by a processor, the above method is implemented.

[0119] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. An image generation method, characterized in that: The method comprises: Acquire a first image, and perform noise reduction processing on the first image to obtain a second image; inputting the second image into a trained image generation model to convert the second image into a third image through the image generation model, wherein the first image and the third image match in multiple quality dimensions representing image quality; The third image is used as an image generated based on the first image.

2. The method according to claim 1, wherein When the first image and the third image match in the multiple quality dimensions, the first image and the third image further satisfy the following condition: The first image and the third image match in multiple aesthetic dimensions representing image beauty.

3. The method according to claim 1 or 2, wherein: The image generation model is trained in the following way: Acquire a first sample image, and perform noise processing on the first sample image to obtain a second sample image; inputting the second sample image into the image generation model to convert the second sample image into a third sample image through the image generation model; Comparing the first sample image with the third sample image to obtain a quality difference value between the first sample image and the third sample image in each quality dimension; If the quality difference value indicates that the first sample image and the third sample image do not match, the image generation model is trained in a first stage according to the quality difference value.

4. The method according to claim 3, wherein The quality dimension includes an instance dimension, the quality difference value includes an instance difference value between the first sample image and the third sample image in the instance dimension, and the first sample image has an instance segmentation annotation; The comparing the first sample image and the third sample image includes: Performing instance segmentation on the third sample image to obtain an instance segmentation result of the third sample image; The difference value between the instance segmentation annotation and the instance segmentation result is used as the instance difference value between the first sample image and the third sample image in the instance dimension.

5. The method according to claim 3, wherein The quality dimension includes an image style dimension, and the quality difference value includes a style difference value between the first sample image and the third sample image in the image style dimension; The comparing the first sample image and the third sample image includes: extracting a first style value representing a style of the first sample image from the first sample image, and extracting a second style value representing a style of the third sample image from the third sample image; The difference value between the first style value and the second style value is used as the style difference value between the first sample image and the third sample image in the image style dimension.

6. The method according to claim 3, wherein After the first stage of training, the image generation model is further trained as follows: Inputting the first sample image and the third sample image into a trained total score model, so as to output a first image score of the first sample image and a second image score of the third sample image through the total score model; performing a second-stage training on the image generation model based on the first image score, so as to increase the first image score of the third sample image output by the image generation model; The total score model is trained based on the first image score and the second image score, so that the first image score output by the total score model is reduced and the second image score is increased.

7. The method according to claim 6, wherein In the first phase training and the second phase training, at least one phase training further includes: Inputting the third sample image into a sub-scoring model, so as to score the third sample image according to a plurality of aesthetic dimensions representing the beauty of the image by the sub-scoring model, thereby obtaining an aesthetic score of the third sample image in each of the aesthetic dimensions; If the aesthetic score does not exceed the score threshold, the image generation model is trained according to the aesthetic score.

8. An image generation system, characterized in that: The system comprises: a noise processing module, configured to acquire a first image and perform noise processing on the first image to obtain a second image; an image conversion module, configured to input the second image into a trained image generation model to convert the second image into a third image through the image generation model, wherein the first image and the third image match in a plurality of quality dimensions representing image quality; An image determination module is configured to use the third image as an image generated based on the first image.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory is used to store a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.