Image generation method and device
By combining the user's preferred image features and text attention denoising in the target database during the image generation process, the randomness of the image generation tool is solved, and an accurate image that meets the user's expectations is achieved.
Patent Information
- Application Number
- CN202510400662.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The existing image generation tools have randomness when generating the pet cat pictures that users expect, and cannot fully meet the user's expectations, resulting in inaccurate images.
By performing image attention denoising using the user preference image features stored in the target database in the image generation request, and performing text attention denoising in combination with prompt information, a picture conforming to the user's preference is generated.
Improves the accuracy of generated images, reduces the probability of regeneration, and ensures that the generated images meet user expectations.
Smart Images

Figure CN120259472A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular, to an image generation method and apparatus. Background Art
[0002] When generating pictures using picture generation algorithms, there is a certain degree of randomness. For example, if a user has a pet cat at home and uses an image generation tool to generate a picture of the cat, the user will expect to generate a picture of the pet cat. However, the image generation tool usually randomly generates pictures of cats, such as Ragdoll cats, orange cats, etc. The finally generated pictures of cats do not fully meet the user's expectations. Summary of the Invention
[0003] In view of this, this application provides an image generation method and apparatus, and the specific solutions are as follows:
[0004] An image generation method includes:
[0005] Obtain an image generation request; determine the prompt information corresponding to the image generation request;
[0006] If it is determined that the target database stores a target label matching the prompt information, determine the target image features stored in the target database corresponding to the target label;
[0007] Perform image attention denoising based on the target image features stored in the target database to obtain a first denoising result;
[0008] Perform text attention denoising based on the prompt information to obtain a second denoising result;
[0009] Generate a target image based on the first denoising result and the second denoising result.
[0010] Further, the performing image attention denoising based on the target image features stored in the target database to obtain a first denoising result includes:
[0011] Extract the target image features stored in the target database;
[0012] Perform vector conversion on the target image features to obtain a target image feature vector;
[0013] Input the target image feature vector into an image attention denoising model structure, and perform image attention denoising on the target image feature vector to obtain a first denoising result.
[0014] Further, the performing text attention denoising based on the prompt information to obtain a second denoising result includes:
[0015] Perform vector conversion on the prompt information to obtain a text feature vector;
[0016] Input the text feature vector into the text attention denoising model structure, and perform text attention denoising on the text feature vector to obtain a second denoising result.
[0017] Further, the generating the target image based on the first denoising result and the second denoising result includes:
[0018] Fuse the first denoising result and the second denoising result to obtain a fusion result;
[0019] Perform image and text joint denoising on the fusion result to obtain a third denoising result;
[0020] Generate the target image using the third denoising result.
[0021] Further, it further includes:
[0022] Display the target image;
[0023] Obtain feedback information of the user on the displayed target image;
[0024] Update the target label and / or the target image feature in the target database based on the feedback information.
[0025] Further, the updating the target label and / or the target image feature in the target database based on the feedback information includes:
[0026] If the feedback information indicates that the target image meets the conditions, update at least part of the image features of the target image to the target image features corresponding to the target label in the target database;
[0027] If the feedback information indicates that the target image does not meet the conditions, delete at least part of the target image features corresponding to the target label in the target database.
[0028] Further, the if the feedback information indicates that the target image does not meet the conditions, deleting at least part of the target image features corresponding to the target label in the target database includes:
[0029] If the feedback information indicates that the target image does not meet the conditions, determine the image features of the target image;
[0030] Delete the features corresponding to the image features of the target image from the target image features corresponding to the target label in the target database.
[0031] Further, it further includes:
[0032] Obtain multiple pre-stored images;
[0033] Analyze each of the multiple images to determine the image features and image categories of each image, where different images may correspond to the same or different image categories;
[0034] Determine labels based on the image categories;
[0035] Cluster the image features of images belonging to the same image category to obtain the clustered image features of the same image category;
[0036] Establish an association relationship between the clustered image features of the same image category and the labels corresponding to the same image category, and form a target database storing the association relationship between the image features and labels corresponding to at least one image category.
[0037] Further, it further includes:
[0038] If it is determined that the target database does not store a target label matching the prompt information, generate the target image according to the prompt information.
[0039] An image generation device, including:
[0040] An application module, configured to receive an image generation request and determine the prompt information corresponding to the image generation request;
[0041] A storage module, configured to store a target database, where the target database records the association relationship between image features and labels;
[0042] A determination module, configured to determine the target image features corresponding to the target label recorded in the target database when it is determined that the target database stored in the storage module records a target label matching the prompt information;
[0043] An image attention denoising model structure, configured to perform image attention denoising based on the target image features recorded in the target database to obtain a first denoising result;
[0044] A text attention denoising model structure, configured to perform text attention denoising based on the prompt information to obtain a second denoising result;
[0045] A generation module, configured to generate a target image based on the first denoising result and the second denoising result. Description of the Drawings
[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the following will briefly introduce the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 Flowchart of an image generation method disclosed in an embodiment of the present application;
[0048] Figure 2 Flowchart of an image generation method disclosed in an embodiment of the present application;
[0049] Figure 3 Flowchart of an image generation method disclosed in an embodiment of the present application;
[0050] Figure 4 Flowchart of an image generation method disclosed in an embodiment of the present application;
[0051] Figure 5 Schematic diagram of a complete image generation method disclosed in an embodiment of the present application;
[0052] Figure 6 Schematic diagram of the structure of an image generation device disclosed in an embodiment of the present application. Detailed implementation manners
[0053] The following describes the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application.
[0054] The following describes the embodiments of the present application in conjunction with the drawings. Those of ordinary skill in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.
[0055] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these process, method, product or device.
[0056] This application discloses an image generation method, and its flowchart is as follows Figure 1 shown, including:
[0057] Step S101, obtain an image generation request, and determine the prompt information corresponding to the image generation request;
[0058] Step S102, if it is determined that there is a target label stored in the target database that matches the prompt information, determine the target image features corresponding to the target label stored in the target database;
[0059] Step S103, perform image attention denoising based on the target image features stored in the target database to obtain a first denoising result;
[0060] Step S104, perform text attention denoising based on the prompt information to obtain a second denoising result;
[0061] Step S105, generate a target image based on the first denoising result and the second denoising result.
[0062] When generating pictures using a picture generation algorithm, there will be a certain degree of randomness. For example, if a user has a pet cat at home and uses an image generation tool to generate a picture of a cat, they will expect to generate a picture of their pet cat. However, the image generation tool usually generates pictures of cats randomly, such as Ragdoll cats, orange cats, etc. The finally generated picture of the cat may not fully meet the user's expectations.
[0063] Based on this, in this solution, when obtaining an image generation request, an image is generated based on the image generation request with reference to the information stored in the target database, and the information stored in the target database is pre-stored by the user and conforms to the user's preferences. Therefore, the generated image meets the image generation request while conforming to the user's preferences, ensuring the accuracy of the generated image.
[0064] Specifically, when obtaining an image generation request, analyze the image generation request to determine the prompt information corresponding to the image generation request. Among them, the prompt information is the key description information of the image to be generated. For example, if the image generation request is: "Please generate an image of a cat", then the corresponding prompt information can be "a cat" or "cat".
[0065] A database can be set up in advance. There can be one or multiple databases. When there is only one database, then this database is the target database. If there are multiple databases, one of the multiple databases needs to be selected as the target database.
[0066] Among them, selecting one from the multiple databases as a target database may be: selecting a database including the category specified in the prompt information from the multiple databases as the target database, such as: if the prompt information is "cat", then it is necessary to select a database including the image features of the category "cat" from the multiple databases as the target database;
[0067] Alternatively, the target database may be determined based on the user of the output image generation request, that is, the user or user identifier of the output image generation request is determined, and a database corresponding to the user or user identifier is selected from multiple databases as the target database. In this case, different databases need to correspond to different users or user identifiers.
[0068] In addition, the image generation method disclosed in this embodiment can be implemented based on a large model. The large model can log in to different accounts. Different users use different accounts. Correspondingly, different users correspond to different databases, and different accounts of the large model correspond to different databases. When the large model logs in to the account of the first user and obtains an image generation request, the large model can determine the database of the first user based on the account of the first user, and determine the database of the first user as the target database; when the large model logs in to the account of the second user and obtains an image generation request, the large model can determine the database of the second user based on the account of the second user, and determine the database of the second user as the target database.
[0069] The database stores multiple correspondences, each of which represents the correspondence between a label and an image feature. A label is a label of a certain category. When a label in the database is determined based on the prompt information, it indicates that the label corresponds to the category of the image to be generated specified in the prompt information. For example, if the prompt information is "cat", and the target label in the target database is determined based on the prompt information, then the target label corresponds to the category of "cat" or "animal", and the target image feature corresponding to the target label is the image feature of the category.
[0070] In addition, the images based on the corresponding relationships stored in the target database are images determined according to user preferences, such as: the target database can be generated based on the user's photo album, and the images in the user's photo album are all images of objects that the user is interested in; or, the target database is generated based on images selected by the user himself. Therefore, when determining the target image features based on the corresponding relationships in the target database and further generating the target image, the user's preferences are taken into consideration for image generation to ensure that the final generated image can meet the user's needs with a high probability, so as to reduce the probability that the generated image does not meet the user's needs and the image needs to be regenerated.
[0071] After determining the target image features in the target database based on the prompt information, image attention denoising can be performed based on the target image features to obtain a first denoising result. The attention mechanism can be used to distinguish noise from effective features, and image reconstruction can be achieved by combining frequency domain analysis, residual learning, and multi-stage optimization.
[0072] During the image attention denoising process, it includes the screening of image features. Among them, the screening of image features can be: frequency domain screening, or threshold screening, etc. If the screening of image features includes frequency domain screening, then frequency domain analysis is performed on the image features to obtain a frequency domain analysis result, and frequency domain screening is performed based on the frequency domain analysis result. For example, high-frequency noise (such as: salt-and-pepper noise, that is, impulse noise in the image, which appears as randomly appearing white or black dots in the image, moiré pattern, that is, high-frequency interference stripes in the image) is usually identified and suppressed in the frequency domain, so the high-frequency noise in the image features is determined and the high-frequency noise is filtered; if the screening of image features includes threshold screening, then a specific threshold is determined, the attention weights of each image feature are determined, and the image features corresponding to the attention weights lower than the specific threshold are screened out.
[0073] In addition, while performing image attention denoising using the target image features, or after determining the prompt information, text attention denoising is performed based on the prompt information to obtain a second denoising result.
[0074] During the text attention denoising process, it includes the screening of text features, and the screening conditions of the text features can include: semantic relevance, attention dynamic weight, noise distribution characteristics, and training objective guidance, etc.
[0075] After performing image attention denoising using the target image features to obtain a first denoising result and performing text attention denoising using the prompt information to obtain a second denoising result respectively, a target image is generated based on the first denoising result and the second denoising result, realizing the separate execution of image attention denoising and text attention denoising. During the separate execution of image attention denoising and text attention denoising, the result obtained through text attention denoising can meet the image generation requirements, and the result obtained through image attention denoising can meet the user's preferences. Thus, it is ensured that the finally generated target image can meet the user's preferences on the basis of meeting the image generation requirements, ensuring the accuracy of the generated image and reducing the probability of regenerating the image.
[0076] The image generation method disclosed in this embodiment, when obtaining an image generation request, determines the prompt information corresponding to the image generation request, and searches in the target database to find whether there is a target label matching the prompt information. If so, it determines the target image features corresponding to the target label stored in the target database, and performs image attention denoising based on the target image features to obtain a first denoising result. In addition, it performs text attention denoising based on the prompt information to obtain a second denoising result. Finally, it generates a target image based on the first denoising result and the second denoising result. In this solution, when there is an image generation request, on the one hand, it uses the target image features stored in the target database to perform image attention denoising, and on the other hand, it uses the prompt information corresponding to the image generation request to perform text attention denoising. Among them, the image features stored in the target database are pre-stored by the user and conform to the user's preferences. By performing image attention denoising and text attention denoising respectively, it is ensured that when generating an image, both the image generation request and the information in the pre-stored target database can be used as reference factors in the image generation process, so that the finally generated image can conform to the user's preferences and improve the accuracy of the generated image.
[0077] This embodiment discloses an image generation method, and its flowchart is as Figure 2 shown, including:
[0078] Step S201, obtain an image generation request and determine the prompt information corresponding to the image generation request;
[0079] Step S202, if it is determined that there is a target label matching the prompt information stored in the target database, determine the target image features corresponding to the target label stored in the target database;
[0080] Step S203, extract the target image features stored in the target database;
[0081] Step S204, perform vector conversion on the target image features to obtain a target image feature vector;
[0082] Step S205, input the target image feature vector into the image attention denoising model structure, and perform image attention denoising on the target image feature vector to obtain a first denoising result;
[0083] Step S206, perform text attention denoising based on the prompt information to obtain a second denoising result;
[0084] Step S207, generate a target image based on the first denoising result and the second denoising result.
[0085] After obtaining an image generation request and determining the corresponding prompt information based on the image generation request, perform Image Attention Denoise using the target image features in the target database and perform Words Attention Denoise using the prompt information, and generate a target image based on the denoising results obtained after denoising respectively, ensuring that the finally generated target image meets both the image generation request and the user preferences represented by the target database.
[0086] Among them, the process of performing image attention denoise using the target image features in the target database can be specifically as follows:
[0087] After determining the target label in the target database that matches the prompt information, determine the target image features corresponding to the target label in the target database. To perform image attention denoise using this target image feature, it is necessary to first extract this target image feature from the target database, and then it can be processed.
[0088] After extracting the target image features, perform vectorization processing on the target image features to obtain a target image feature vector. This target image feature vector is an expression in the form of noise. Only when expressed in the form of noise can denoising processing be performed.
[0089] Finally, input the target image feature vector into the image attention denoise model structure, and use the image attention denoise model structure to perform image attention denoise to obtain a first denoising result.
[0090] Among them, the process of the image attention denoise model structure performing image attention denoise can be: based on the self-attention or hybrid attention mechanism, dynamically evaluate the importance of the target image feature vector to adjust the attention weight of the target image feature vector; then, filter some target image feature vectors based on filtering conditions and strategies, such as: performing frequency domain analysis and performing frequency domain screening based on the frequency domain analysis results, or: performing threshold screening based on a weight threshold, or using a loss function to guide the screening of the target image feature vector; then, strengthen the retained effective target image feature vectors, such as: retaining shallow details through residual learning.
[0091] If an image can be directly generated after performing image attention denoise only using the target image features, however, the image generated only using the target image features may not fully meet the image generation request input by the user because the prompt information is not considered. Therefore, it is also necessary to add denoising based on the prompt information and combine the result obtained by denoising based on the target image features with the result obtained by denoising based on the prompt information to finally obtain a target image that meets both the user's preferences and the image generation request.
[0092] Specifically, performing text attention denoising based on the prompt information to obtain a second denoising result can be specifically as follows:
[0093] Perform vector conversion on the prompt information to obtain a text feature vector; input the text feature vector into the text attention denoising model structure, and perform text attention denoising on the text feature vector to obtain a second denoising result.
[0094] To perform text attention denoising, the corresponding text needs to be vectorized, that is, perform vector conversion on the prompt information to obtain a text feature vector. This step can be achieved through pre-trained large language models, such as: BERT, GPT, etc. These models will learn the semantic and syntactic information in the text and map us into a high-dimensional vector space to obtain the text feature vector;
[0095] After that, input the text feature vector into the text attention denoising model structure, and perform text attention denoising through the text attention denoising model structure to obtain a second denoising result.
[0096] Among them, in the process of the text attention denoising model structure performing text attention denoising, a certain amount of noise is added to the text feature vector to simulate the imperfection and uncertainty of the data in actual applications. The type and intensity of the noise can be adjusted according to specific situations; use the attention mechanism to denoise the text feature vector after adding noise. The attention mechanism can help the model focus on the key information in the text, ignore the noise or irrelevant information. By calculating the attention weights at each position in the text, the model can adaptively adjust the attention to different parts of the text, so as to better restore the text features disturbed by the noise.
[0097] It is possible to generate an image only using the text feature vector. However, if an image is generated only using the text feature vector, it is very likely that the finally generated image does not conform to the user's preferences. For example, if the image generation request is "Please generate an image of a cat", if only the text feature vector is used, it will randomly generate an image of a cat, and most of the image features related to cats stored in the target database are the image features of Ragdoll cats. Then, it will cause the randomly generated cat to not match the image features in the target database with a high probability. Since the image features in the target database are obtained based on the user's personal album, and the user's personal album conforms to the user's preferences, then the randomly generated cat will very likely not conform to the user's preferences. Therefore, in this embodiment, it is also necessary to combine the image features in the target database to jointly generate the target image.
[0098] Furthermore, combining the first denoising result obtained based on the target image features and the second denoising result obtained based on the prompt information to generate the target image can be specifically as follows:
[0099] Fuse the first denoising result and the second denoising result to obtain a fused result; perform joint image and text denoising on the fused result to obtain a third denoising result; generate a target image using the third denoising result.
[0100] In the image generation method disclosed in this embodiment, neither the first denoising result nor the second denoising result is a completed denoising result. That is, the first denoising result is the result obtained after performing partial image attention denoising, and the second denoising result is the result obtained after performing partial text attention denoising.
[0101] Perform image attention denoising based on the target image feature vector to obtain a first denoising result. Among them, performing denoising processing on the target image feature vector requires M steps. In this embodiment, the image attention denoising model structure outputs the first denoising result after only performing K1 steps on the target image feature vector, where K1 < M. For example, performing image attention denoising on the image feature vector requires 8 steps, and in this embodiment, only after performing the 1st and 2nd steps of image attention denoising, the image attention denoising is stopped and the first denoising result is directly obtained.
[0102] Correspondingly, perform text attention denoising based on the text feature vector to obtain a second denoising result. Performing denoising processing on the text feature vector requires N steps. In this embodiment, the text attention denoising model structure outputs the second denoising result after only performing K2 steps on the text feature vector, where K2 < N. For example, performing text attention denoising on the text feature vector requires 7 steps, and in this embodiment, only after performing the 1st, 2nd, and 3rd steps of text attention denoising, the text attention denoising is stopped and the second denoising result is directly obtained.
[0103] That is, whether it is text attention denoising or image attention denoising, only partial steps are performed. After obtaining the second denoising result and the first denoising result, fuse the second denoising result and the first denoising result to obtain a fused result. Then, perform the remaining joint image and text denoising on the fused result to obtain a third denoising result, and generate a target image using the third denoising result.
[0104] Among them, fusing the second denoising result and the first denoising result is to combine the intermediate result of text attention denoising and the intermediate result of image attention denoising. For the remaining parts, whether it is the unperformed part of text attention denoising or the unperformed part of image attention denoising, they are no longer denoised separately through text attention denoising and image attention denoising, but through joint image and text denoising to complete text attention denoising and image attention denoising simultaneously.
[0105] Among them, the process of performing joint denoising on the fusion result makes full use of the complementary information between the second denoising result corresponding to the text feature vector and the first denoising result corresponding to the target image feature to achieve a better denoising effect. That is, by using the semantic information about the image content contained in the text features, more explicit semantic constraints are provided for the image features, so that the target image features in the third denoising result are more consistent with the text description, reducing the noise and irrelevant information in the target image features; the target image features can help verify and correct the semantic information in the text features to help the model better understand the true meaning of the text, thereby denoising and correcting the text features so that the text features in the third denoising result are more consistent with the visual cues of the image features.
[0106] The goal of joint denoising is to obtain text features and image features that better meet the requirements, providing a data basis for subsequent image generation.
[0107] After simultaneously performing text attention denoising and image attention denoising through joint denoising, a third denoising result is obtained. This third denoising result is the final denoising result. Based on the final denoising result, a target image is generated using the VAE Decoder (Variational Autoencoder Decoder), that is, the VAE Decoder generates a corresponding image using the final denoising result (i.e., the latent vector). Among them, different latent vectors can lead to different images generated by the VAE Decoder.
[0108] For the image generation method disclosed in this embodiment, after obtaining an image generation request, the corresponding prompt information for the image generation request is determined. If a target label matching the prompt information is stored in the target database, the target image features corresponding to the target label in the target database are determined, the target image features are extracted, the target image features are vector-converted to obtain a target image feature vector, and it is input into the image attention denoising model structure to implement the image attention denoising process for the target image feature vector; in addition, text attention denoising needs to be performed based on the prompt information, and the image attention denoising result and the text attention denoising result are combined to generate a target image. In this embodiment, the use of the target image features and the image attention denoising model structure to perform image attention denoising ensures the accuracy of the denoising result; in addition, the image features stored in the target database are pre-stored by the user and are in line with the user's preferences. By separately performing image attention denoising and text attention denoising, it is ensured that when generating an image, both the image generation request and the information in the pre-stored target database can be used as reference factors in the image generation process, so that the finally generated image can meet the user's preferences and improve the accuracy of the generated image.
[0109] This embodiment discloses an image generation method, and its flowchart is as Figure 3 shown, including:
[0110] Step S301: Obtain an image generation request and determine the prompt information corresponding to the image generation request;
[0111] Step S302: If it is determined that there is a target label stored in the target database that matches the prompt information, determine the target image features stored in the target database corresponding to the target label;
[0112] Step S303: Perform image attention denoising based on the target image features stored in the target database to obtain a first denoising result;
[0113] Step S304: Perform text attention denoising based on the prompt information to obtain a second denoising result;
[0114] Step S305: Generate a target image based on the first denoising result and the second denoising result;
[0115] Step S306: Display the target image;
[0116] Step S307: Obtain the feedback information of the user on the displayed target image;
[0117] Step S308: Update the target label and / or target image features in the target database based on the feedback information.
[0118] After obtaining the image generation request, it is necessary to determine the prompt information corresponding to the image generation request based on the image generation request. After that, it is necessary to determine whether there is a target label stored in the target database that matches the prompt information. If so, it is necessary to determine the target image features stored in the target database corresponding to the target label and perform image attention denoising based on the target image features to obtain a first denoising result; at the same time, perform text attention denoising based on the prompt information to obtain a second denoising result; after that, generate a target image based on the first denoising result and the second denoising result.
[0119] After generating the target image, in order to respond to the image generation request, it is necessary to display the target image so that the user can determine whether the displayed target image meets the image generation request input by the user, that is, the user determines whether the displayed target image meets the user's needs.
[0120] Therefore, for the image generation method disclosed in this embodiment, it is also possible to obtain the feedback information of the user on the displayed target image. The feedback information can be: operations such as the user liking, downloading, or applying the displayed target image. At this time, it can be determined that the currently displayed target image meets the user's needs;
[0121] Alternatively, the feedback information can also be: after the target image is displayed, the user re-enters the same image generation request to request that the image generation device or model corresponding to the image generation method disclosed in this embodiment can regenerate an image based on the image generation request. Or, if the user inputs content such as "This is not the image I want" based on the displayed target image, it can be determined that the currently displayed target image does not meet the user's requirements.
[0122] When obtaining the feedback information of the user on the displayed target image, the relevant data recorded in the target database can be updated based on the feedback information. The relevant data is the target label and / or target image feature used to generate the target image, so that the updated target label and / or target image feature in the target database can conform to the user's current preference. That is, the data recorded in the target database is updated in a timely manner based on the user's feedback information to ensure that the data stored in the target database can always conform to the user's current preference. So that after the target database is updated, when using the target database to continue image generation, the generated image is generated based on the user's current preference, so as to ensure that the generated image can always follow the user's preference, thereby reducing the probability of regenerating the image.
[0123] Furthermore, updating the target label and / or target image feature in the target database based on the feedback information can be specifically:
[0124] If the feedback information indicates that the target image meets the conditions, at least part of the features of the target image are updated to the target image feature corresponding to the target label in the target database; if the feedback information indicates that the target image does not meet the conditions, at least part of the target image features corresponding to the target label in the target database are deleted.
[0125] The feedback information indicating that the target image meets the conditions means that the target image generated by the image generation system or model corresponding to the image generation method disclosed in this embodiment meets the user's requirements, that is, the generated target image not only matches the image generation request but also meets the user's preference; correspondingly, the feedback information indicating that the target image does not meet the conditions means that the target image generated by the image generation system or model corresponding to the image generation method disclosed in this embodiment does not meet the user's requirements, that is, the generated target image does not meet the user's preference.
[0126] When the feedback information indicates that the target image meets the conditions, at least part of the features of the target image can be used to update the target image feature in the target database, so that the target image feature stored in the target database can better conform to the user's preference.
[0127] Among them, the target image features corresponding to the target tags stored in the target database are image features of multiple different types in the same category. The target image generated based on the target image features and the prompt information is generated based on the image features of a certain type in the target image features and the prompt information. Then, the image features of this type must also be included in the image features of the generated target image. Therefore, the image features of this type are updated to the target image features in the target database to increase the weight of the image features of this type in the target image features corresponding to the target tag.
[0128] For example, the image features corresponding to the tag of the "cat" category in the target database may include: the image features of Ragdoll cats, the image features of orange cats, the image features of Persian cats, etc. If the obtained image generation request is "Please generate an image of a cat", randomly select the image features of a certain type of cat from the image features corresponding to the tag of the "cat" category in the target database, such as: the image features of orange cats, and generate an image of an orange cat. After displaying it, if the feedback information obtained from the user is a like, it indicates that the generated image of the orange cat meets the image generation request and the user's preference. Therefore, the image features of the orange cat in the generated image of the orange cat can be stored in the image features corresponding to the tag of the "cat" category in the target database, so that the image features of the orange cat in the image features corresponding to the tag of the "cat" category in the target database increase, thereby achieving an increase in the weight of the image features of the orange cat in the image features corresponding to the tag of the "cat" category in the target database. This is actually a process of correcting or updating the user's preference.
[0129] In addition, when the feedback information indicates that the generated target image meets the conditions, obtain the image features of the target image. The image features corresponding to the target tag in the image features of the target image can be all updated to the target database to increase the quantity of the image features of this type in the target image features corresponding to the target tag in the target database; or, not all the image features corresponding to the target tag in the target image are updated to the target database, but only a part of them are updated. For example, a first target quantity is preset. Each time image features are added to the target database, only the first target quantity of image features are added. This will make the different types of image features corresponding to each tag in the target database related to the number of times the generated target image meets or does not meet the conditions, rather than directly adding a large number of image features to the target database through a target image that meets the conditions once, which will affect the accuracy of judging the user's preference.
[0130] When the feedback information indicates that the target image does not meet the conditions, at least part of the image features corresponding to the target label in the target database can be directly deleted, so that the target image features stored in the target database can better conform to the user's preferences.
[0131] Since the image features corresponding to each label stored in the target database may be multiple types of image features, then, when the target image generated based on a certain type of image feature corresponding to a certain label does not meet the conditions, the image features of this type corresponding to the label in the target database can be directly deleted, while the other types of image features corresponding to the label in the target database are not deleted; or, only part of the image features of this type corresponding to the label in the target database are deleted, rather than directly deleting all the image features of this type.
[0132] That is: If the feedback information indicates that the target image does not meet the conditions, determine the image features of the target image; delete the features corresponding to the image features of the target image in the target image features corresponding to the target label in the target database.
[0133] For example: The image features corresponding to the label of the category "cat" in the target database may include: the image features of Ragdoll cats, the image features of orange cats, the image features of Persian cats, etc. If the obtained image generation request is "Please generate an image of a cat", randomly select a certain type of cat's image features from the image features corresponding to the label of the category "cat" in the target database, such as: the image features of Persian cats, and generate an image of a Persian cat. After displaying it, if the user's feedback information is the input of "This is not what I want", it indicates that the generated image of the Persian cat does not conform to the user's preferences. Therefore, the image features of Persian cats in the target database can be determined, and part of the image features of Persian cats in the database can be deleted. The second target quantity of Persian cat image features can be deleted, and the second target data can be 10 or 5, etc. For example: Delete 10 of the image features of Persian cats to reduce the weight of the image features of Persian cats in the image features corresponding to the label of the category "cat" in the target database, so that when the image generation request of "Please generate an image of a cat" is obtained again, the probability of generating an image of a Persian cat can be reduced. This is actually a process of correcting or updating the user's preferences.
[0134] The image generation method disclosed in this embodiment, after obtaining an image generation request and determining the prompt information corresponding to the image generation request, if it is determined that the target database stores a target tag that matches the prompt information, determines the target image features corresponding to the target tag, performs image attention denoising based on the target image features respectively to obtain a first denoising result, and performs text attention denoising based on the prompt information to obtain a second denoising result. After generating a target image based on the first denoising result and the second denoising result and displaying it, it can receive feedback information of the user on the displayed target image, and can update the target tag and / or target image features in the target database based on the feedback information. By updating the target database in a timely manner based on the feedback information of the user, it is ensured that when generating an image using the information in the target database subsequently, an image can be generated using the updated data, so as to ensure that the generated image can meet the latest preferences of the user, reduce the probability of regenerating the image, and improve the user experience.
[0135] This embodiment discloses an image generation method, and its flowchart is as Figure 4 shown, including:
[0136] Step S401, obtain a plurality of pre-stored images;
[0137] Step S402, analyze each image in the plurality of images, determine the image features and image categories of each image, and different images correspond to the same or different image categories;
[0138] Step S403, determine tags based on the image categories;
[0139] Step S404, cluster the image features of the images belonging to the same category to obtain the image features after clustering of the same image category;
[0140] Step S405, establish an association relationship between the image features after clustering of the same image category and the tags corresponding to the same image category, and form a target database storing the association relationship between the image features and tags corresponding to at least one image category;
[0141] Step S406, obtain an image generation request and determine the prompt information corresponding to the image generation request;
[0142] Step S407, if it is determined that the target database stores a target tag that matches the prompt information, determine the target image features stored in the target database corresponding to the target tag;
[0143] Step S408, perform image attention denoising based on the target image features stored in the target database to obtain a first denoising result;
[0144] Step S409: Perform text attention denoising based on the prompt information to obtain a second denoising result;
[0145] Step S410: Generate a target image based on the first denoising result and the second denoising result.
[0146] When an image generation request is obtained, in order to respond to the image generation request to generate a target image, the prompt information corresponding to the image generation request can be determined first. Then, query the target database to check if there is a target label that matches the prompt information. If so, it is necessary to determine the target image features corresponding to the target label in the target database, and perform image attention denoising based on the target image features to obtain a first denoising result; at the same time, perform text attention denoising based on the prompt information to obtain a second denoising result; finally, generate a target image based on the first denoising result and the second denoising result.
[0147] Among them, the target database is pre-obtained based on multiple images. Specifically, multiple images can be pre-stored. The multiple images can be pre-stored in the user's photo album, or selected or stored by the user from the network, etc.
[0148] After obtaining multiple images, each image is analyzed to determine the image features and image categories of each image. For example, if there is a tree in an image, its image category can be determined as "plant"; another example, if an image is in the style of Van Gogh, its image category can be determined as "Van Gogh style", etc.
[0149] Among the image categories corresponding to each image in the obtained multiple images, there are the same image categories. Cluster the image features of the images belonging to the same image category to obtain the image features after clustering for this image category. For example, if there are 30 images among the multiple images whose image category is "plant", then cluster the image features corresponding to these 30 images whose image category is "plant" to obtain the image features after clustering for the image category of "plant"; another example, if there are 5 images among the multiple images whose image category is "Van Gogh style", then cluster the image features corresponding to these 5 images whose image category is "Van Gogh style" to obtain the image features after clustering for the image category of "Van Gogh style".
[0150] Among them, when clustering the image features of images of the same image category, the image features of images of the same category can also be clustered according to different types to obtain different types of image features corresponding to the same image category. For example, for the image category of "plants", it may include multiple types, such as "flowers", "trees", etc. All the image features corresponding to the type of "flowers" in the image category of "plants" can be clustered to obtain the image features after clustering of the type of "flowers" under the image category of "plants". All the image features corresponding to the type of "trees" in the image category of "plants" can be clustered to obtain the image features after clustering of the type of "trees" under the image category of "plants", and so on.
[0151] Of course, it is also possible to cluster the image features of all the images corresponding to the same image category together to obtain an image feature set, which may include various different types of image features.
[0152] In addition, after determining the image category corresponding to each image in the multiple pre-stored images, the labels respectively corresponding to each image category can be determined to facilitate marking the image categories corresponding to different image features.
[0153] The labels respectively corresponding to each image category can specifically be vectorized labels. The vectorized label is a four-dimensional vector, which precisely locates the image category through the four-dimensional vector, and the distance between different labels can be determined through this four-dimensional vector. For example, determine the vectorized label of "cat" and the vectorized label of "dog". Based on the vectorized labels, if the distance between "cat" and "dog" is determined to be 1, then the distance between "Ragdoll cat" and "Orange cat" must be less than 1.
[0154] In addition, after determining the prompt information based on the image generation request, the prompt information can be vector-converted to obtain a text feature vector (i.e., vectorized prompt information). Based on the comparison between this text feature vector and each vectorized label in the target database, to determine which vectorized label in the target database has a closer distance to the text feature vector, then select this vectorized label as the target vectorized label to facilitate determining the target image features corresponding to the target vectorized label in the target database. By vectorizing the labels in the target database and vectorizing the prompt information, and comparing the vectorized prompt information and labels to determine the target label, the accuracy of determining the target label is improved, thereby increasing the probability that the finally generated target image meets the user's requirements.
[0155] After determining the label based on the image category and determining the image features after clustering of the same image category based on the image features of the images belonging to the same image category, an association relationship between the label of the same image category and the image features after clustering can be established and stored in the target database. Multiple association relationships between the labels corresponding to different image categories and the image features after clustering can be stored in the target database to increase the probability of querying the label and image features corresponding to a certain image category in the target database when an image generation request requests to generate an image of a certain image category, so as to increase the probability that the generated target image meets the user's requirements.
[0156] Further, in the image generation method disclosed in this embodiment, it may further include:
[0157] If it is determined that the number of image features corresponding to a certain image category stored in the target database is lower than the third target number, then the association relationship between the label corresponding to the image category and the image features after clustering will not be generated in the target database, or directly do not cluster the image features corresponding to the image category, so as to ensure that the association relationship between the label corresponding to the image category and the image features after clustering will not be stored in the target database.
[0158] Or, if for a certain image category, the association relationship between the label corresponding to the image category and the image features after clustering is recorded in the target database, and during the target image generation process, a deletion operation is performed on at least some of the image features corresponding to the image category in the target database (for example: the feedback information of the user indicates that the target image does not meet the user's requirements. At this time, at least some of the target image features used to generate the target image in the target database can be deleted), resulting in the number of image features corresponding to the image category recorded in the target database being lower than the third target number. At this time, the label corresponding to the image category can be directly deleted. When an image belonging to the image category is regenerated in response to the image generation request and the feedback information of the generated image indicates that the generated image meets the conditions, the image features corresponding to the generated image are added to the image features corresponding to the image category. If it is determined that the number of image features corresponding to the image category in the target database reaches the third target number after the addition, then at this time, the label can be re-determined based on the image category, and the association relationship between the label of the image category and the image features after clustering of the image category can be re-determined.
[0159] The image generation method disclosed in this embodiment, when generating a target image in response to an image generation request, needs to perform image attention denoising using the target image features obtained from the target database and perform text attention denoising using the prompt information corresponding to the image generation request, and generate the target image based on the two denoising results. Among them, the target database is pre-generated. When generating, it analyzes a plurality of pre-stored images to determine the image features of each image and the image category of each image, and determines the label based on the image category. After clustering the image features of the images belonging to the same image category, the image features corresponding to the image category are obtained, and the label corresponding to the image category is associated with the image features corresponding to the image category, thereby forming the target database, so that when an image generation request is obtained, the image features corresponding to the relevant image category can be queried from the target database according to the label, so as to ensure that the finally generated target image conforms to the image features of the image category in the target database, that is, to ensure that the finally generated target image conforms to the images pre-stored by the user, so as to characterize that the target image conforms to the user's preference and improve the efficiency of generating the target image that meets the user's needs.
[0160] Furthermore, if the label corresponding to the image category of the image requested to be generated in the image generation request is not stored in the target database, the target image can be generated only based on the prompt information corresponding to the image generation request.
[0161] That is: if the image category of the image to be generated corresponding to the image generation request is category a, and the target database does not store the association relationship between the label and the image features of the images of category a, then the label corresponding to the images of category a cannot be queried from the target database, and image attention denoising based on the image features in the target database is no longer considered, but the image is generated only based on the prompt information corresponding to the image generation request.
[0162] At this time, the prompt information can be input into the image attention denoising model structure and the text attention denoising model structure at the same time, so as to perform image attention denoising and text attention denoising respectively using the prompt information, and then generate the target image based on the two obtained denoising results; or, directly input the prompt information into the image and text joint denoising model structure to directly obtain the final denoising result, and generate the target image based on the final denoising result; or, it can also be: only input the prompt information into the text attention denoising model structure, only perform text attention denoising to obtain the denoising result, and generate the target image based on this.
[0163] Based on this, the schematic diagram of the complete image generation method disclosed in this embodiment can be as Figure 5 shown, including:
[0164] Obtain a plurality of pre-stored images;
[0165] Obtain an image category based on an image;
[0166] Obtain a vectorized label based on the image category;
[0167] Obtain image features based on an image;
[0168] Associate the vectorized labels and image features of the same image category to construct a target database;
[0169] Obtain an image generation request;
[0170] Obtain a prompt corresponding to the image generation request;
[0171] Vectorize the prompt to obtain a text feature vector;
[0172] Match the text feature vector with the vectorized label;
[0173] If a match is found, perform image attention denoising using the target image features corresponding to the vectorized label to obtain a first denoising result;
[0174] Perform text attention denoising using the text feature vector to obtain a second denoising result;
[0175] Obtain a target image based on the first denoising result and the second denoising result;
[0176] If no match is found, perform denoising based on the text feature vector to obtain a denoising result;
[0177] Obtain a target image based on the denoising result;
[0178] Obtain feedback information on the target image;
[0179] If the feedback information indicates that the condition is satisfied, update the target image features in the target database;
[0180] If the feedback information indicates that the condition is not satisfied, delete at least part of the target image features in the target database.
[0181] This embodiment discloses an image generation device, the structural schematic diagram of which is as Figure 6 shown, including:
[0182] An application module 61, a storage module 62, a determination module 63, an image attention denoising model structure 64, a text attention denoising model structure 65, and a generation module 66.
[0183] Among them, the application module 61 is used to receive an image generation request and determine the prompt information corresponding to the image generation request;
[0184] The storage module 62 is used to store the target database, and the association relationship between the image features and the labels is recorded in the target database;
[0185] The determination module 63 is used to determine that the target label matching the prompt information is recorded in the target database stored by the storage module, and determine the target image features corresponding to the target label recorded in the target database;
[0186] The image attention denoising model structure 64 is used to perform image attention denoising based on the target image features recorded in the target database to obtain a first denoising result;
[0187] The text attention denoising model structure 65 is used to perform text attention denoising based on the prompt information to obtain a second denoising result;
[0188] The generation module 66 is used to generate a target image based on the first denoising result and the second denoising result.
[0189] Furthermore, the image attention denoising model structure is used to:
[0190] Extract the target image features stored in the target database; perform vector conversion on the target image features to obtain a target image feature vector; input the target image feature vector into the image attention denoising model structure to perform image attention denoising on the target image feature vector to obtain a first denoising result.
[0191] Furthermore, the text attention denoising model structure is used to:
[0192] Perform vector conversion on the prompt information to obtain a text feature vector; input the text feature vector into the text attention denoising model structure to perform text attention denoising on the text feature vector to obtain a second denoising result.
[0193] Furthermore, the generation module is used to:
[0194] Fuse the first denoising result and the second denoising result to obtain a fusion result; perform image and text joint denoising on the fusion result to obtain a third denoising result; use the third denoising result to generate a target image.
[0195] Furthermore, the image generation device disclosed in this embodiment may further include:
[0196] A feedback module, which is used to display the target image; obtain the feedback information of the user on the displayed target image; update the target label and / or target image features in the target database based on the feedback information.
[0197] Furthermore, the feedback module is used to:
[0198] If the feedback information indicates that the target image meets the conditions, update at least part of the image features of the target image to the target image features corresponding to the target label in the target database; if the feedback information indicates that the target image does not meet the conditions, delete at least part of the target image features corresponding to the target label in the target database.
[0199] Further, the feedback module is configured to:
[0200] If the feedback information indicates that the target image does not meet the conditions, determine the image features of the target image; delete the features corresponding to the image features of the target image from the target image features corresponding to the target label in the target database.
[0201] Further, the image generation device disclosed in this embodiment may further include:
[0202] A construction module, configured to obtain a plurality of pre-stored images; analyze each image in the plurality of images to determine the image features and image categories of each image, where different images may correspond to the same or different image categories; determine labels based on the image categories; cluster the image features of the images belonging to the same image category to obtain the image features after clustering for the same image category; establish an association relationship between the image features after clustering for the same image category and the labels corresponding to the same image category, and form a target database storing the association relationship between the image features and labels corresponding to at least one image category.
[0203] Further, the generation module is further configured to:
[0204] If it is determined that the target database does not store a target label matching the prompt information, generate a target image according to the prompt information.
[0205] It should be noted that the image generation device disclosed in this embodiment may specifically be an image generation model.
[0206] The image generation device disclosed in this embodiment is implemented based on the image generation method disclosed in the above embodiment, and details are not described herein again.
[0207] The image generation device disclosed in this embodiment, when obtaining an image generation request, determines the prompt information corresponding to the image generation request, and searches in the target database to find out whether there is a target tag matching the prompt information. If so, it determines the target image features corresponding to the target tag stored in the target database, and performs image attention denoising based on the target image features to obtain a first denoising result. In addition, it performs text attention denoising based on the prompt information to obtain a second denoising result. Finally, it generates a target image based on the first denoising result and the second denoising result. In this solution, when there is an image generation request, on the one hand, it uses the target image features stored in the target database to perform image attention denoising, and on the other hand, it uses the prompt information corresponding to the image generation request to perform text attention denoising. Among them, the image features stored in the target database are pre-stored by the user and conform to the user's preferences. By performing image attention denoising and text attention denoising respectively, it is ensured that when generating an image, both the image generation request and the information in the pre-stored target database can be used as reference factors in the image generation process, so that the finally generated image can conform to the user's preferences and improve the accuracy of the generated image.
[0208] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution in this embodiment. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0209] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the present application, software program implementation is a better embodiment in more cases. Based on such an understanding, the technical solution of the present application, in essence, or the part that makes a contribution to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0210] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0211] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).
Claims
1. An image generation method, comprising: Obtaining an image generation request; Determining prompt information corresponding to the image generation request; If it is determined that a target label matching the prompt information is stored in a target database, determining target image features stored in the target database corresponding to the target label; Performing image attention denoising based on the target image features stored in the target database to obtain a first denoising result; Performing text attention denoising based on the prompt information to obtain a second denoising result; Generating a target image based on the first denoising result and the second denoising result.
2. The method according to claim 1, wherein the performing image attention denoising based on the target image features stored in the target database to obtain a first denoising result comprises: Extracting the target image features stored in the target database; Performing vector conversion on the target image features to obtain a target image feature vector; Inputting the target image feature vector into an image attention denoising model structure to perform image attention denoising on the target image feature vector to obtain a first denoising result.
3. The method according to claim 1, wherein the performing text attention denoising based on the prompt information to obtain a second denoising result comprises: Performing vector conversion on the prompt information to obtain a text feature vector; Inputting the text feature vector into a text attention denoising model structure to perform text attention denoising on the text feature vector to obtain a second denoising result.
4. The method according to claim 1, wherein the generating a target image based on the first denoising result and the second denoising result comprises: Fusing the first denoising result and the second denoising result to obtain a fusion result; Performing image and text joint denoising on the fusion result to obtain a third denoising result; Generating the target image using the third denoising result.
5. The method according to claim 1, further comprising: Displaying the target image; Obtaining feedback information of a user on the displayed target image; Updating the target label and / or the target image features in the target database based on the feedback information.
6. The method according to claim 5, wherein the updating the target label and / or the target image features in the target database based on the feedback information comprises: If the feedback information indicates that the target image meets the condition, updating at least part of the image features of the target image to the target image features corresponding to the target label in the target database; If the feedback information indicates that the target image does not meet the condition, deleting at least part of the target image features corresponding to the target label in the target database.
7. The method according to claim 6, wherein the if the feedback information indicates that the target image does not meet the condition, deleting at least part of the target image features corresponding to the target label in the target database comprises: If the feedback information indicates that the target image does not meet the condition, determining the image features of the target image; Delete the features corresponding to the image features of the target image in the target image features corresponding to the target label in the target database.
8. The method according to claim 1, further comprising: Obtain a plurality of pre-stored images; Analyze each of the plurality of images to determine the image features and image categories of each image, where different images correspond to the same or different image categories; Determine labels based on the image categories; Cluster the image features of the images belonging to the same image category to obtain the image features after clustering of the same image category; Establish an association relationship between the image features after clustering of the same image category and the labels corresponding to the same image category, and form a target database storing the association relationship between the image features and labels corresponding to at least one image category.
9. The method according to claim 1, further comprising: If it is determined that the target database does not store a target label matching the prompt information, generate the target image according to the prompt information.
10. An image generation device, comprising: An application module, configured to receive an image generation request and determine the prompt information corresponding to the image generation request; A storage module, configured to store a target database, where the target database records the association relationship between image features and labels; A determination module, configured to, when determining that the target database stored in the storage module records a target label matching the prompt information, determine the target image features recorded in the target database corresponding to the target label; An image attention denoising model structure, configured to perform image attention denoising based on the target image features recorded in the target database to obtain a first denoising result; A text attention denoising model structure, configured to perform text attention denoising based on the prompt information to obtain a second denoising result; A generation module, configured to generate a target image based on the first denoising result and the second denoising result.