Image generation method and image generation model training method and device
By extracting the characteristics of the reference image and prompt text and iteratively updating the noisy image, the problem of difficulty in generating target images similar to the reference image in the prior art is solved, high-quality image generation is achieved, and semantic understanding ability is maintained.
Patent Information
- Application Number
- CN202410309735.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-03-18
AI Technical Summary
The prior art is difficult to generate a target image based on the reference image while maintaining semantic understanding ability, so that the objects in the target image are similar to those in the reference image. Especially in literary-picture scenes, it is impossible to take into account this requirement and maintain semantic understanding ability.
A training method for image generation and image generation model is provided, by extracting the features of reference images and prompt text, and iteratively updates the noise image to generate a target image. The method includes a feature extraction network, a literary graph model and an image decoding network. By fusing the reference image features and text features, the noise image is updated to ensure that the target image is similar to the reference image and match the prompt text.
While maintaining semantic understanding capabilities, the objects in the generated target image are similar to those in the reference image, improving the authenticity of user experience and image generation.
Smart Images

Figure CN118015144B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, specifically to the fields of deep learning, image processing, natural language processing and computer vision, and can be applied to scenarios such as artificial intelligence generated content. Background Art
[0002] With the development of computer technology and network technology, deep learning models are being used more and more widely, and deep learning models have made breakthrough progress in various fields. Among them, AI generated content (AIGC) is an important direction of deep learning, and the key points to focus on include: how to make the generated content more in line with user needs. Summary of the invention
[0003] The present disclosure aims to provide an image generation method and an image generation model training method, apparatus, device, medium and program product that are beneficial to improving the authenticity of generated images and improving user experience.
[0004] According to a first aspect of the present disclosure, an image generation method is provided, comprising: extracting features of a reference image to obtain reference image features; the reference image includes a first target object; extracting features of a prompt text to obtain text features; using a random noise image as an initial image of the noise image, and iteratively updating the noise image based on text features and reference image features to generate a target image; the target image includes a second target object similar to the first target object, and the target image matches the prompt text; wherein, iteratively updating the noise image based on the text features and the reference image features comprises, in at least one updating process: fusing the reference image features and the text features with the current noise image respectively to obtain fused features; and updating the current noise image based on the fused features, wherein the target image is obtained by decoding the final noise image.
[0005] According to a second aspect of the present disclosure, a training method for an image generation model is provided, wherein the image generation model includes a feature extraction network and a Vincent graph model; the Vincent graph model includes a text understanding network, an image information creation network and an image decoding network; the method includes: using a feature extraction network to extract features of a reference image to obtain reference image features; the reference image includes a first target object; using a text understanding network to extract features of a sample text in first sample data to obtain text features; using an image information creation network, taking a random noise image as an initial image of the noise image, iteratively updating the noise image according to text features and reference image features to obtain an updated noise image; using an image decoding network to decode the updated noise image to obtain a target image, the target image includes a second target object similar to the first target object, and the target image matches the sample text; and training the Vincent graph model according to the target image and the first sample image in the first sample data.
[0006] According to a third aspect of the present disclosure, an image generating device is provided, comprising: an image feature extraction module, used to extract features of a reference image to obtain reference image features; the reference image comprises a first target object; a text feature extraction module, used to extract features of a prompt text to obtain text features; a noise image updating module, used to use a random noise image as an initial image of the noise image, and iteratively update the noise image according to text features and reference image features to generate a target image; the target image comprises a second target object similar to the first target object, and the target image matches the prompt text; wherein the noise image updating module comprises: a fusion submodule, used to fuse the reference image features and the text features with the current noise image respectively in at least one updating process to obtain fused features; and an updating submodule, used to update the current noise image based on the fused features, wherein the target image is obtained by decoding the final noise image.
[0007] According to a fourth aspect of the present disclosure, a training device for an image generation model is provided, wherein the image generation model includes a feature extraction network and a Vincent graph model; the Vincent graph model includes a text understanding network, an image information creation network and an image decoding network; the device includes: an image feature extraction module, which is used to extract features of a reference image using a feature extraction network to obtain reference image features; the reference image includes a first target object; a text feature extraction module, which is used to extract features of a sample text in first sample data using a text understanding network to obtain text features; a noise image update module, which is used to use an image information creation network, take a random noise image as an initial image of the noise image, iteratively update the noise image according to text features and reference image features, and obtain an updated noise image; a decoding module, which is used to decode the updated noise image using an image decoding network to obtain a target image, the target image includes a second target object similar to the first target object, and the target image matches the sample text; and a first training module, which is used to train the Vincent graph model according to the target image and the first sample image in the first sample data.
[0008] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image generation method or the image generation model training method provided by the present disclosure.
[0009] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the image generation method or the image generation model training method provided by the present disclosure.
[0010] According to a seventh aspect of the present disclosure, a computer program product is provided, comprising a computer program / instruction, wherein the computer program / instruction is stored on at least one of a readable storage medium and an electronic device, and wherein the computer program / instruction, when executed by a processor, implements the image generation method or the image generation model training method provided in the present disclosure.
[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0013] Figure 1It is a schematic diagram of an application scenario of the image generation method and the training method and device of the image generation model according to the embodiment of the present disclosure;
[0014] Figure 2 is a flowchart of an image generating method according to an embodiment of the present disclosure;
[0015] Figure 3 is a schematic diagram of the principle of the image generation method according to an embodiment of the present disclosure;
[0016] Figure 4 is a schematic diagram of the principle of the fusion feature according to the first embodiment of the present disclosure;
[0017] Figure 5 is a schematic diagram of the principle of the fusion feature according to the second embodiment of the present disclosure;
[0018] Figure 6 is a flowchart of a method for training an image generation model according to an embodiment of the present disclosure;
[0019] Figure 7 is a structural block diagram of an image generating device according to an embodiment of the present disclosure;
[0020] Figure 8 is a structural block diagram of a training device for an image generation model according to an embodiment of the present disclosure; and
[0021] Fig. 9 It is a block diagram of an electronic device used to implement the image generation method or the image generation model training method of the embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0023] Under the general direction of AIGC, the technologies of generating images from text (text2img, referred to as text-generated images), generating images from images (img2img), generating videos from images (img2video), and generating videos from text (text2video) are becoming more and more popular. This has raised more requirements, such as the need to generate images with similar composition and color to the reference image. However, in the text-generated image scenario, it is usually impossible to take into account this requirement and the need to maintain the original semantic understanding ability of the text-generated image network at the same time.
[0024] In order to solve the above problems, the present disclosure provides an image generation method and an image generation model training method, apparatus, device, medium and program product, so as to generate a target image based on a reference image while maintaining the semantic understanding ability, so that the object in the target image is similar to the object in the reference image. Figure 1 The application scenarios of the method and device provided by the present disclosure are described.
[0025] Figure 1 It is a schematic diagram of application scenarios of the image generation method and the training method and device of the image generation model according to the embodiments of the present disclosure.
[0026] like Figure 1 As shown, the application scenario 100 may include an electronic device 110. The electronic device 110 may be any electronic device with processing capability, such as a smart phone, a tablet computer, a portable computer, a desktop computer, or a server.
[0027] In one embodiment, the electronic device 110 may generate a target image 103 according to the prompt text 101 and the reference image 102 provided by the user. The target image 103 may be similar to the reference image 102 in terms of picture composition and color, and the target image 103 may match the prompt text 101. Specifically, the target object in the target image 103 may be similar to the target object in the reference image 102. The target object may be any object such as a person, a cat, a dog, etc., and the present disclosure does not limit this.
[0028] Exemplarily, the prompt text 101 may be “a boy wearing a white T-shirt”, the reference image 102 may be an image of boy a wearing a brown windbreaker, and the generated target image 103 may include, for example, a boy wearing a T-shirt and having a similar appearance to boy a.
[0029] In one embodiment, the electronic device 110 may be installed with various client applications, such as image processing applications, instant messaging applications, web browsing applications, intelligent generation applications, etc. The electronic device 110 may use the installed intelligent generation application to process the prompt text 101 and the reference image 102 provided by the user, thereby generating a target image 103.
[0030] In one embodiment, the electronic device 110 may use a pre-generated image generation model 104 to process the prompt text 101 and the reference image 102, thereby generating a target image 103. The image generation model 104 here may be provided with a feature extraction network for extracting image features of the reference image 102, and may also be provided with an improved text graph model to generate the target image 103 based on the features extracted by the feature extraction network and the prompt text 101.
[0031] In one embodiment, the application scenario 100 may further include a server 120. The server 120 may, for example, train the image generation model 104 based on data in a public data set, and directly or indirectly send the image generation model 104 to the electronic device 110. For example, the server 120 may be a background management server that provides support for the operation of the intelligent generation application installed in the electronic device 110, or may be any server that is in communication with the background management server, which is not limited in the present disclosure.
[0032] It should be noted that the image generation method provided by the present disclosure can be executed by the electronic device 110. Accordingly, the image generation device provided by the present disclosure can be set in the electronic device 110. The image generation model training method provided by the present disclosure can be executed by the server 120. Accordingly, the image generation model training device provided by the present disclosure can be set in the server 120.
[0033] It should be understood that Figure 1 The number and type of the electronic devices 110 and the server 120 in the embodiment are only for illustration. According to the implementation requirements, any number and type of the electronic devices 110 and the server 120 may be provided.
[0034] The following will be combined Figure 2 to Figure 5 The image generation method provided by the present disclosure is described in detail.
[0035] Figure 2 It is a flowchart of an image generating method according to an embodiment of the present disclosure.
[0036] like Figure 2 As shown, the image generating method 200 of this embodiment may include operations S210 to S230, wherein, in the iterative updating process of implementing operation S230, part of the iterative updating process may be implemented through operations S231 to S232.
[0037] In operation S210, features of a reference image are extracted to obtain reference image features.
[0038] For example, the reference image may include a first target object, which may be a person, a cat, a dog, etc., and the reference image may be any image in a public data set, or an image provided by a user with permission, which is not limited in the present disclosure.
[0039] For example, any feature extraction network can be used to process the reference image to extract the reference image features. The feature extraction network may include ResNet, a Transformer-based backbone network, DenseNet or U-Net, etc. The feature extraction network may include, for example, an encoder and / or a decoder, which is not limited in the present disclosure.
[0040] In operation S220, features of the prompt text are extracted to obtain text features.
[0041] For example, the prompt text may be text input into a text model in the text graph technology, and the text content may be determined in response to user input, for example, the prompt text may be "a boy, wearing a white T-shirt" and the like.
[0042] For example, a word segmenter may be used to segment the prompt text, and each word obtained by the segmentation may be converted into a token to obtain a token sequence. Each token in the token sequence may then be converted into an embedding vector (e.g., a 768-dimensional vector). Finally, the multiple embedding vectors obtained by the conversion may be concatenated into a matrix to obtain text features. For example, a text transformer (e.g., a Text transformer) may also be used to process the concatenated matrix, and the processing result may be used as a text feature.
[0043] For example, the text understanding network included in the text-generated graph model can be used to extract text features. The text understanding network is similar to the text understanding network included in the Chinese-generated graph model in the related art, and will not be described in detail here.
[0044] In operation S230 , the random noise image is used as an initial image of the noise image, and the noise image is iteratively updated according to the text feature and the reference image feature to generate a target image.
[0045] According to an embodiment of the present disclosure, the implementation process of operation S230 can be understood as a Wenshengtu process. Among them, the Wenshengtu process refers to a process in which, given a prompt text, an AI image that matches the text can be fed back. The operation S230 can be understood as a reverse diffusion process, that is, starting from a noisy, undisputed image (random noise image), reverse diffusion is performed to restore the image to obtain an image including the target object. The opposite of the reverse diffusion process is the forward diffusion process, which is the process of adding noise to an image and gradually transforming the image into a frame of featureless noise image.
[0046] For example, a random noise image can be understood as a random tensor. The random noise image can be generated by a random number generator by setting a seed in the random number generator according to actual needs. For example, the random noise image can obey a Gaussian distribution.
[0047] For reverse diffusion, a neural network model can be pre-trained to predict the added noise. In steady-state diffusion, this neural network model is called a noise predictor. The noise predictor is trained as follows: a training image is selected (for example, an image with a cat); a random noise image is then generated and added to the training image; the noise image is added for a certain number of steps to corrupt the training image; and finally the noise predictor is taught to give feedback on how much noise was added. This is done by adjusting the weights of the noise predictor and showing it the correct answer. After training, a noise predictor is obtained that is able to estimate the noise added to the image.
[0048] Among them, steady-state diffusion is a latent diffusion model (LatentDiffusionModel), which first compresses the image into a latent space, which is 48 times smaller than the pixel space. Steady-state diffusion can be achieved through a variational autoencoder (VAE). The variational autoencoder consists of two parts, an encoder and a decoder. The encoder compresses an image into a lower-dimensional representation in the latent space, and the decoder recovers the image from the latent space. During the training process, steady-state diffusion does not generate noisy images, but generates random tensors (latent noise) in the latent space. The process of adding noise is not the process of destroying the image with noise, but the process of destroying the expression of the image in the latent space with latent noise.
[0049] In operation S230, the text features and the reference image features can be used as prompt information to guide the noise predictor to predict noise. After the predicted noise is subtracted from the random noise image, a desired image (eg, an image of a cat or an image of a dog) can be obtained through conversion.
[0050] In one embodiment, the overall process of operation S230 may be as follows:
[0051] In the first step, stable diffusion generates random tensors in the latent space as the initial images of the noise images.
[0052] In the second step, the noise predictor takes the latent noisy image, text features, and reference image features as input and predicts the noise in the latent space (a 4*64*64 tensor).
[0053] In the third step, the potential noise is subtracted from the current noisy image to obtain a new noisy image.
[0054] Repeat the second step and the third step to perform a certain number of samplings, for example, 20 times. The number of samplings can be set according to actual needs, and the present disclosure does not limit this.
[0055] In the fourth step, the decoder in the variational autoencoder converts the final noisy image back to the pixel space to obtain the image after stable diffusion, that is, the target image.
[0056] In one embodiment, when a certain number of samplings are performed, that is, in the process of iteratively updating the noise image, reference image features may be introduced in part or all of the iterative processes to guide the noise predictor to predict the noise.
[0057] Specifically, at least one updating process may include operations S231 to S232.
[0058] In operation S231, the reference image features and the text features are fused with the current noise image to obtain fused features.
[0059] The operation S231 may be, for example, inputting the reference image feature, the text feature and the current noise image into a noise predictor, and the noise predictor predicts a potential noise, which may be understood as a fusion feature.
[0060] In operation S231, the reference image features and text features are fused with the current noise image respectively, so that the noise prediction process can equally consider the reference image features and text features. For example, the reference image features and the current noise image can be fused synchronously, and the text features and the current noise image can be fused synchronously, and finally the fused features are obtained based on the fusion results of the two parts. Alternatively, the reference image features and the text features can be first concat()ed, that is, spliced in the channel dimension, and then fused with the current noise image.
[0061] In operation S232, the current noise image is updated based on the fusion feature. For example, this embodiment can subtract the fusion feature from the current noise image to obtain an updated noise image, thereby completing the update of the current noise image.
[0062] The disclosed embodiment introduces a reference image in the image generation process and updates the noise image based on the reference image and the prompt text, so that the target image obtained by decoding the noise image can not only match the prompt text, but also the generation process of the target image can refer to the reference image, so that the target object in the generated target image is similar to the target object in the reference image. Furthermore, by fusing the reference image features and the text features with the current noise image in at least one iteration, the noise prediction process can consider the influence of the prompt text and the reference image on the predicted noise, which is conducive to improving the accuracy of the predicted noise, and thus helps to improve the user's satisfaction with the final generated target image and improve the image generation effect.
[0063] In one embodiment, reference image features may be introduced in some iterative update processes to improve image generation efficiency. Reference image features may also be introduced in each iterative update process so that each noise prediction process considers the reference image features, which is beneficial to improving the similarity between the restored target image and the reference image.
[0064] For example, the reference image feature can be introduced in the iterative update process after a predetermined number of iterations in multiple iterations, that is, the aforementioned at least one update process includes an update process after a predetermined number of iterations in the process of iteratively updating the noise image. For example, if the total number of updates is set to 20, the reference image feature can be introduced in each update process starting from the 6th update process. In this way, the image generation efficiency can be improved on the basis of ensuring the similarity between the restored target image and the reference image. This is because the noise is large in the first few update processes, and even if the reference image feature is introduced, the guidance provided by the reference image feature is limited.
[0065] Figure 3 It is a schematic diagram of the principle of the image generating method according to an embodiment of the present disclosure.
[0066] In one embodiment, the feature extraction network for extracting reference image features may include, for example, multiple sampling layers, each sampling layer being used to sample the reference image at different resolutions. The aforementioned noise predictor may, for example, create a network for image information in a text graph model, and the noise predictor may be a multi-layer network structure for performing feature processing (e.g., feature fusion) step by step.
[0067] In one embodiment, the number of sampling layers in the feature extraction network may be equal to the number of network levels in the noise predictor. The multiple reference image features extracted by the multiple sampling layers of the feature extraction network may guide the noise predictor to perform feature processing layer by layer, or the reference image features extracted by some sampling layers may be used to guide the network of the corresponding level in the noise predictor to perform feature processing.
[0068] Taking the multiple reference image features extracted by multiple sampling layers of the feature extraction network to guide the noise predictor to perform feature processing layer by layer as an example, the number of multiple sampling layers in the feature extraction network and the number of network layers in the noise predictor (i.e., the image information creation network) are both set to N. In this embodiment, in a single update process in which reference image features need to be introduced, the i-th network layer in the image information creation network can process the features output by the (i-1)-th network layer (i.e., the features of the input i-th network layer obtained by processing the current noise image, i.e., the i-th image features), the i-th reference image features and text features extracted by the i-th sampling layer in the feature extraction network. Specifically, the i-th network layer can fuse the i-th reference image features and text features with the i-th image features, respectively. For the first network layer in the image information creation network, the input features are the noise images that need to be updated in the current iteration. It can be understood that in the layer-by-layer guidance technical solution, the value range of i is all integers greater than or equal to 1 and less than or equal to N. In the technical solution that does not require layer-by-layer guidance, the value range of i is some integers among the integers greater than or equal to 1 and less than or equal to N.
[0069] like Figure 3 As shown, in the embodiment 300, the feature extraction network 310 may be a U-Net including an encoder and a decoder, and the feature extraction network 310 may include four downsampling layers (as encoders) and four upsampling layers (as decoders). The reference image 301 is input into the feature extraction network 310, and 8 reference image features may be extracted layer by layer. The image information creation network 320 may correspondingly include four downsampling layers and four upsampling layers, and the input of the i-th sampling layer in the image information creation network 320 includes: the output feature of the previous sampling layer (the i-th image feature), the i-th reference image feature output by the i-th sampling layer in the feature extraction network 310, and the text feature 303 obtained by extracting the prompt text 302. The i-th sampling layer in the image information creation network 320 may fuse the i-th reference image feature and the text feature with the i-th image feature respectively, and the obtained fused feature may be used as the (i+1)-th image feature. It can be understood that the first image feature is the noise image 331 that needs to be updated in this iteration. The feature finally output by the image information creation network 320 may be a fusion feature used as an update basis, specifically, the predicted noise. The predicted noise is subtracted from the noise image 331 to obtain an updated noise image 332. At this point, an iterative update can be completed.
[0070] This embodiment predicts noise in a layer-by-layer guiding manner, which can improve the correlation between the predicted noise and the reference image, and can enable the reference image to guide the image generation process in terms of both high-level semantic features and low-level visual features, so that the generated target image can better retain the characteristics of the reference image in composition and color (such as the characteristics of the target object).
[0071] In one embodiment, during the iterative process of introducing reference image features, a cross attention mechanism may be used to fuse the reference image features and text features with the current noise image, respectively, so as to improve the image fusion effect.
[0072] Figure 4 It is a schematic diagram of the principle of the fusion feature according to the first embodiment of the present disclosure.
[0073] According to an embodiment of the present disclosure, when performing feature fusion, for example, a cross-attention mechanism may be first used to fuse the reference image feature with the current noise image to obtain a first sub-fusion feature. At the same time, a cross-attention mechanism may be used to fuse the text feature with the current noise image in any order to obtain a second sub-fusion feature. Finally, the final fusion feature is obtained based on the first sub-fusion feature and the second sub-fusion feature.
[0074] When the cross attention mechanism is adopted, for example, the current noise image can be used as a query, and the reference image feature and the text feature can be used as the key and value to perform a cross attention operation. In this embodiment, a concat() operation can be performed on the first sub-fusion feature and the second sub-fusion feature to obtain the final fusion feature.
[0075] It can be understood that, when the aforementioned feature extraction network and noise predictor are a single-layer structure, the fused feature can be used as the noise predicted in a single iteration.
[0076] like Figure 4 As shown, in the embodiment 400 in which the feature extraction network 410 and the noise predictor 420 are multi-layer structures and the reference image features extracted by the feature extraction network 410 guide the noise generation process layer by layer, the reference image features extracted from the reference image 401 by the i-th sampling layer in the feature extraction network 410 and the text features 403 extracted from the prompt text 402 can be input into the i-th layer network in the noise predictor 420. The i-th layer network in the noise predictor 420 first uses a cross attention mechanism to fuse the input reference image features and the current noise image, and fuse the input text features and the current noise image. Then, based on the obtained first sub-fusion features and the second sub-fusion features, a fusion feature is obtained.
[0077] It can be understood that, in the feature fusion process, the current noise image of the first layer network of the input noise predictor is the noise image 431 that needs to be updated in the current update process. Afterwards, the fused features output by the i-th layer network in the noise predictor can be used as the current noise image (i.e., the (i+1)th image feature) of the (i+1)th layer in the input noise predictor. Similarly, the fused features output by the last layer in the noise predictor are finally used as the predicted noise. By subtracting the predicted noise from the noise image 431, the updated noise image 432 can be obtained.
[0078] It can be understood that, when the reference image features extracted by the feature extraction network guide the noise generation process through layer selection, the unselected network layers in the noise predictor can only use the cross-attention mechanism to fuse the text features and the image features input thereto.
[0079] In one embodiment, when a fusion feature is obtained based on the first sub-fusion feature and the second sub-fusion feature, for example, a weight can be added to the first sub-fusion feature obtained by fusion based on the reference image feature, and as the number of iterations increases, the weight added to the first sub-fusion feature is increased accordingly. Specifically, the first sub-fusion feature can be weighted by a predetermined weight to obtain a weighted feature. The value of the predetermined weight is positively correlated with the number of iterations of the current iteration. The weighted feature and the second sub-fusion feature are then spliced to obtain a fusion feature. The process of splicing features can be understood as a process of performing a concat() operation. In this way, as the number of iterations increases, the guiding role of the reference image feature in the noise prediction process can be gradually increased, which is conducive to improving the satisfaction of the target image finally generated. This is because as the number of iterations increases, the content expressed by the noise image gradually increases. Setting a larger weight for the reference image feature can make the updated noise image more similar to the reference image.
[0080] Figure 5 It is a schematic diagram of the principle of the fusion feature according to the second embodiment of the present disclosure.
[0081] According to an embodiment of the present disclosure, when performing feature fusion, for example, an overall guiding feature can be first obtained based on the reference image feature and the text feature, and then a cross-attention mechanism can be used to fuse the overall guiding feature and the current noise image to obtain a fused feature.
[0082] Among them, when the cross attention mechanism is adopted, for example, the current noise image can be used as a query, and the overall guiding feature can be used as a key and value to perform a cross attention operation. This embodiment can obtain an overall guiding feature by first converting the reference image feature to the feature space where the text feature is located, and then splicing the converted feature with the text feature. The process of splicing features can be understood as the process of performing a concat() operation.
[0083] It can be understood that, when the aforementioned feature extraction network and noise predictor are a single-layer structure, the fused features obtained by fusing the concatenated features and the current noise image can be used as the noise predicted in a single iteration.
[0084] like Figure 5 As shown, in the embodiment 500 in which the feature extraction network 510 and the noise predictor 520 are multi-layer structures and the reference image features extracted by the feature extraction network 510 guide the noise generation process layer by layer, the reference image features extracted from the reference image 501 by the i-th sampling layer in the feature extraction network 510 and the text features 503 extracted from the prompt text 502 can be input into the i-th network of the noise predictor 520. The i-th network in the noise predictor 520 first converts the input reference image features, and splices the converted features and the text features to obtain spliced features. Finally, the spliced features and the current noise image are fused using a cross attention mechanism to obtain fused features.
[0085] It can be understood that, in the feature fusion process, the current noise image of the first layer network of the input noise predictor is the noise image 531 that needs to be updated in the current update process. Afterwards, the fused features output by the i-th layer network in the noise predictor can be used as the current noise image (i.e., the (i+1)th image feature) of the (i+1)th layer in the input noise predictor. Similarly, the fused features output by the last layer in the noise predictor are finally used as the predicted noise. By subtracting the predicted noise from the noise image 531, the updated noise image 532 can be obtained.
[0086] It can be understood that, when the reference image features extracted by the feature extraction network guide the noise generation process through layer selection, the unselected network layers in the noise predictor can only use the cross-attention mechanism to fuse the text features and the image features input thereto.
[0087] In one embodiment, when an overall guiding feature is obtained based on the reference image feature and the text feature, for example, a weight can be added to the reference image feature first, and as the number of iterations increases, the weight added to the reference image feature is increased accordingly. Specifically, the reference image feature can be weighted using a predetermined weight to obtain a weighted feature. The weighted feature is then converted to the feature space where the text feature is located to obtain a converted feature. Finally, the converted feature and the text feature are spliced to obtain a spliced feature as an overall guiding feature. In this way, as the number of iterations increases, the guiding effect of the reference image feature on the noise prediction process can be gradually increased, which is conducive to improving the satisfaction of the target image finally generated. This is because as the number of iterations increases, the content expressed by the noise image gradually increases. Setting a larger weight for the reference image feature can make the updated noise image more similar to the reference image.
[0088] In one embodiment, in order to facilitate the implementation of the image generation method described above, the embodiment of the present disclosure also provides an improved image generation model. The image generation model includes a text understanding component, a feature extraction network and an image generator. Among them, the text understanding component is used to convert the prompt text into a digital representation, and can specifically include the aforementioned word segmenter, an embedding network for converting a token sequence into an embedding vector, and the aforementioned text converter. The feature extraction network can be, for example, the aforementioned U-Net including an encoder and a decoder. The image generator can include an image information creation network and an image decoder. Among them, the image information creation network includes a network layer for introducing text features extracted by the text understanding component and a network layer for introducing reference image features extracted by the feature extraction network. The network layer for introducing text features and the network layer for introducing reference image features in the image information creation network can be, for example, an attention layer, which is used to fuse the introduced features and the noise image using a cross-attention mechanism. The image information creation network can also include a calculation layer for performing a concat() operation on the fused features obtained by fusing the text features and the noise image and the fused features obtained by fusing the reference image features and the noise image to obtain predicted noise. Alternatively, the image information creation network may include a network layer for converting reference image features and concatenating reference image features and text features, and may also include an attention layer for performing cross-attention operations on the aforementioned overall guide features and noise image to obtain predicted noise. The image information creation network may also include a network layer for updating the noise image according to the predicted noise. By running the image information creation network multiple times, a final noise image can be obtained. The image decoder is used to decode the final noise image to generate a target image.
[0089] In order to make the improved image generation model generate images better, the embodiment of the present disclosure can also train the improved image generation model. After training, the weights of the attention layer in the improved image generation model are not shared with the weights of the attention layer in the Chinese image generation model in the related art, so that the setting of the weights in the improved image generation model is more in line with the actual scene, so as to improve the effect of generating images. Figure 6 The training of the image generation model is described in detail.
[0090] Figure 6 It is a flowchart of a method for training an image generation model according to an embodiment of the present disclosure.
[0091] like Figure 6 As shown, the training method of the image generation model of the embodiment 600 includes operations S610 to S650. The image generation model includes a feature extraction network and a text graph model. The feature extraction network can be the aforementioned ResNet, a Transformer-based backbone network, DenseNet or U-Net, etc. The text graph model includes a text understanding network, an image information creation network and an image decoding network, wherein the image information creation network is the aforementioned network with the corresponding network layer added, which will not be described in detail here.
[0092] In operation S610, a feature extraction network is used to extract features of a reference image to obtain reference image features.
[0093] The reference image may be any image provided by a user or any image obtained randomly, and the present disclosure does not limit this, as long as the reference image includes the first target object. The implementation principle of operation S610 is similar to the implementation principle of operation S210 described above, and will not be repeated here.
[0094] In operation S620, a text understanding network is used to extract features of a sample text in the first sample data to obtain text features.
[0095] According to an embodiment of the present disclosure, the text understanding network may include a word segmenter, an embedding vector converter, and a text converter (e.g., a Text transformer), etc. The implementation principle of operation S620 is similar to the implementation principle of operation S220 described above, and will not be repeated here.
[0096] In operation S630, a network is created using image information, a random noise image is used as an initial image of a noise image, and the noise image is iteratively updated according to text features and reference image features to obtain an updated noise image.
[0097] In operation S640, an image decoding network is used to decode the updated noise image to obtain a target image, where the target image includes a second target object similar to the first target object, and the target image matches the sample text.
[0098] According to an embodiment of the present disclosure, the principle of updating the noise image in operation S630 is similar to the principle of updating the noise image in the aforementioned operation S230, except that operation S630 does not involve a process of decoding the updated noise image. Operation S640 is used to implement a process of decoding the updated noise image. The decoding principle is similar to the decoding principle of the image decoding network in the related art Chinese raw image model, and will not be described in detail here.
[0099] In one embodiment, in the iterative update process of implementing operation S630, part of the iterative update process can be implemented by the following operations. For example, in the iterative update process involved in operation S630, at least one update process may include the following operations: fusing the reference image features and the text features with the current noise image to obtain fused features; and updating the current noise image based on the fused features. The implementation principles of these two parts of the operation are similar to the implementation principles of operations S231 to S232 described above, and will not be repeated here.
[0100] In operation S650, a Vincent graph model is trained according to the target image and the first sample image in the first sample data.
[0101] According to an embodiment of the present disclosure, the first sample data may include, for example, a text-image data pair obtained from a public data set, where text is a sample text and image is a sample image that matches the sample text. This embodiment may compare the first sample image in the first sample data with the target image, and train the text graph model with the goal of minimizing the difference between the two. During the training process, the weights in the image information creation network may be adjusted. In one embodiment, the principle of training the text graph model is similar to the related art.
[0102] In one embodiment, before training the text graph model, for example, the feature extraction network may be trained first. Specifically, image data may be obtained from a public data set as a second sample image, and the feature extraction network may be used to extract features of the second sample image, and an image may be predicted based on the extracted features, and the feature extraction network may be trained based on the difference between the predicted image and the original image (i.e., the second sample image).
[0103] Specifically, the feature extraction network can be used to extract the features of the second sample image to obtain the predicted image features. The predicted image features can be understood as the image features output by the last sampling layer in the feature extraction network. Subsequently, based on the predicted image features, a predicted image of the second sample image is generated. For example, this embodiment can use a model similar to the Wensheng graph model in the related art as a graph model, which includes the aforementioned feature extraction network. The difference between the graph model and the Wensheng graph model in the related art is that there is no attention layer in the graph model. Then the predicted image here can be understood as the image output by the graph model with the second sample image as input. Finally, the feature extraction network can be trained based on the second sample image and the predicted image. Specifically, the loss of the graph model can be calculated using the L2 loss function based on the second sample image and the predicted image. Subsequently, the gradient descent algorithm is used to adjust the network parameters in the graph model with the goal of minimizing the loss, thereby realizing the training of the feature extraction network.
[0104] According to an embodiment of the present disclosure, the training of the image generation model as a whole may include, for example, two stages. The first stage is a stage of training the feature extraction network. The second stage is a stage of training the text graph model while keeping the network parameters of the feature extraction network unchanged.
[0105] In one embodiment, the text graph model may include, for example, a stable diffusion model (StableDiffusionModel), and the image generation model of the embodiment of the present disclosure is constructed based on the stable diffusion model.
[0106] In one embodiment, the structure of the image information creation network in the image generation model can be seen, for example, in Figure 4 Noise predictor 420 in or see Figure 5 Noise predictor 520 in .
[0107] Based on the image generation method provided by the present disclosure, the present disclosure also provides an image generation device. Figure 7 The device is described in detail.
[0108] Figure 7 is a structural block diagram of an image generating device according to an embodiment of the present disclosure.
[0109] like Figure 7 As shown, the image generation device 700 of this embodiment may include a first image feature extraction module 710, a first text feature extraction module 720, and a first noise image update module 730. The first noise image update module 730 may include a first fusion submodule 731 and a first update submodule 732.
[0110] The first image feature extraction module 710 is used to extract features of a reference image to obtain reference image features; the reference image includes a first target object. In one embodiment, the first image feature extraction module 710 can be used to perform the operation S210 described above, which will not be described in detail herein.
[0111] The first text feature extraction module 720 is used to extract features of the prompt text to obtain text features. In one embodiment, the first text feature extraction module 720 can be used to perform the operation S220 described above, which will not be described in detail here.
[0112] The first noise image update module 730 is used to use a random noise image as an initial image of the noise image, and iteratively update the noise image according to the text features and the reference image features to generate a target image. The target image includes a second target object similar to the first target object, and the target image matches the prompt text. The target image is obtained by decoding the final noise image. In one embodiment, the first noise image update module 730 can be used to perform the operation S230 described above, which will not be repeated here.
[0113] The first fusion submodule 731 is used to fuse the reference image features and the text features with the current noise image in at least one updating process to obtain fused features. In one embodiment, the first fusion submodule 731 can be used to perform the operation S231 described above, which will not be described in detail here.
[0114] The first updating submodule 732 is used to update the current noise image based on the fusion feature. In one embodiment, the first updating submodule 732 can be used to perform the operation S232 described above, which will not be described in detail here.
[0115] According to an embodiment of the present disclosure, the first image feature extraction module 710 can be specifically used to input a reference image into a feature extraction network composed of multiple sampling layers, and use the features output by each of the multiple sampling layers as a reference image feature. The first fusion submodule 731 can include: a feature acquisition unit, used to input the current noise image into the image information creation network in the text graph model, and obtain the features of the i-th network layer in the input image information creation network as the i-th image feature; and a fusion unit, used to fuse the features and text features output by the i-th sampling layer in the feature extraction network with the i-th image feature respectively. Wherein, i is an integer greater than or equal to 1, and the features of the first network layer in the input image information creation network are the current noise image.
[0116] According to an embodiment of the present disclosure, the fusion unit may include a first fusion subunit, a second fusion subunit and a feature acquisition subunit. The first fusion subunit is used to fuse the features output by the i-th sampling layer with the i-th image features using a cross-attention mechanism to obtain a first sub-fusion feature. The second fusion subunit is used to fuse the text features with the i-th image features using a cross-attention mechanism to obtain a second sub-fusion feature. The feature acquisition subunit is used to obtain the i+1-th image feature based on the first sub-fusion feature and the second sub-fusion feature.
[0117] According to an embodiment of the present disclosure, the fusion unit may include a conversion subunit, a splicing subunit and a fusion subunit. The conversion subunit is used to convert the features output by the i-th sampling layer to the feature space where the text features are located to obtain the converted features. The splicing subunit is used to splice the converted features and the text features to obtain the splicing features. The fusion subunit is used to fuse the splicing features with the i-th image features using a cross-attention mechanism to obtain the i+1-th image features.
[0118] According to an embodiment of the present disclosure, the number of sampling layers in the feature extraction network and the number of network layers in the image information creation network are both N, where N is an integer greater than 1. i is a partial or complete integer among all integers greater than or equal to 1 and less than or equal to N; the fused feature is the feature output by the Nth network layer in the image information creation network.
[0119] According to an embodiment of the present disclosure, the at least one updating process includes a later updating process in the process of iteratively updating the noise image.
[0120] According to an embodiment of the present disclosure, the feature acquisition subunit is used to: weight the first sub-fusion feature with a predetermined weight to obtain a weighted feature; and splice the weighted feature and the second sub-fusion feature to obtain the i+1th image feature, wherein the value of the predetermined weight is positively correlated with the number of iterations of the current iteration.
[0121] Based on the training method of the image generation model provided by the present disclosure, the present disclosure also provides a training device for the image generation model. Figure 8 The device is described in detail.
[0122] Figure 8 It is a structural block diagram of a training device for an image generation model according to an embodiment of the present disclosure.
[0123] like Figure 8As shown, the training device 800 of the image generation model of this embodiment may include a second image feature extraction module 810, a second text feature extraction module 820, a second noise image update module 830, a decoding module 840 and a first training module 850. The image generation model includes a feature extraction network and a text graph model, and the text graph model includes a text understanding network, an image information creation network and an image decoding network.
[0124] The second image feature extraction module 810 is used to extract the features of the reference image using a feature extraction network to obtain reference image features. The reference image includes the first target object. In one embodiment, the second image feature extraction module 810 can be used to perform the operation S610 described above, which will not be repeated here.
[0125] The second text feature extraction module 820 is used to extract features of the sample text in the first sample data using a text understanding network to obtain text features. In one embodiment, the second text feature extraction module 820 can be used to perform the operation S620 described above, which will not be described in detail here.
[0126] The second noise image updating module 830 is used to create a network using image information, use a random noise image as an initial image of the noise image, iteratively update the noise image according to the text features and the reference image features, and obtain an updated noise image. In one embodiment, the second noise image updating module 830 can be used to perform the operation S630 described above, which will not be repeated here.
[0127] In one embodiment, the second noise image update module 830 may include a second fusion submodule and a second update submodule. The second fusion submodule is used to fuse the reference image features and the text features with the current noise image to obtain fused features during at least one update process. The second update submodule is used to update the current noise image based on the fused features.
[0128] The decoding module 840 is used to decode the updated noise image using an image decoding network to obtain a target image. The target image includes a second target object similar to the first target object, and the target image matches the sample text. In one embodiment, the decoding module 840 can be used to perform the operation S640 described above, which will not be repeated here.
[0129] The first training module 850 is used to train the Wensheng graph model according to the target image and the first sample image in the first sample data. In one embodiment, the first training module 850 can be used to perform the operation S650 described above, which will not be described in detail here.
[0130] According to an embodiment of the present disclosure, the second image feature extraction module 810 may also be used to: extract features of the second sample image using a feature extraction network to obtain predicted image features. The training device 800 of the image generation model may also include an image generation module and a second training module. The image generation module is used to generate a predicted image of the second sample image based on the predicted image features. The second training module is used to train the feature extraction network based on the second sample image and the predicted image.
[0131] It should be noted that the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved in the technical solution of this disclosure are in compliance with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and they do not violate public order and good morals. In the technical solution of this disclosure, the user's authorization or consent is obtained before obtaining or collecting user personal information.
[0132] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0133] Fig. 9 is a block diagram of an electronic device 900 for implementing the image generation method or the training method of the image generation model of the embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0134] like Fig. 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0135] A number of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0136] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as an image generation method or a training method for an image generation model. For example, in some embodiments, the image generation method or the training method for an image generation model may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the image generation method or the training method for the image generation model described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the image generation method or the image generation model training method in any other appropriate manner (for example, by means of firmware).
[0137] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0138] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0139] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0141] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0142] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. Among them, the server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0143] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0144] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for generating an image, comprising: Extracting features of a reference image to obtain reference image features; the reference image includes a first target object; Extract the features of the prompt text to obtain text features; Using a random noise image as an initial image of a noise image, iteratively updating the noise image according to the text features and the reference image features to generate a target image; the target image includes a second target object similar to the first target object, and the target image matches the prompt text; Wherein, iteratively updating the noise image according to the text feature and the reference image feature comprises, in at least one updating process: The reference image feature and the text feature are respectively fused with the current noise image to obtain a fused feature, including: using a cross attention mechanism to fuse the reference image feature and the current noise image to obtain a first sub-fused feature; using the cross attention mechanism to fuse the text feature and the current noise image to obtain a second sub-fused feature; obtaining the fused feature based on the first sub-fused feature and the second sub-fused feature; and Based on the fusion feature, updating the current noise image, The target image is obtained by decoding the final noise image.
2. The method according to claim 1, wherein: The step of extracting the features of the reference image to obtain the features of the reference image includes: inputting the reference image into a feature extraction network composed of multiple sampling layers, and taking the features output by each sampling layer in the multiple sampling layers as a reference image feature.
3. The method according to claim 2, wherein: The fusing the reference image feature and the text feature with the current noise image to obtain the fused feature comprises: Input the current noise image into the image information creation network in the Vincent graph model, and obtain the feature of the i-th network layer input into the image information creation network as the i-th image feature; and The features output by the i-th sampling layer in the feature extraction network and the text features are respectively fused with the i-th image features, Wherein, i is an integer greater than or equal to 1, and the feature of the first network layer in the network created by inputting the image information is the current noise image.
4. The method according to claim 3, wherein: The fusing the features output by the i-th sampling layer in the feature extraction network and the text features with the i-th image features respectively comprises: Using a cross attention mechanism to fuse the features output by the i-th sampling layer with the i-th image features, to obtain a first sub-fusion feature; Using a cross attention mechanism to fuse the text feature with the i-th image feature to obtain a second sub-fusion feature; and Based on the first sub-fusion feature and the second sub-fusion feature, an (i+1)th image feature is obtained.
5. The method according to claim 3, wherein: The fusing the features output by the i-th sampling layer in the feature extraction network and the text features with the i-th image features respectively comprises: Convert the features output by the i-th sampling layer to the feature space where the text features are located to obtain converted features; splicing the converted feature and the text feature to obtain a spliced feature; and The splicing feature and the i-th image feature are fused using a cross attention mechanism to obtain the i+1-th image feature.
6. The method according to claim 3, wherein: The number of sampling layers in the feature extraction network and the number of network layers in the image information creation network are both N, where N is an integer greater than 1; i is part or all of all integers greater than or equal to 1 and less than or equal to N; the fusion feature is the feature output by the Nth network layer in the image information creation network.
7. The method according to claim 1, wherein: The at least one updating process includes an updating process after a predetermined number of iterations in the process of iteratively updating the noise image.
8. The method according to claim 4, wherein: The obtaining of the (i+1)th image feature based on the first sub-fusion feature and the second sub-fusion feature comprises: weighting the first sub-fusion features by using a predetermined weight to obtain a weighted feature; and The weighted feature and the second sub-fusion feature are concatenated to obtain the i+1th image feature. The value of the predetermined weight is positively correlated with the number of iterations of the current iteration.
9. A method for training an image generation model, wherein: The image generation model includes a feature extraction network and a text graph model; the text graph model includes a text understanding network, an image information creation network and an image decoding network; the method includes: The feature extraction network is used to extract features of a reference image to obtain reference image features; the reference image includes a first target object; Using the text understanding network to extract features of sample text in the first sample data to obtain text features; Using the image information to create a network, using a random noise image as an initial image of a noise image, iteratively updating the noise image according to the text features and the reference image features, to obtain an updated noise image; The image information creation network is used in at least one updating process: The reference image feature and the text feature are respectively fused with the current noise image to obtain a fused feature, including: using a cross attention mechanism to fuse the reference image feature and the current noise image to obtain a first sub-fused feature; using the cross attention mechanism to fuse the text feature and the current noise image to obtain a second sub-fused feature; obtaining the fused feature based on the first sub-fused feature and the second sub-fused feature; and Based on the fusion feature, updating the current noise image; Decoding the updated noise image using the image decoding network to obtain a target image, wherein the target image includes a second target object similar to the first target object, and the target image matches the sample text; and The Vincent graph model is trained according to the target image and a first sample image in the first sample data.
10. The method according to claim 9, further comprising, before training the text graph model: extracting features of the second sample image using the feature extraction network to obtain predicted image features; generating a predicted image of the second sample image according to the predicted image feature; and The feature extraction network is trained according to the second sample image and the predicted image.
11. An image generating device, comprising: An image feature extraction module, used to extract features of a reference image to obtain reference image features; the reference image includes a first target object; A text feature extraction module is used to extract the features of the prompt text to obtain text features; A noise image updating module, configured to use a random noise image as an initial image of a noise image, and iteratively update the noise image according to the text features and the reference image features to generate a target image; the target image includes a second target object similar to the first target object, and the target image matches the prompt text; Wherein, the noise image updating module comprises: A fusion submodule is used to fuse the reference image feature and the text feature with the current noise image in at least one updating process to obtain a fused feature, including: fusing the reference image feature with the current noise image using a cross attention mechanism to obtain a first fused sub-feature; fusing the text feature with the current noise image using the cross attention mechanism to obtain a second fused sub-feature; obtaining the fused feature based on the first fused sub-feature and the second fused sub-feature; and An updating submodule, used for updating the current noise image based on the fusion feature, The target image is obtained by decoding the final noise image.
12. The device according to claim 11, wherein The image feature extraction module is used for: The reference image is input into a feature extraction network composed of multiple sampling layers, and the feature output by each sampling layer in the multiple sampling layers is used as a reference image feature.
13. The device according to claim 12, wherein: The fusion submodule includes: a feature acquisition unit, configured to input the current noise image into the image information creation network in the Vincent graph model, and obtain the feature of the i-th network layer input into the image information creation network as the i-th image feature; and A fusion unit, used to fuse the features output by the i-th sampling layer in the feature extraction network and the text features with the i-th image features respectively, Wherein, i is an integer greater than or equal to 1, and the feature of the first network layer in the network created by inputting the image information is the current noise image.
14. The device according to claim 13, wherein: The fusion unit comprises: A first fusion subunit is used to fuse the feature output by the i-th sampling layer with the i-th image feature by adopting a cross attention mechanism to obtain a first sub-fusion feature; A second fusion subunit is used to fuse the text feature with the i-th image feature by using a cross attention mechanism to obtain a second sub-fusion feature; and The feature acquisition subunit is used to obtain the (i+1)th image feature based on the first sub-fusion feature and the second sub-fusion feature.
15. The device according to claim 13, wherein: The fusion unit comprises: A conversion subunit, used for converting the features output by the i-th sampling layer into the feature space where the text features are located, to obtain converted features; a concatenation subunit, configured to concatenate the converted feature and the text feature to obtain a concatenated feature; and The fusion subunit is used to fuse the splicing feature with the i-th image feature by adopting a cross attention mechanism to obtain the i+1-th image feature.
16. The apparatus of claim 13, wherein: The number of sampling layers in the feature extraction network and the number of network layers in the image information creation network are both N, where N is an integer greater than 1; i is part or all of all integers greater than or equal to 1 and less than or equal to N; the fusion feature is the feature output by the Nth network layer in the image information creation network.
17. The device according to claim 11, wherein: The at least one updating process includes an updating process after a predetermined number of iterations in the process of iteratively updating the noise image.
18. The device according to claim 14, wherein: The feature acquisition subunit is used for: weighting the first sub-fusion features by using a predetermined weight to obtain a weighted feature; and The weighted feature and the second sub-fusion feature are concatenated to obtain the i+1th image feature. The value of the predetermined weight is positively correlated with the number of iterations of the current iteration.
19. A training device for an image generation model, wherein: The image generation model includes a feature extraction network and a text graph model; the text graph model includes a text understanding network, an image information creation network and an image decoding network; the device includes: An image feature extraction module, configured to extract features of a reference image using the feature extraction network to obtain reference image features; the reference image includes a first target object; A text feature extraction module, used to extract features of a sample text in the first sample data using the text understanding network to obtain text features; A noise image updating module, used to create a network using the image information, use a random noise image as an initial image of the noise image, iteratively update the noise image according to the text features and the reference image features, and obtain an updated noise image; Wherein, the noise image updating module comprises: A fusion submodule is used to fuse the reference image feature and the text feature with the current noise image in at least one updating process to obtain a fused feature, including: fusing the reference image feature with the current noise image using a cross attention mechanism to obtain a first fused sub-feature; fusing the text feature with the current noise image using the cross attention mechanism to obtain a second fused sub-feature; obtaining the fused feature based on the first fused sub-feature and the second fused sub-feature; and An updating submodule, used for updating the current noise image based on the fusion feature; a decoding module, configured to decode the updated noise image using the image decoding network to obtain a target image, wherein the target image includes a second target object similar to the first target object, and the target image matches the sample text; and The first training module is used to train the Vincent graph model according to the target image and the first sample image in the first sample data.
20. The apparatus of claim 19, wherein: The image feature extraction module is further used to: extract features of the second sample image using the feature extraction network to obtain predicted image features; The device also includes: an image generating module, configured to generate a predicted image of the second sample image according to the predicted image feature; and The second training module is used to train the feature extraction network according to the second sample image and the predicted image.
21. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method of any one of claims 1 to 8, or execute the method of any one of claims 9 to 10.
22. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 8, or to execute the method according to any one of claims 9 to 10.
23. A computer program product, comprising a computer program / instruction, wherein the computer program / instruction is stored on at least one of a readable storage medium and an electronic device, and when the computer program / instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented, or the steps of the method according to any one of claims 9 to 10 are implemented.
Citation Information
Patent Citations
Image generation method and device, electronic equipment, storage medium and program product
CN117437317A