Image generation method and device, electronic equipment and storage medium

Through the cross attention mechanism, the constrained text and image encoding in image generation are fused, combined with the inverse diffusion processing, the problem of poor controllability of image generation under multi-constraint conditions in the prior art is solved, and more efficient and accurate image generation is achieved.

CN120182403APending Publication Date: 2025-06-20SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311763780.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, when processing image generation based on multiple constraints, it is difficult to coordinately process different constraints, resulting in poor controllability of the generated images and it is difficult to generate images that meet multiple constraints.

Method used

By obtaining the constrained text and the constrained image, the encoding process is performed separately, the text encoding and image encoding are obtained, and then the two are fused based on the cross attention mechanism to obtain the target fusion feature. Finally, the noise image is reverse diffusion based on the feature to generate the target image.

Benefits of technology

The accuracy and controllability of image generation are improved, and the target image can be generated that meets the constraint conditions reflected by the constraint text and the constraint image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182403A_ABST
    Figure CN120182403A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of image processing, and provides an image generation method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a constraint text and a constraint image, carrying out the coding processing of the constraint text to obtain a text code, and carrying out the coding processing of the constraint image to obtain an image code, and performing fusion processing on the text code and the image code based on a cross attention mechanism to obtain a target fusion feature, and performing inverse diffusion processing on the noise image based on the target fusion feature to obtain a target image. According to the invention, the image generation accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to an image generation method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] Artificial Intelligence Generated Content (AIGC) refers to various forms of content and data such as new text, images, music, videos, 3D interactive content (such as virtual avatars, virtual items, virtual environments), etc., which are autonomously generated in the manner of Generative AI (GAI) based on training data and a generation algorithm model.

[0003] Among them, image generation technology is a popular research direction in AIGC. Currently, the image generation technology in AIGC can already handle simple tasks such as image super-resolution reconstruction, denoising, and brightness enhancement well. However, when dealing with image generation based on multiple constraints, it is difficult to collaboratively process different constraints, resulting in poor controllability of the generated images and difficulty in generating images that meet multiple constraints. Summary of the Invention

[0004] Embodiments of this application provide an image generation method, apparatus, electronic device, and storage medium, which can improve the accuracy of image generation.

[0005] In a first aspect, embodiments of this application provide an image generation method, including:

[0006] Obtain a constraint text and a constraint image;

[0007] Perform encoding processing on the constraint text to obtain a text encoding, and perform encoding processing on the constraint image to obtain an image encoding;

[0008] Perform fusion processing on the text encoding and the image encoding based on a cross-attention mechanism to obtain a target fusion feature;

[0009] Perform inverse diffusion processing on a noise image based on the target fusion feature to obtain a target image.

[0010] In a second aspect, embodiments of this application provide an image generation apparatus, including:

[0011] A condition acquisition module, configured to obtain a constraint text and a constraint image;

[0012] An encoding module, configured to perform encoding processing on the constraint text to obtain a text encoding, and perform encoding processing on the constraint image to obtain an image encoding;

[0013] A fusion module, configured to fuse the text encoding and the image encoding based on a cross-attention mechanism to obtain a target fusion feature;

[0014] An inverse diffusion module, configured to perform inverse diffusion processing on a noise image based on the target fusion feature to obtain a target image.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the image generation method described in the first aspect above are implemented.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium. The computer storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the image generation method described in the first aspect above are implemented.

[0017] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device is caused to execute the image generation method described in any one of the first aspects above.

[0018] The beneficial effects of the embodiments of the present application compared with the prior art are as follows:

[0019] In the embodiments of the present application, after obtaining the constrained text and the constrained image, they are first encoded respectively, and then based on the cross-attention mechanism, the obtained text encoding and image encoding are fused to obtain a target fusion feature. The cross-attention mechanism can simultaneously focus on the constraint conditions reflected in the constrained text and the constrained image of different modalities, so that the obtained target fusion feature can better embed the constraint conditions in text form and the constraint conditions in image form. At the same time, since inverse diffusion processing can reverse-derive from noise and gradually eliminate the noise to reverse-generate an image, therefore, performing inverse diffusion processing on the noise image based on the target fusion feature can eliminate the noise in the noise image based on the target fusion feature and convert the noise image into a target image that simultaneously meets the constraint conditions reflected by the constrained text and the constraint conditions reflected by the constrained image, improving the accuracy of image generation. Moreover, by first fusing the constrained text and the constrained image reflecting different constraint conditions based on the cross-attention mechanism, only one target fusion feature obtained by fusion is needed for image generation, and it is not necessary to consider the collaborative processing of constraint conditions of different modalities during the image generation process, which can better control the generation of the image, that is, improve the controllability of image generation. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art.

[0021] Figure 1 It is a schematic flowchart of an image generation method provided by an embodiment of the present application;

[0022] Figure 2 It is a schematic flowchart of an inverse diffusion process provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic flowchart of an inverse diffusion process provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic structural diagram of an image generation device provided by an embodiment of the present application;

[0025] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0026] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0027] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0028] It should also be understood that the term "and / or" used in the specification and the appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0029] In addition, in the description of the specification and the appended claims of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0030] References to "one embodiment" or "some embodiments" in the description of the present application mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way.

[0031] Example 1:

[0032] When jointly generating images according to multiple constraints such as text, etc., it is difficult to jointly process multiple generation conditions. Usually, dozens or even hundreds of images need to be generated, and then manually select the images that better meet the expected requirements from the large number of generated images. The efficiency and accuracy of image generation are relatively low. Especially when generating images jointly according to the constraint conditions of two different modalities of text and image, the difference in the constraint condition modalities further increases the difficulty of jointly processing multiple generation conditions, and it is usually difficult to generate images that meet the expected requirements.

[0033] In order to improve the efficiency and accuracy of image generation, an embodiment of the present application provides an image generation method.

[0034] In the image generation method provided by the embodiment of the present application, after obtaining the constraint text and the constraint image, the constraint text and the constraint image are first respectively encoded to obtain corresponding text encoding and image encoding, then the text encoding and the image encoding are fused based on the cross-attention mechanism to obtain a target fusion feature, and finally, the noise image is inversely diffused based on the target fusion feature to obtain the required target image.

[0035] Since the cross-attention mechanism can simultaneously focus on the constraint conditions reflected in the constraint text and the constraint image of different modalities, the constraint conditions in text form and image form can be better embedded in the obtained target fusion feature. And the inverse diffusion process can reverse-derive from the noise image and gradually eliminate the noise to reverse-generate the image. Therefore, inversely diffusing the noise image based on the target fusion feature can better eliminate the noise in the noise image based on the target fusion feature and generate a target image that simultaneously meets the constraint conditions reflected by the constraint text and the constraint conditions reflected by the constraint image, improving the accuracy of image generation. Moreover, by fusing the constraint text and the constraint image of different modalities based on the cross-attention mechanism, only one target fusion feature obtained by fusion needs to be used to inversely diffuse the noise image to generate the target image, and there is no need to consider the joint processing of the constraint conditions of different modalities during the generation process, which can better control the generation of the image, that is, improve the controllability of image generation.

[0036] It can be understood that the image generation method provided by the embodiments of the present application can be implemented independently by an electronic device such as a laptop computer, or can be implemented in cooperation by a terminal such as a mobile phone or a smart watch and a server. For example, the user inputs a constraint text and a constraint image on the mobile phone. After the mobile phone obtains the constraint text and the constraint image input by the user, it sends the constraint text and the constraint image to the server. The server encodes the received constraint text and constraint image respectively to obtain corresponding text encoding and image encoding, then fuses the text encoding and the image encoding based on the cross-attention mechanism to obtain a target fusion feature, and then performs inverse diffusion processing on the noise image based on the target fusion feature to obtain a target image. Finally, the target image is sent to the corresponding terminal, so that the user can obtain the required target image through the mobile phone.

[0037] Figure 1 The flowchart of an image generation method provided by an embodiment of the present invention is shown and described in detail as follows:

[0038] Step S101, obtain a constraint text and a constraint image.

[0039] The above-mentioned constraint text refers to the text used to reflect the constraint conditions of the finally required generated image, that is, the text used to guide the generation of the image. For example, assuming the constraint text is "a red car", then the corresponding image generated based on this constraint text contains a red car.

[0040] The above-mentioned constraint image refers to the image used to reflect the constraint conditions of the finally required generated image, and is also used to guide the generation of the image. The above-mentioned constraint image can be an image such as a style image, a draft image, and a structure image.

[0041] In some embodiments, when obtaining the constraint image, multiple constraint images can be obtained. The multiple obtained constraint images can be constraint images of the same constraint type (for example, all belong to structure images), or constraint images of different constraint types (for example, including structure images and style images). The specific setting can be determined according to the actual application scenario, and the embodiments of the present application do not limit this.

[0042] Step S102, perform encoding processing on the above-mentioned constraint text to obtain text encoding, and perform encoding processing on the above-mentioned constraint image to obtain image encoding.

[0043] In some embodiments, when performing encoding processing on the constraint text, the constraint text can be encoded by means such as one-hot encoding or word embedding encoding to obtain text encoding in vector form, so that the text encoding in vector form can be better fused with the image encoding subsequently.

[0044] In some embodiments, in order to further improve the accuracy of image generation, when encoding the constraint text and / or the constraint image, an attention mechanism such as self-attention or cross-attention can be used for encoding, so as to better embed the information in the constraint text and / or the constraint image into its corresponding encoding, that is, to make the text encoding and / or the image encoding more accurately reflect its corresponding constraint conditions and improve the accuracy of the obtained text encoding and / or image encoding.

[0045] For example, the constraint image and the noise image can be used as the inputs of an encoder based on the cross-attention mechanism. The encoder first converts the constraint image into a constraint image feature vector and converts the noise image into a noise vector. Then, the noise vector is used as the query matrix, and the constraint image feature vector is used as the key matrix and the value matrix. The attention weights are calculated based on the query matrix and the key matrix, and then the attention weights are multiplied by the value matrix to obtain a feature vector embedding the information of the noise image, that is, the image encoding. That is, the constraint image is corrected by the noise image, so that there is a certain correlation between the corrected image encoding and the noise image. Similarly, the constraint text feature vector can also be used as the key matrix and the value matrix for the same processing to obtain the text encoding corrected by the noise image, thereby improving the accuracy of the target image obtained by subsequent inverse diffusion processing of the noise image based on the text encoding and the image encoding.

[0046] In the embodiments of the present application, the constraint text and the constraint image in different modalities are first encoded respectively, so that subsequent fusion processing can be better performed based on the encoded constraint text and the constraint image, improving the accuracy of the obtained target fusion feature, and thus improving the accuracy of image generation.

[0047] Step S103: Based on the cross-attention mechanism, fuse the above text encoding and the above image encoding to obtain a target fusion feature.

[0048] The above cross-attention mechanism refers to an attention mechanism that calculates attention on two different sequences. It is necessary to split the input tensor into two parts (such as X1 and X2), use one part (such as X1) as the query matrix (i.e., the query), use the other part (such as X2) as the key matrix and the value matrix (i.e., the key-value), and then calculate the attention according to the query and the key-value, thereby introducing the mutual dependence relationship between the query and the key-value (i.e., X1 and X2).

[0049] In the embodiments of the present application, the text encoding and the image encoding are first fused based on the cross-attention mechanism, and the cross-attention mechanism can introduce the mutual dependence between the two, so that the obtained target fusion feature can better embed the text constraints reflected by the text encoding and the image constraints reflected by the image encoding. At the same time, all the constraints are embedded into a fusion feature encoding, and subsequently, the generation of the image can be based only on the target fusion image, without considering the collaborative processing of multiple constraints of different modalities, thereby being able to better control the generation of the image, that is, improving the controllability of image generation.

[0050] Step S104: Perform inverse diffusion processing on the noise image based on the above target fusion feature to obtain the target image.

[0051] The above noise image can be a pre-generated fixed noise image or a randomly generated noise image. Optionally, when generating the noise image, the noise image can be generated based on noises such as Gaussian noise, salt-and-pepper noise, or random noise.

[0052] The above inverse diffusion processing is the reverse diffusion processing. In the diffusion processing of the image, usually noise is gradually added to the image until the image is diffused into a pure noise, and the reverse diffusion processing is to gradually remove the noise from the noise to restore the required image, and the constraint condition is used to constrain the restored image during the process of removing the noise from the noise to restore the image, so that the finally obtained image meets the constraint condition.

[0053] It should be noted that the result obtained by performing inverse diffusion processing on the noise image based on the target fusion feature can be an image, that is, the target image is directly obtained by performing inverse diffusion processing on the noise image based on the target fusion feature. The result obtained by performing inverse diffusion processing on the noise image based on the target fusion feature can also be the image encoding in the hidden layer space. At this time, it is necessary to use a pre-trained decoder to decode the image encoding in the hidden layer space into an image to obtain the target image.

[0054] In the embodiments of the present application, the inverse diffusion processing is directly performed on the noise image based on a target fusion feature that integrates various constraint conditions, so that it is not necessary to consider the collaborative processing of constraint conditions of different modalities during the inverse diffusion processing, nor to separately fuse multiple constraint conditions into the inverse diffusion processing process, which can improve the accuracy of the generated target image and the image generation efficiency at the same time.

[0055] In the embodiments of the present application, after encoding the text encoding corresponding to the constraint text and the image encoding corresponding to the constraint image, based on the cross-attention mechanism, the obtained text encoding and image encoding are fused to obtain the target fusion feature. The cross-attention mechanism can simultaneously focus on the constraint conditions reflected in the constraint text and constraint image of different modalities, and embed the information in the constraint text and constraint image into the obtained target fusion feature better, so that the target fusion feature can more accurately reflect all the constraint conditions. At the same time, since the inverse diffusion process can reverse-derive from noise, gradually eliminate the noise to reverse-generate an image, and the constraint conditions can guide the generated image during the process of eliminating the noise, making the generated image conform to the constraint conditions. Therefore, performing inverse diffusion processing on the noise image based on the target fusion feature, that is, eliminating the noise in the noise image based on the target fusion feature, can convert the noise image into a target image that simultaneously conforms to the constraint conditions reflected by the constraint text and the constraint conditions reflected by the constraint image, improving the information in image generation. Only one target fusion feature obtained by fusion is required for image generation, and it is not necessary to consider the collaborative processing of constraint conditions of different modalities during the image generation process, which can better control the image generation, that is, improve the controllability of image generation.

[0056] It can be understood that the type of the above-mentioned constraint image can be determined according to the actual application scenario, and the encoders or networks corresponding to the encoding process, fusion process, and inverse diffusion process involved in the above image generation method can be integrated into the same model. For example, when applying the above image generation method to a portrait generator to generate the required portrait painting of a person, a person contour map and a style map can be obtained as the constraint images. That is, when a portrait painting of a person needs to be generated, two constraint images, namely a person contour map and a style map, need to be obtained, and the constraint text is obtained. Then, the constraint text, the person contour map, and the style map are used as the inputs of the portrait generator, and a target image that simultaneously conforms to the constraint text, the person contour map, and the style map output by the portrait generator can be obtained. It can be understood that the portrait generator can integrate a text encoder, an image encoder, a fusion network, and an inverse diffusion network. After inputting the constraint text, the person contour map, and the style map into the portrait generator, the text encoder and the image encoder in the portrait generator respectively perform encoding processing on their corresponding constraint conditions to obtain a text encoding and an image encoding. Then, the fusion network fuses the text encoding and the image encoding to obtain the target fusion feature, and this target fusion feature is used as the input of the inverse diffusion network. The inverse diffusion network performs inverse diffusion processing on the noise image according to this target fusion feature, obtains the target image and outputs it, so as to obtain the required portrait painting of a person.

[0057] For another example, when applying the above image generation method to an image reconstruction model to reconstruct a low-resolution image and wait for the image to be reconstructed, the constraint image may include a style map and a structure map determined according to the image to be reconstructed. For example, the style features of the image to be reconstructed can be extracted by a pre-trained style extraction model to obtain the style map, and the structure features of the image to be reconstructed can be extracted by a pre-trained structure extraction model to obtain the structure map. Then, the constraint text, the structure map, and the style map obtained based on the image to be reconstructed are used as the input of the image reconstruction model. The text encoder and the image encoder in the image reconstruction model respectively encode their corresponding constraint conditions to obtain text encoding and image encoding. Then, the fusion network fuses the text encoding and the image encoding to obtain a target fusion feature. This target fusion feature is used as the input of the inverse diffusion network. The inverse diffusion network performs inverse diffusion processing on the noise image according to this target fusion feature to obtain a target image and output it. This target image is the image reconstructed based on the image to be reconstructed, which conforms to the constraint text, the structure map, and the style map at the same time.

[0058] In some embodiments, before the above step S102, it further includes:

[0059] When the number of the above constraint images is greater than or equal to 2, perform a merging process on the above image encodings corresponding to each of the above constraint images to obtain a fused image encoding;

[0060] Correspondingly, the above step S102 includes:

[0061] Perform a fusion process on the above text encoding and the above fused image encoding based on the above cross-attention mechanism to obtain the above target fusion feature.

[0062] Specifically, when the number of the obtained constraint images is greater than or equal to 2, that is, when multiple constraint images are obtained, since each constraint image belongs to a two-dimensional image, the image encodings corresponding to each constraint image can be first merged to obtain a fused image encoding that can reflect the constraint conditions corresponding to each constraint image. When performing a fusion process on the text encoding and the image encoding based on the attention mechanism later, only one fused image encoding that reflects the constraint conditions of each constraint image needs to be fused with the text encoding, improving the processing efficiency.

[0063] It can be understood that in the actual application scenario, the number of constraint texts is usually 1. If the number of the obtained constraint texts is greater than or equal to 2, the text encodings corresponding to each constraint text can also be concatenated or merged to obtain a fused text encoding.

[0064] For example, when the semantic differences between the respective constraint texts are relatively large, the text encodings can be concatenated so that the fused text encoding can reflect the constraint conditions corresponding to the respective constraint texts. When the semantic differences between the respective constraint texts are relatively small, the text encodings can be fused to reduce the size and complexity of the obtained fused text encoding, thereby reducing the difficulty of subsequent processing. Similarly, when the number of constraint images is greater than or equal to 2, images with relatively high similarity can be merged, and images with relatively low similarity can be concatenated.

[0065] In the embodiments of the present application, the image encodings corresponding to the obtained multiple constraint images are merged, so that only the fused image encoding obtained by fusing multiple image encodings needs to be processed subsequently, reducing the complexity of subsequent processing, thereby improving the image generation efficiency.

[0066] In some embodiments, the above steps perform a merging process based on the above image encodings corresponding to the respective above constraint images to obtain a fused image encoding, including:

[0067] Perform a normalization process on the above image encodings corresponding to the respective above constraint images to obtain the normalized above image encodings.

[0068] Merge the normalized above image encodings pixel by pixel to obtain the above fused image encoding.

[0069] Specifically, since the resolutions of different constraint images may be different, resulting in different resolutions of the corresponding image encodings, in the embodiments of the present application, before merging the respective image encodings, a normalization process is first performed on the respective image encodings. Through the normalization process, the respective image encodings are converted into image encodings of the same size, so that each pixel between the respective image encodings can correspond one by one, so as to merge the respective image encodings pixel by pixel, that is, the pixels corresponding to each pixel position in each image encoding are merged into one pixel. After the pixels at each pixel position are merged, the required fused image encoding is obtained.

[0070] In the embodiments of the present application, the respective image encodings are first normalized into image encodings of the same size, so that the pixels at each pixel position between the respective image encodings can correspond one by one, thereby reducing the processing difficulty of image encoding merging and ensuring the rationality of the generated fused image encoding.

[0071] In some embodiments, the above target fusion feature includes a first fusion feature and a second fusion feature, and the attention weights corresponding to the above first fusion feature and the above second fusion feature are different. The above S103 includes:

[0072] A. Based on the above cross-attention mechanism, perform fusion processing with different attention weights on the above text encoding and the above image encoding respectively to obtain the above first fusion feature and the above second fusion feature.

[0073] Correspondingly, the above step S104 includes:

[0074] Perform inverse diffusion processing on the above noise image based on the above first fusion feature and the above second fusion feature to obtain the above target image.

[0075] Specifically, in order to fully embed the constraint conditions of different modalities, when performing fusion processing according to text encoding and image encoding, different weighted attentions can be used to perform at least two fusion processes, that is, at least two fusion processes are performed. Among them, the attention weights used in the two fusion processes are different, and they respectively focus on the constraint conditions reflected by the text and the constraint conditions reflected by the image, and obtain their corresponding fusion features, that is, the first fusion feature and the second fusion feature focus on the constraint conditions of different modalities. Correspondingly, when performing inverse diffusion processing to obtain the target image, the inverse diffusion process of the noise image will be jointly guided by the first fusion feature and the second fusion feature, so that the generated target image is more consistent with the constraint text and the constraint image.

[0076] In the embodiments of the present application, two fusion processes are performed on the text encoding and the image encoding with different attention weights to obtain the first fusion feature and the second fusion feature that focus on the constraint conditions of different modalities, and both the first fusion feature and the second fusion feature embed all the constraint conditions. Performing inverse diffusion processing based on the first fusion feature and the second fusion feature to obtain the target image can consider the characteristics of the constraint conditions of different modalities respectively while better generating images by coordinating the constraint images of different modalities, thereby improving the accuracy of image generation.

[0077] In some embodiments, the above step A includes:

[0078] A1. Determine a first query matrix based on the above text encoding, and determine a first key matrix and a first value matrix based on the above image encoding.

[0079] A2. Determine a first attention weight according to the above first query matrix and the above first key matrix.

[0080] A3. Determine the above first fusion feature according to the above first attention weight and the above first value matrix.

[0081] A4. Determine a second query matrix based on the above image encoding, and determine a second key matrix and a second value matrix based on the above text encoding.

[0082] A5. Determine a second attention weight according to the above second query matrix and the above second key matrix.

[0083] A6. Determine the second fusion feature according to the above second attention weight and the above second value matrix.

[0084] It should be noted that the above steps A1 - A3 are used to determine the first fusion feature, and steps A4 - A6 are used to determine the second fusion feature. In practical applications, the relationship between the above steps A1 - A3 and steps A4 - A6 can be a parallel execution relationship or a serial execution relationship, that is, the first fusion feature and the second fusion feature can be determined simultaneously, or one of the fusion features can be determined first (for example, first execute steps A4 - A6 to determine the second fusion feature, and then execute steps A1 - A3 to determine the second fusion feature), and no specific limitation is made here.

[0085] Specifically, in order to further improve the accuracy of image generation, the text encoding and the image encoding can be used as queries respectively, and the other as the key and value for calculating the attention weight, and then based on the different calculated attention weights, a fusion feature that focuses on reflecting the constraints of different modalities is obtained.

[0086] That is, when determining the first fusion feature, the first query matrix (i.e., the query) is determined according to the text encoding, the first key matrix and the first value matrix (i.e., the key and value) are determined according to the image encoding, then the dot product calculation is performed using the first query matrix and the first key matrix to obtain the attention score, and the obtained attention score is normalized through the softmax function to obtain the first attention weight, and then the first attention weight is multiplied by the first value matrix to obtain the first fusion feature. Similarly, when determining the second fusion feature, after determining the second query matrix according to the image encoding and the second key matrix and the second value matrix according to the text encoding, the same calculation is performed to obtain the second fusion feature, which will not be elaborated here.

[0087] In some embodiments, the above first fusion feature is based on the following form:

[0088]

[0089] C text is the image encoding, C im is the text encoding, Q(C text ) is the first query matrix, and the softmax function is used to map the input features to a probability distribution (i.e., the attention score), that is is the first attention weight, K(C im ) is the first key matrix, d is the dimension of the first key matrix, K(C im ) T is the transpose of the first key matrix, and V(C im ) is the first value matrix.

[0090] The above second fusion feature is represented in the following form:

[0091]

[0092] Q(C im ) is the second query matrix, is the second attention weight, K(C text ) is the second key matrix, d is the dimension of the second key matrix, K(C text ) T is the transpose of the second key matrix, V(C text ) is the second value matrix.

[0093] It can be understood that the calculated first attention weight and second attention weight can be expressed in matrix form.

[0094] It can be understood that the text encoding and the image encoding are encodings with different shapes. The text encoding and the image encoding can be mapped through linear transformation or different attention weight matrices to obtain text encoding and image encoding with the same shape and dimension, so as to fuse them.

[0095] In the embodiments of the present application, in the two fusion processes, the text encoding and the image encoding are respectively used as keys and values, and the other encoding is used as a query to obtain the first fusion feature that focuses on constraining the constraints of the image and the second fusion feature that focuses on constraining the constraints of the text, introducing the mutual dependence relationship between the constraints of the text and the constraints of the image. While obtaining the fusion feature embedding all the constraints, different modal constraints are emphasized, so that the generated target image can fit the constrained text and the constrained image at the same time.

[0096] In some embodiments, the above step S104 includes:

[0097] B1. Using the above noise image, the above first fusion feature, and the above second fusion feature as the input of the inverse diffusion network, and obtaining the denoised above noise image output by the inverse diffusion network. The inverse diffusion network is used to denoise the above noise image based on the above cross-attention mechanism through the above first fusion feature and the above second fusion feature to obtain the denoised noise image.

[0098] B2. Determining the above target image based on the denoised above noise image.

[0099] The above inverse diffusion network can be a denoising diffusion probabilistic model (DDPM), a segmentation network with a UNet structure, or other neural networks, and the embodiments of the present application do not limit this.

[0100] Specifically, when performing inverse diffusion processing on the noisy image based on the first fusion feature and the second fusion feature, a pre-trained inverse diffusion network can be used to better fuse different first fusion features and second fusion features into the inverse diffusion process based on the cross-attention mechanism, generating a target image with higher accuracy.

[0101] Among them, when generating the target image through the inverse diffusion network, the noisy image, the first fusion feature, and the second fusion feature are input into the inverse diffusion network. The inverse diffusion network uses the first fusion feature and the second fusion feature together to perform denoising processing on the noisy image, obtaining the denoised noisy image.

[0102] In some embodiments, since the denoised noisy image obtained by one-time denoising processing may still have noise, that is, the noisy image obtained by one-time denoising processing may not meet the requirements. Therefore, in order to improve the accuracy of the obtained target image, the noisy image can be denoised multiple times to further improve the accuracy of the obtained target image. For example, when the denoised noisy image does not meet the user's requirements (such as the accuracy reaches 0.9), the denoised noisy image is input into the inverse diffusion network again for denoising processing. Another example is that N (N>1) cascaded inverse diffusion networks can be set, that is, the noisy image is denoised N times through N inverse diffusion networks.

[0103] In some embodiments, when performing denoising processing on the denoised noisy image (i.e., the second to the Nth denoising processing), the constraint image and the constraint text can be re-encoded, and then the denoised noisy image can be denoised based on the new text encoding and image encoding. For example, as Figure 2 shown, after obtaining the current denoised noisy image, the text encoding and the image encoding are re-obtained to denoise the current denoised noisy image, and during the encoding process, the constraint text and the constraint image are corrected in combination with the current denoised noisy image, that is, the information of the denoised noisy image is embedded into the constraint image based on the cross-attention mechanism, and the information of the denoised noisy image is embedded into the constraint text based on the cross-attention mechanism, obtaining the text encoding and the image encoding embedded with the information of the current denoised noisy image, and re-generating the target fusion feature, so that during each denoising processing, the actual situation of the current noisy image can be combined to perform denoising processing on it, further improving the accuracy of the denoised noisy image.

[0104] It should be noted that in some other embodiments, it is also possible to select to perform only encoding processing on the constrained text or only on the constrained image according to the actual situation of the current denoised noise image, and then perform denoising processing on the denoised noise image only based on text encoding or only based on image encoding. At the same time, in each denoising process after the first denoising process, the input features can be different. For example, in the second denoising process, denoising processing is performed on the encoded noise image based on the new image encoding, in the third denoising process, denoising processing is performed on the noise image based on the new text encoding and image encoding, and in the fourth denoising process, denoising processing is performed on the noise image based on the new image encoding. The constraint conditions for denoising processing are determined based on the actual situation of the current denoised image, improving the controllability and accuracy of the finally generated target image.

[0105] For example, assume that N is 5, that is, there are 5 cascaded inverse diffusion networks. During the denoising process based on the inverse diffusion network, the noise image, the first fusion feature, and the second fusion feature are input into the first inverse diffusion network, and the first inverse diffusion network performs denoising processing on the noise image and outputs the denoised noise image; for the second inverse diffusion network, first perform encoding processing on the constrained text and the constrained image respectively based on the cross-attention mechanism and the denoised noise image, and generate new first and second fusion features based on the obtained text encoding and image encoding. Then, use the current denoised noise image, the new first fusion feature, and the second fusion feature as the input of the second inverse diffusion network. The second inverse diffusion network performs denoising processing on the denoised noise image and outputs the denoised noise image, and so on, to obtain the denoised noise image output by the Nth inverse diffusion network, that is, the target image.

[0106] Use the above noise image, the above first fusion feature, and the above second fusion feature as the input of the inverse diffusion network to obtain the denoised above noise image output by the inverse diffusion network. The inverse diffusion network is used to perform denoising processing on the noise image based on the above cross-attention mechanism through the above first fusion feature and the above second fusion feature to obtain the denoised noise image.

[0107] In the embodiments of the present application, since both the first fusion feature and the second fusion feature are features of the same model embedded with all the constraint conditions, and respectively focus on the constraint conditions of the constrained text and the constraint conditions of the constrained image, therefore, generating the target image based on the common features of the first fusion feature and the second fusion can better synergistically process all the constraint conditions while making the generated target image more conform to the constraint conditions of the constrained text and the constraint conditions of the constrained image, that is, improving the accuracy of the target image.

[0108] In some embodiments, the inverse diffusion network includes a first network and a second network, and step B1 includes:

[0109] Taking the noise image and the first fusion feature as the input of the first network, obtaining an intermediate noise image output by the first network, where the first network is used to denoise the noise image through the first fusion feature based on the cross-attention mechanism to obtain the intermediate noise image;

[0110] Taking the intermediate noise image and the second fusion feature as the input of the second network, obtaining the denoised noise image output by the second network, where the second network is used to denoise the intermediate noise image through the second fusion feature based on the cross-attention mechanism to obtain the denoised noise image.

[0111] Figure 3 The flowchart of generating the denoised noise image based on the inverse diffusion network is shown. As Figure 3 shown, first, the first fusion feature is used as the key matrix (i.e., key, K) and the value matrix (i.e., value, V), and the noise image is used as the query matrix (i.e., query, Q). After calculating the attention weights based on the query matrix and the key matrix, the attention weights are then multiplied by the value matrix to obtain the intermediate noise image for denoising processing, and the resulting matrix is the intermediate noise image. It can be understood that when the second network denoises the intermediate noise image, the intermediate noise image is used as the query matrix, and the second fusion feature is used as the key matrix and the value matrix for cross-attention calculation to obtain the denoised noise image.

[0112] In some embodiments, the intermediate noise image can be expressed in the following form:

[0113]

[0114] Z t is the noise image, C1 is the first fusion feature, Q(Z t ) is the query matrix, K(C1) is the key matrix, K(C1) T is the transpose of the key matrix, V(C1) is the value matrix, and d is the dimension of the key matrix K(C1).

[0115] The denoised noise image can be expressed in the following form:

[0116]

[0117] Z t+1 is the denoised noise image, C2 is the second fusion feature, is the query matrix, K(C2) is the key matrix, K(C2) Tis the transpose of the key matrix, V(C2) is the value matrix, and d is the dimension of the key matrix K(C2).

[0118] It should be noted that the first network can also perform denoising on the noisy image based on the cross-attention mechanism through the second fusion feature to obtain an intermediate noisy image, and then the second network performs denoising on the intermediate noisy image based on the cross-attention mechanism through the first fusion feature to obtain the denoised noisy image.

[0119] It can be understood that when the noisy image is denoised N times through the inverse diffusion network, Z t i.e., the current noisy image (i.e., the denoised noisy image obtained after the t-th denoising process), Z t+1 i.e., the noisy image after denoising the current noisy image, and Z N (i.e., the noisy image after the N-th denoising process) is the finally obtained target image.

[0120] In the embodiments of the present application, the noisy image is denoised successively through the first fusion feature and the second fusion feature. The first fusion feature and the second fusion feature focus on different modal constraint conditions and both embed all the constraint conditions, which can better describe the constraint conditions, so that the generated denoised noisy image better conforms to each constraint condition, improving the accuracy of the finally obtained target image.

[0121] In some embodiments, one or more of the text encoder (i.e., the encoder for encoding the constraint text), the image encoder (i.e., the encoder for encoding the constraint image), the feature fusion network (the network for fusing text encoding and image encoding based on the attention mechanism), the inverse diffusion network, and the first network and the second network in the inverse diffusion network can be pre-trained to obtain networks with better performance to further improve the accuracy of image generation.

[0122] For example, multiple sample images can be obtained, and the text descriptions corresponding to the sample images can be generated through a pre-trained text extraction model or manually as constraint texts. Various types of images such as structure diagrams, style diagrams, or contour diagrams in the sample images can be generated through a pre-trained extraction model or manually as constraint images. Then, for the obtained constraint texts and corresponding constraint images, after encoding them in sequence through the corresponding encoders, the encoded features are fused based on a feature fusion network to obtain target fusion features (which may include first fusion features and second fusion features). Then, the target fusion features are input into an inverse diffusion network for processing to obtain the generated image output by the inverse diffusion network. After obtaining the generated image, the network parameters of the first network and the second network in the inverse diffusion network can be adjusted according to the difference between the generated image and the corresponding sample image. Alternatively, the encoding parameters of the encoder and / or the network parameters of the feature fusion network can also be adjusted according to the difference between the generated image and the corresponding sample image. For example, when the style difference between the generated image and the corresponding sample image is large, the encoding parameters of the encoder can be adjusted according to the difference, so that the encoder can better extract more accurate text encodings and image encodings, thereby improving the difference between the generated image generated based on the text encoding and the image encoding and the corresponding sample image, that is, improving the accuracy of the generated image.

[0123] For another example, in order to generate images that are more accurately satisfactory to users, the constraint text and constraint image can also be directly obtained. After encoding and fusing the constraint text and constraint image, and generating the corresponding generated image through the inverse diffusion network, the user scores the generated image to obtain the score of the generated image. Then, the network parameters of the inverse diffusion network or the encoder, fusion feature network, etc. are adjusted according to the difference between the score and the maximum score. For example, the parameters to be adjusted can be determined according to the level of the difference. For example, when the maximum score is 100 and the score of the generated image is 92, the difference between the score and the maximum score is 8, belonging to the first level (the difference is within 0 - 10). At this time, the encoding parameters of the encoder and the network parameters of the feature fusion network can be adjusted based on the difference. Another example is that when the score of the generated image is 72 and the difference from the maximum score is 28, belonging to the second level (the difference is greater than 10 and less than 30), the network parameters of the inverse diffusion network can be adjusted based on the difference.

[0124] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0125] Embodiment 2:

[0126] Corresponding to the image generation method described in the above embodiment, Figure 4The block diagram of the image generation device provided by the embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.

[0127] Referring to Figure 4 , the device includes: a condition acquisition module 41, an encoding module 42, a fusion module 43, and an inverse diffusion module 44. Among them,

[0128] The condition acquisition module is used to acquire constraint text and constraint images.

[0129] The encoding module is used to perform encoding processing on the above-mentioned constraint text to obtain text encoding, and perform encoding processing on the above-mentioned constraint image to obtain image encoding.

[0130] The fusion module is used to perform fusion processing on the above-mentioned text encoding and the above-mentioned image encoding based on the cross-attention mechanism to obtain a target fusion feature.

[0131] The inverse diffusion module is used to perform inverse diffusion processing on the noise image based on the above-mentioned target fusion feature to obtain a target image.

[0132] In some embodiments, the above-mentioned target fusion feature includes a first fusion feature and a second fusion feature, and the attention weights corresponding to the above-mentioned first fusion feature and the above-mentioned second fusion feature are different. The fusion module includes:

[0133] The fusion unit is used to perform fusion processing with different attention weights on the above-mentioned text encoding and the above-mentioned image encoding respectively based on the above-mentioned cross-attention mechanism to obtain the above-mentioned first fusion feature and the above-mentioned second fusion feature.

[0134] Correspondingly, the above-mentioned inverse diffusion module is used to perform inverse diffusion processing on the above-mentioned noise image based on the above-mentioned first fusion feature and the above-mentioned second fusion feature to obtain the above-mentioned target image.

[0135] In some embodiments, the above-mentioned fusion module includes:

[0136] The first calculation unit is used to determine a first query matrix based on the above-mentioned text encoding, and determine a first key matrix and a first value matrix based on the above-mentioned image encoding.

[0137] The first attention weight acquisition unit is used to determine a first attention weight according to the above-mentioned first query matrix and the above-mentioned first key matrix.

[0138] The first fusion unit is used to determine the above-mentioned first fusion feature according to the above-mentioned first attention weight and the above-mentioned first value matrix.

[0139] The second calculation unit is used to determine a second query matrix based on the above-mentioned image encoding, and determine a second key matrix and a second value matrix based on the above-mentioned text encoding.

[0140] A second attention weight acquisition unit, configured to determine a second attention weight according to the above-mentioned second query matrix and the above-mentioned second key matrix.

[0141] A second fusion unit, configured to determine the above-mentioned second fusion feature according to the above-mentioned second attention weight and the above-mentioned second value matrix.

[0142] In some embodiments, the above-mentioned inverse diffusion module includes:

[0143] An inverse diffusion processing unit, configured to use the above-mentioned noise image, the above-mentioned first fusion feature, and the above-mentioned second fusion feature as inputs of an inverse diffusion network, and obtain the denoised above-mentioned noise image output by the inverse diffusion network. The inverse diffusion network is configured to perform denoising processing on the above-mentioned noise image based on the above-mentioned cross-attention mechanism through the above-mentioned first fusion feature and the above-mentioned second fusion feature to obtain the denoised noise image.

[0144] A target image acquisition unit, configured to determine the above-mentioned target image based on the denoised above-mentioned noise image.

[0145] In some embodiments, the above-mentioned inverse diffusion network includes a first network and a second network, and the above-mentioned inverse diffusion module includes:

[0146] A first network processing unit, configured to use the above-mentioned noise image and the above-mentioned first fusion feature as inputs of the above-mentioned first network, and obtain an intermediate noise image output by the above-mentioned first network. The above-mentioned first network is configured to perform denoising processing on the above-mentioned noise image based on the above-mentioned cross-attention mechanism through the above-mentioned first fusion feature to obtain the above-mentioned intermediate noise image.

[0147] A second network processing unit, configured to use the above-mentioned intermediate noise image and the above-mentioned second fusion feature as inputs of the above-mentioned second network, and obtain the denoised above-mentioned noise image output by the above-mentioned second network. The above-mentioned second network is configured to perform denoising processing on the above-mentioned intermediate noise image based on the above-mentioned cross-attention mechanism through the above-mentioned second fusion feature to obtain the denoised above-mentioned noise image.

[0148] In some embodiments, the above-mentioned image generation device further includes:

[0149] A merging module, configured to perform a merging process on the above-mentioned image encodings corresponding to each of the above-mentioned constraint images to obtain a fused image encoding when the number of the above-mentioned constraint images is greater than or equal to 2.

[0150] Correspondingly, the above-mentioned fusion module is configured to perform a fusion process on the above-mentioned text encoding and the above-mentioned fused image encoding based on the above-mentioned cross-attention mechanism to obtain the above-mentioned target fusion feature.

[0151] In some embodiments, the above image generation device further includes:

[0152] A normalization module, configured to perform normalization processing based on the above image encodings corresponding to each of the above constraint images, so as to obtain the normalized above image encodings.

[0153] A pixel-by-pixel merging module, configured to merge the normalized above image encodings pixel by pixel to obtain the above fusion image encoding.

[0154] It should be noted that for the information interaction, execution process, etc. between the above devices / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought, reference can be specifically made to the method embodiment part, and details are not described herein again.

[0155] Embodiment Three:

[0156] Figure 5 This is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As Figure 5 shown, the electronic device 5 in this embodiment includes: at least one processor 50 ( Figure 5 only one processor is shown in the figure), a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50. When the processor 50 executes the computer program 52, the steps in any of the above method embodiments are implemented.

[0157] The electronic device 5 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art can understand that Figure 5 merely examples of the electronic device 5 are given, which do not constitute a limitation on the electronic device 5. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0158] The so-called processor 50 may be a central processing unit (CPU), and the processor 50 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0159] In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. In other embodiments, the memory 51 may also be an external storage device of the electronic device 5, such as a plug-in hard disk equipped on the electronic device 5, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 51 may also include both the internal storage unit and the external storage device of the electronic device 5. The memory 51 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program. The memory 51 may also be used to temporarily store data that has been output or will be output.

[0160] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0161] The embodiments of the present application also provide a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, the steps in any of the foregoing method embodiments are implemented.

[0162] The embodiments of the present application also provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in any of the foregoing method embodiments can be implemented.

[0163] The embodiments of the present application provide a computer program product. When the computer program product runs on an electronic device, the electronic device is caused to implement the steps in any of the foregoing method embodiments.

[0164] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0165] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0166] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0167] In the embodiments provided by the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0168] The unit described as the separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0169] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. An image generation method, characterized in that, Including: Obtaining constraint text and constraint images; Performing encoding processing on the constraint text to obtain text encoding, and performing encoding processing on the constraint images to obtain image encoding; Performing fusion processing on the text encoding and the image encoding based on a cross-attention mechanism to obtain a target fusion feature; Performing inverse diffusion processing on a noise image based on the target fusion feature to obtain a target image.

2. The image generation method according to claim 1, characterized in that, The target fusion feature includes a first fusion feature and a second fusion feature, and the attention weights corresponding to the first fusion feature and the second fusion feature are different. The performing fusion processing on the text encoding and the image encoding based on the cross-attention mechanism to obtain a target fusion feature includes: Performing fusion processing with different attention weights on the text encoding and the image encoding respectively based on the cross-attention mechanism to obtain the first fusion feature and the second fusion feature; Correspondingly, the performing inverse diffusion processing on the noise image based on the target fusion feature to obtain a target image includes: Performing inverse diffusion processing on the noise image based on the first fusion feature and the second fusion feature to obtain the target image.

3. The image generation method according to claim 2, characterized in that, The performing fusion processing with different attention weights on the text encoding and the image encoding respectively based on the cross-attention mechanism to obtain the first fusion feature and the second fusion feature includes: Determining a first query matrix based on the text encoding, and determining a first key matrix and a first value matrix based on the image encoding; Determining a first attention weight according to the first query matrix and the first key matrix; Determining the first fusion feature according to the first attention weight and the first value matrix; Determining a second query matrix based on the image encoding, and determining a second key matrix and a second value matrix based on the text encoding; Determining a second attention weight according to the second query matrix and the second key matrix; Determining the second fusion feature according to the second attention weight and the second value matrix.

4. The image generation method according to claim 2, characterized in that, The performing inverse diffusion processing on the noise image based on the first fusion feature and the second fusion feature to obtain the target image includes: Taking the noise image, the first fusion feature, and the second fusion feature as inputs of an inverse diffusion network to obtain the denoised noise image output by the inverse diffusion network. The inverse diffusion network is used to perform denoising processing on the noise image based on the cross-attention mechanism through the first fusion feature and the second fusion feature to obtain the denoised noise image; Determining the target image based on the denoised noise image.

5. The image generation method according to claim 4, characterized in that, The inverse diffusion network includes a first network and a second network. The taking the noise image, the first fusion feature, and the second fusion feature as inputs of the inverse diffusion network to obtain the denoised noise image output by the inverse diffusion network includes: Taking the noise image and the first fusion feature as the input of the first network, an intermediate noise image output by the first network is obtained. The first network is used to denoise the noise image through the first fusion feature based on the cross-attention mechanism to obtain the intermediate noise image; Taking the intermediate noise image and the second fusion feature as the input of the second network, the denoised noise image output by the second network is obtained. The second network is used to denoise the intermediate noise image through the second fusion feature based on the cross-attention mechanism to obtain the denoised noise image.

6. The image generation method according to any one of claims 1 to 5, characterized in that, Before fusing the text encoding and the image encoding based on the cross-attention mechanism to obtain the target fusion feature, it further includes: When the number of the constraint images is greater than or equal to 2, performing a merging process on the image encodings corresponding to the respective constraint images to obtain a fused image encoding; Correspondingly, the fusing the text encoding and the image encoding based on the cross-attention mechanism to obtain the target fusion feature includes: Fusing the text encoding and the fused image encoding based on the cross-attention mechanism to obtain the target fusion feature.

7. The image generation method according to claim 6, characterized in that, The performing a merging process on the image encodings corresponding to the respective constraint images to obtain a fused image encoding includes: Performing a normalization process on the image encodings corresponding to the respective constraint images to obtain the normalized image encodings; Merging the normalized image encodings pixel by pixel to obtain the fused image encoding.

8. An image generation device, characterized in that, It includes: A conditional acquisition module, configured to acquire a constraint text and constraint images; An encoding module, configured to perform an encoding process on the constraint text to obtain a text encoding, and perform an encoding process on the constraint images to obtain image encodings; A fusion module, configured to fuse the text encoding and the image encodings based on the cross-attention mechanism to obtain a target fusion feature; An inverse diffusion module, configured to perform an inverse diffusion process on a noise image based on the target fusion feature to obtain a target image.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.