Image generation method, device, computer device and storage medium

By extracting and fusing the style features of multiple reference images and combining content features to generate target images, the limitations of image generation in the prior art are solved, and the solution of zero-sample style transmission is realized, and the flexibility and accuracy of generating images are improved.

CN119417931BActive Publication Date: 2025-06-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510013340.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-06-24
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

When generating images of a specified style, the prior art is limited to using the trained model to generate images of this style, and it is impossible to generate images of other styles, resulting in greater limitations in image generation.

Method used

By acquiring content information and multiple reference images, extracting content features corresponding to content information and reference style features in the reference image, fusing the reference style features of multiple reference images, self-attention encoding is performed, and the target image is generated so that its content matches the content indicated by the content information, and the style is the same as the specified style.

Benefits of technology

The solution of generating images with zero sample style transfer is realized, and images that match the specified style and content can be generated, which improves the flexibility and universality of the image generation method and ensures the accuracy of the generated image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119417931B_ABST
    Figure CN119417931B_ABST
Patent Text Reader

Abstract

The present application discloses an image generation method, apparatus, computer device, and storage medium, belonging to the field of image generation. The method includes: obtaining content information and a plurality of reference images, where the plurality of reference images have the same reference style; extracting content features corresponding to the content information, and extracting reference style features respectively corresponding to each of the plurality of reference images; fusing the reference style features respectively corresponding to the plurality of reference images to obtain a first fused style feature; performing a self-attention encoding operation on the first fused style feature to obtain a second fused style feature; guiding by the reference style characterized by the second fused style feature, and generating a target image based on the content features, where the content of the target image matches the content indicated by the content information. By providing a plurality of reference images of any style, an image of this style can be generated, improving the flexibility and versatility of the image generation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of image generation, and particularly to an image generation method, apparatus, computer device, and storage medium. Background Art

[0002] With the continuous development of computer technology, image processing technology has become increasingly rich. In some image processing scenarios, it is necessary to generate images in a specified style.

[0003] In the related art, if you want to generate an image in a specified style, a large number of sample images of this specified style are used to train a model for generating images of this specified style.

[0004] In this method, only the trained model can be used to generate images of this specified style, and images of other styles cannot be generated, resulting in great limitations in image generation. Summary of the Invention

[0005] Embodiments of the present application provide an image generation method, apparatus, computer device, and storage medium, which can improve the efficiency of image generation. The technical solutions include the following aspects.

[0006] On the one hand, an image generation method is provided, and the method includes:

[0007] Obtain content information and a plurality of reference images, the plurality of reference images having the same reference style, and the content information being used to indicate the content of the image to be generated;

[0008] Extract the content features corresponding to the content information, and extract the reference style features respectively corresponding to each of the plurality of reference images;

[0009] Fuse the reference style features respectively corresponding to the plurality of reference images to obtain a first fused style feature;

[0010] Perform a self-attention encoding operation on the first fused style feature to obtain a second fused style feature, and the second fused style feature is used to characterize the reference style;

[0011] Guided by the reference style characterized by the second fused style feature, based on the content features, generate a target image, and the content of the target image matches the content indicated by the content information.

[0012] On the other hand, an image generation apparatus is provided, and the apparatus includes:

[0013] An obtaining module, configured to obtain content information and a plurality of reference images, the plurality of reference images having the same reference style, and the content information being used to indicate the content of the image to be generated;

[0014] A feature extraction module, configured to extract content features corresponding to the content information, and extract reference style features respectively corresponding to each of the multiple reference images;

[0015] A feature fusion module, configured to fuse the reference style features respectively corresponding to the multiple reference images to obtain a first fused style feature;

[0016] A self-attention encoding module, configured to perform a self-attention encoding operation on the first fused style feature to obtain a second fused style feature, where the second fused style feature is used to represent the reference style;

[0017] An image generation module, configured to generate a target image based on the content features, guided by the reference style represented by the second fused style feature, where the content of the target image matches the content indicated by the content information.

[0018] Optionally, the first image generation model includes an encoding network and a first downsampling network; the feature extraction module includes:

[0019] An encoding unit, configured to perform an encoding operation on each reference image through the encoding network to obtain image encoding features respectively corresponding to each reference image;

[0020] A downsampling unit, configured to perform a downsampling operation on the image encoding features respectively corresponding to each reference image through the first downsampling network to obtain reference style features respectively corresponding to each reference image.

[0021] Optionally, the first downsampling network includes a first residual layer and a feature encoding layer; the downsampling unit is configured to:

[0022] Map the image encoding features respectively corresponding to each reference image to first residual features respectively corresponding to each reference image through the first residual layer;

[0023] Perform an attention encoding operation on the first residual features respectively corresponding to each reference image through the feature encoding layer to obtain reference style features respectively corresponding to each reference image.

[0024] Optionally, the feature encoding layer includes a first self-attention layer and a second self-attention layer; the downsampling unit is configured to:

[0025] Perform a self-attention encoding operation on the first residual features respectively corresponding to each reference image through the first self-attention layer to obtain first intermediate features respectively corresponding to each reference image;

[0026] Through the second self-attention layer, perform a self-attention encoding operation on the first intermediate features respectively corresponding to each of the reference images to obtain the reference style features respectively corresponding to each of the reference images.

[0027] Optionally, the first downsampling network includes k first downsampling layers connected in sequence, where k is an integer greater than 1; the downsampling unit is used for:

[0028] Through the first first downsampling layer among the k first downsampling layers, perform a downsampling operation on the image encoding features respectively corresponding to each of the reference images to obtain the first reference style features respectively corresponding to each of the reference images;

[0029] Through the m-th first downsampling layer among the k first downsampling layers, perform a downsampling operation on the (m - 1)-th first reference style features respectively corresponding to each of the reference images to obtain the m-th reference style features respectively corresponding to each of the reference images, where m is an integer greater than 1 and not greater than k.

[0030] Optionally, the first image generation model includes a feature processing network, and the feature processing network includes k feature fusion layers; the feature fusion module is used for:

[0031] Through the first feature fusion layer among the k feature fusion layers, fuse the first reference style features respectively corresponding to the multiple reference images to obtain the first first fusion style feature;

[0032] Through the m-th feature fusion layer among the k feature fusion layers, fuse the m-th reference style features respectively corresponding to the multiple reference images to obtain the m-th first fusion style feature.

[0033] Optionally, the feature processing network further includes k third self-attention layers; the self-attention encoding module is used for:

[0034] Through the first third self-attention layer among the k third self-attention layers, perform a self-attention encoding operation on the first first fusion style feature to obtain the first second fusion style feature;

[0035] Through the m-th third self-attention layer among the k third self-attention layers, perform a self-attention encoding operation on the m-th first fusion style feature to obtain the m-th second fusion style feature.

[0036] Optionally, the first image generation model includes a decoding network and a first upsampling network; the image generation module includes:

[0037] An upsampling unit, configured to perform an upsampling operation on the content feature and the second fused style feature through the first upsampling network to obtain a target image feature;

[0038] A decoding unit, configured to perform a decoding operation on the target image feature through the decoding network to obtain the target image.

[0039] Optionally, the first upsampling network includes a second residual layer and an attention layer; the upsampling unit is configured to:

[0040] Map a first noise feature, which is used to represent random noise, to a second residual feature through the second residual layer;

[0041] Perform an attention encoding operation on the content feature, the second fused style feature, and the second residual feature through the attention layer to obtain the target image feature.

[0042] Optionally, the attention layer includes a fourth self-attention layer, a first cross-attention layer, and a second cross-attention layer; the upsampling unit is configured to:

[0043] Perform a self-attention encoding operation on the second residual feature through the fourth self-attention layer to obtain a third residual feature;

[0044] Perform a cross-attention encoding operation on the third residual feature and the second fused style feature through the first cross-attention layer to obtain a second intermediate feature;

[0045] Perform a cross-attention encoding operation on the content feature and the second intermediate feature through the second cross-attention layer to obtain the target image feature.

[0046] Optionally, the first image generation model further includes a first downsampling network, the first downsampling network includes k first downsampling layers connected in sequence, the first upsampling network includes k first upsampling layers connected in sequence, the number of the second fused style features is k, and k is an integer greater than 1; the upsampling unit is configured to:

[0047] Perform an upsampling operation on a first noise feature, the content feature, and the first second fused style feature through the first first upsampling layer among the k first upsampling layers to obtain a first target image feature, where the first second fused style feature is determined based on the feature output by the first first downsampling layer, and the first noise feature is used to represent random noise;

[0048] Through the m-th first upsampling layer among the k first upsampling layers, perform an upsampling operation on the (m - 1)-th target image feature, the content feature, and the m-th second fusion style feature to obtain the m-th target image feature, where the m-th second fusion style feature is determined based on the feature output by the m-th first downsampling layer, and m is an integer greater than 1 and not greater than k.

[0049] Optionally, the decoding unit is configured to:

[0050] Perform a decoding operation on the k-th target image feature through the decoding network to obtain the target image.

[0051] Optionally, the apparatus further includes a model training module, configured to:

[0052] Obtain sample content information and a sample label image, where the content of the sample label image matches the content indicated by the sample content information;

[0053] Extract a sample content feature corresponding to the sample content information through a second image generation model, extract a first sample reference feature corresponding to the sample label image, and add a second noise feature to the first sample reference feature to obtain a second sample reference feature, where the second noise feature is used to represent random noise;

[0054] Generate a sample target image based on the sample content feature with the second sample reference feature as a guide through the second image generation model;

[0055] Train the second image generation model based on the sample target image and the sample label image to obtain the trained second image generation model;

[0056] The apparatus further includes a model construction module, configured to:

[0057] Construct the first image generation model based on the trained second image generation model.

[0058] Optionally, the trained first image generation model includes an encoding network, a second downsampling network, a second upsampling network, and a decoding network connected in sequence. The second downsampling network includes a first residual layer, a first self-attention layer, and a third cross-attention layer connected in sequence. The second upsampling network includes a second residual layer, a fourth self-attention layer, and a first cross-attention layer connected in sequence;

[0059] The model construction module is configured to:

[0060] Copy the first self-attention layer to obtain a second self-attention layer; copy the first cross-attention layer to obtain a second cross-attention layer;

[0061] Duplicate the fourth self-attention layer to obtain the third self-attention layer; construct a feature processing network based on the feature fusion layer and the third self-attention layer;

[0062] In the trained second image generation model, replace the third cross-attention layer of the second downsampling network with the second self-attention layer to obtain a first downsampling network, add the second cross-attention layer after the first cross-attention layer of the second upsampling network to obtain a first upsampling network, and add the feature processing network between the first downsampling network and the first upsampling network to obtain the first image generation model.

[0063] On the other hand, a computer device is provided. The computer device includes a processor and a memory. At least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor to implement the operations performed by the image generation method described in the above aspect.

[0064] On the other hand, a computer-readable storage medium is provided. At least one computer program is stored in the computer-readable storage medium. The at least one computer program is loaded and executed by a processor to implement the operations performed by the image generation method described in the above aspect.

[0065] On the other hand, a computer program product is provided, including a computer program. The computer program is loaded and executed by a processor to implement the operations performed by the image generation method described in the above aspect.

[0066] The solution provided in the embodiments of the present application, when an image of a specified style needs to be generated, provides multiple reference images of the specified style. First, the style features of multiple different reference images are separately extracted in parallel, and then the style features of the multiple different reference images are fused and self-attention encoded, so as to capture the common and key style features in the multiple different reference images, avoid being affected by the semantic features of the reference images themselves, and ensure the global relevance of the extracted style features. Furthermore, the target image is generated by using the content features of the content information and the extracted style features, so that it can be ensured that the content of the target image is the same as the content indicated by the content information, and the style of the target image is the same as the specified style. Therefore, the present application realizes the zero-shot style transfer image generation solution. Only by providing multiple reference images of any specified style can an image of this specified style be generated, which is not limited to only generating an image of a specific style, improving the flexibility and versatility of the image generation method. And, guided by the style features extracted from multiple reference images, the accuracy of the generated image can be ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0068] Figure 1 is a schematic diagram of a computer system provided by an embodiment of the present application;

[0069] Figure 2 is a flowchart of an image generation method provided by an embodiment of the present application;

[0070] Figure 3 is a schematic structural diagram of a first image generation model provided by an embodiment of the present application;

[0071] Figure 4 is a flowchart of another image generation method provided by an embodiment of the present application;

[0072] Figure 5 is a flowchart of a reference style feature determination method provided by an embodiment of the present application;

[0073] Figure 6 is a schematic structural diagram of a downsampling network provided by an embodiment of the present application;

[0074] Figure 7 is a schematic structural diagram of an upsampling network and a feature processing network provided by an embodiment of the present application;

[0075] Figure 8 is a flowchart of a target image feature determination method provided by an embodiment of the present application;

[0076] Figure 9 is a comparison diagram of a reference image and a target image provided by an embodiment of the present application;

[0077] Figure 10 is a flowchart of yet another image generation method provided by an embodiment of the present application;

[0078] Figure 11 is a schematic structural diagram of another first image generation model provided by an embodiment of the present application;

[0079] Figure 12 is a flowchart of a model training method provided by an embodiment of the present application;

[0080] Figure 13 is a schematic structural diagram of a second image generation model provided by an embodiment of the present application;

[0081] Figure 14It is a schematic structural diagram of a third image generation model provided by an embodiment of the present application;

[0082] Figure 15 It is a flowchart of a method for constructing a first image generation model provided by an embodiment of the present application;

[0083] Figure 16 It is a flowchart of another image generation method provided by an embodiment of the present application;

[0084] Figure 17 It is a schematic structural diagram of an image generation device provided by an embodiment of the present application;

[0085] Figure 18 It is a schematic structural diagram of another image generation device provided by an embodiment of the present application;

[0086] Figure 19 It is a schematic structural diagram of a terminal provided by an embodiment of the present application;

[0087] Figure 20 It is a schematic structural diagram of a server provided by an embodiment of the present application. Specific Embodiments

[0088] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0089] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe various concepts, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the present application, the first fusion style feature may be referred to as the second fusion style feature, and similarly, the second fusion style feature may be referred to as the first fusion style feature.

[0090] Among them, at least one means one or more than one. For example, at least one reference image may be one reference image, two reference images, three reference images, etc., any integer greater than or equal to one. A plurality means two or more than two. For example, a plurality of reference images may be two reference images, three reference images, etc., any integer greater than or equal to two. Each means each one of at least one. For example, each reference image refers to each reference image among a plurality of reference images. If there are 3 reference images in a plurality of reference images, then each reference image refers to each of the 3 reference images.

[0091] It should be noted that the information involved in this application (including but not limited to user device information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.), and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) are all fully authorized by the user or relevant parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions. For example, the content information and reference images involved in this application are all fully authorized by the user or relevant parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0092] For ease of understanding, the concepts involved in the embodiments of this application are explained below.

[0093] Stable Diffusion: It is a variant of the diffusion model called "Latent Diffusion Model (LDM)". The purpose of the diffusion model is to eliminate the Gaussian noise continuously applied to the training images, and it can be regarded as a series of denoising autoencoders. Stable Diffusion consists of three parts: VAE (Variational Auto-Encoder), U-Net (U-shaped network), and text encoder. The VAE is used to transform the image into a low-dimensional latent space, and the process of adding and removing Gaussian noise is applied to the low-dimensional latent space, and then the denoised features are decoded into the pixel space. In the forward diffusion process, Gaussian noise is iteratively applied to the compressed latent features. Each denoising step is completed by a U-Net architecture containing a residual neural network, and the latent features are obtained by denoising from the forward diffusion to the reverse direction. Finally, the VAE decoder generates an image by transforming the latent features back into the pixel space.

[0094] U-Net: It is one of the algorithms for semantic segmentation using fully convolutional networks, and it is processed using a symmetric U-shaped structure that includes a contracting path and an expanding path. At the same time, U-Net is also the basic backbone model of the Stable Diffusion model. Self-attention and cross-attention modules are added to U-Net, enabling the construction of the model's noise reduction function. U-Net is further divided into down blocks and up blocks. In fact, the U-Net in the Stable Diffusion model is divided into three parts: down block, mid block, and up block. Each block in these three parts has the same composition, but only the trends of the feature sizes and feature dimensions are different, thus distinguishing the down block, mid block, and up block.

[0095] Clip (Contrastive Language-Image Pre-Training) model: This is a deep learning model that can understand text and images. It is a multimodal model that can match text and images and learn to recognize the content in images and the text describing the images. The main goal is to learn to match images and text through contrastive learning. During training, the Clip model learns to encode images and text into a unified vector space, which enables it to understand the relationship between them in terms of language and vision. In this way, the Clip model can identify elements such as objects, scenes, and actions in images, and at the same time, it can also understand the text related to the images, such as labels, descriptions, and captions.

[0096] latent: In the Stable Diffusion model, "latent" usually refers to latent features, that is, features that are not directly observed in the model, which are the most basic features in the entire model processing flow. These latent features are very important for explaining the observed data and the behavior of the model. In the Stable Diffusion model, latent features may represent unobserved factors, such as an individual's characteristics, behaviors, styles, and compositions, which may affect the diffusion process.

[0097] Cross attention: "Cross attention" in the Stable Diffusion model refers to the cross-attention mechanism in the model. This cross-attention mechanism is typically used to process the associations between multiple input sequences in order to capture important information interactions between different sequences in the model. In the Stable Diffusion model, the cross-attention mechanism helps the model better understand the associations between different parts, thereby more accurately predicting the changes and impacts during the diffusion process.

[0098] Self attention: "Self attention" in the Stable Diffusion model refers to the self-attention mechanism in the model, which is an important supplementary part in the U-Net of the Stable Diffusion model. The input of its calculation mechanism is the backbone feature latent of the Stable Diffusion model. The role of self attention is to enable further fusion of all semantic representations in the latent features, and at the same time make the performance of the generated object clearer, enhancing the latent noise reduction ability of the U-Net.

[0099] Zero-shot: It refers to the zero-shot learning mode, which allows the model to perform on data of a certain category during the inference stage even though it has not seen data of this category during the training stage. In the embodiments of this application, the ability of zero-shot means that under the solution of this application, the model can achieve good performance in the style transfer function when facing unknown style images.

[0100] The image generation method provided by the embodiments of this application can be used in computer devices. Optionally, the computer device is a terminal or a server. The embodiments of this application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc.

[0101] Figure 1 It is a schematic diagram of a computer system provided by the embodiments of this application. Refer to Figure 1 , this computer system includes: a first terminal 101 and a server 102. The first terminal 101 and the server 102 are connected through a wireless or wired network.

[0102] The first terminal 101 installs and runs a client 111, which can be of various types such as a social application client, an online payment client, an online shopping client, a game client, a medical service client, an app store client, a video client, etc. When the first terminal 101 runs the client 111, the user interface of the client 111 is displayed on the screen of the first terminal 101. The first terminal 101 is the terminal used by the user 121.

[0103] Optionally, the first terminal 101 can generally refer to one of multiple terminals. The first terminal 101 includes: smartphones, tablets, laptops, desktop computers, smart speakers, smart watches, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircrafts, VR (Virtual Reality) devices, AR (Augmented Reality) devices, etc., but is not limited thereto.

[0104] Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminals can be only one, or the above terminals can be 6 or 8 or more in number. The embodiments of the present application do not limit the number and device types of the terminals.

[0105] Figure 1 Only one terminal is shown in the figure, but in different embodiments, there are multiple other second terminals 103 that can access the server 102. Optionally, there is also one or more second terminals 103 that are the terminals corresponding to developers. A client development and editing platform is installed on the second terminals 103. The developers can edit and update the client on the second terminals 103, and transmit the updated client installation package to the server 102 through a wired or wireless network. The first terminal 101 can download the client installation package from the server 102 to update the client.

[0106] The first terminal 101 and other second terminals 103 are connected to the server 102 through a wired network or a wireless network.

[0107] The server 102 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. The server 102 is used to provide background services for the client. Optionally, the server 102 undertakes the main computing work, and the first terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the first terminal 101 undertakes the main computing work; or, a distributed computing architecture is adopted between the server 102 and the first terminal 101 for collaborative computing.

[0108] In an embodiment of the present application, the first terminal 101 sends an image generation request to the server 102. The image generation request carries content information and multiple reference images. After receiving the image generation request, the server 102 uses the method provided in the embodiment of the present application to generate a target image based on the content information and the multiple reference images. The target image has the same content as the content information, and the target image has the same style as the multiple reference images. The server 102 returns the target image to the first terminal 101.

[0109] It should be noted that the above implementation environment is only an example. The method provided in the embodiment of the present application can also be executed independently by the first terminal 101 or the server 102, or jointly executed by the first terminal 101 and the server 102, or executed by other computer devices. The embodiment of the present application does not limit this.

[0110] Figure 2 is a flowchart of an image generation method provided in an embodiment of the present application. The embodiment of the present application is executed by a computer device. See Figure 2 and the method includes the following steps.

[0111] 201. The computer device obtains content information and multiple reference images. The multiple reference images have the same reference style, and the content information is used to indicate the content of the image to be generated.

[0112] The computer device obtains content information, which is used to indicate the content of the image to be generated. The content information can be any type of information. For example, the content information can be a content image, and the image to be generated needs to include the content in the content image. For another example, the content information can be a descriptive text, and the image to be generated needs to include the content described by the descriptive text. For still another example, the content information can be a descriptive voice, and the image to be generated needs to include the content described by the descriptive voice.

[0113] The computer device obtains multiple reference images, where the multiple reference images are different images, but the multiple reference images have the same reference style, that is, the multiple reference images correspond to the same style label. The style label of the reference image is used to represent the reference style of the reference image. For example, the reference styles of the multiple reference images are all anime style, pixel style, stick figure style, picture book style, etc.

[0114] Among them, the multiple reference images refer to at least two reference images, such as 3 reference images, 5 reference images or more reference images. The embodiment of the present application does not limit the number of reference images.

[0115] 202. The computer device extracts the content features corresponding to the content information, and extracts the reference style features corresponding to each of the multiple reference images respectively.

[0116] For each of the multiple reference images, the computer device extracts features from the reference image to obtain the reference style features corresponding to the reference image, thereby obtaining the reference style features corresponding to each reference image respectively. It should be noted that although the multiple reference images have the same reference style, since it is difficult to directly and accurately extract the features in the reference image that are only related to the reference style, the extracted reference style features will include some features that are not related to the reference style. Therefore, the reference style features of the multiple reference images are not necessarily exactly the same.

[0117] The computer device extracts the content features corresponding to the content information, and the content features are used to represent the semantics of the content information and can reflect the content of the content information.

[0118] 203. The computer device fuses the reference style features corresponding to each of the multiple reference images to obtain the first fused style feature.

[0119] After the computer device obtains the reference style features corresponding to each of the multiple reference images, it fuses the reference style features corresponding to each of the multiple reference images to obtain the first fused style feature. Among them, the first fused style feature contains the style features of the multiple reference images and is a global expression of the reference style of the multiple reference images.

[0120] 204. The computer device performs a self-attention encoding operation on the first fused style feature to obtain a second fused style feature, and the second fused style feature is used to characterize the reference style.

[0121] After the computer device obtains the first fused style feature, it uses the self-attention mechanism to perform a self-attention encoding operation on the first fused style feature to obtain the second fused style feature, thereby capturing the key parts in the first fused style feature, reducing the influence of the features that are not related to the style, and making the obtained second fused style feature pay more attention to the parts related to the style. Therefore, the second fused style feature can more accurately characterize the reference style.

[0122] In the embodiments of the present application, by fusing the style features of multiple different reference images with the same style, the style features of multiple different reference images can be balanced, the common style features can be extracted from multiple different reference images, and then the extracted style features are further purified through self-attention encoding, so as to accurately extract the style features of multiple different reference images, and improve the expression ability of the second fused style feature in terms of image style.

[0123] 205. The computer device generates a target image based on the content features, guided by the reference style characterized by the second fused style feature, and the content of the target image matches the content indicated by the content information.

[0124] After obtaining the content feature and the second fusion style feature, the computer device generates a target image guided by the second fusion style feature in the style dimension and based on the content feature in the content dimension.

[0125] Since the content feature can reflect the content of the content information and guides the image generation process in the content dimension through the content feature, it can ensure that the content of the generated target image matches the content indicated by the content information. For example, if the content information is a running child, then the target image includes a running child. Since the second fusion style feature can reflect the reference styles of multiple reference images and guides the image generation process in the style dimension through the second fusion style feature, it can ensure that the generated target image has the same reference style as the multiple reference images. For example, if the styles of the multiple reference images are anime styles, then the style of the target image is an anime style. Therefore, under the joint action of the content feature and the second fusion style feature, it can be ensured that both the content and style of the generated target image match the expectations.

[0126] It should be noted that the content information used in the embodiments of the present application can include any content, and the reference styles of multiple reference images can be any styles. Therefore, by using the method of the embodiments of the present application, images of different styles can be flexibly generated. Only multiple reference images of the specified style need to be provided, and there is no need to train different models for different styles, realizing a zero-shot style transfer image generation process.

[0127] The method provided by the embodiments of the present application, when an image of a specified style needs to be generated, provides multiple reference images of the specified style. First, the style features of multiple different reference images are respectively extracted in parallel, and then the style features of multiple different reference images are fused and self-attention encoded to capture the common and key style features in multiple different reference images, avoiding being affected by the semantic features of the reference images themselves and ensuring the global relevance of the extracted style features. Furthermore, by using the content feature of the content information and the extracted style feature to generate the target image, it can be ensured that the content of the target image is the same as the content indicated by the content information, and the style of the target image is the same as the specified style. Therefore, the present application realizes a zero-shot style transfer image generation solution. Only multiple reference images of any specified style need to be provided to generate an image of this specified style, not limited to only generating an image of a specific style, improving the flexibility and versatility of the image generation method. And, guided by the style features extracted from multiple reference images, the accuracy of the generated image can be ensured.

[0128] The above Figure 2The embodiments are only a brief description of the image generation method. In some embodiments, the computer device may also execute the image generation process through the first image generation model.

[0129] Figure 3 FIG. 4 is a schematic structural diagram of a first image generation model shown in the embodiments of the present application. As Figure 3 shown, the first image generation model includes an encoding network 301, a first downsampling network 302, a feature processing network 303, a first upsampling network 304, a decoding network 305, and a feature extraction network 306. The encoding network 301 is connected to the first downsampling network 302, the first downsampling network 302 is connected to the feature processing network 303, the feature processing network 303 is connected to the first upsampling network 304, the feature extraction network 306 is connected to the first upsampling network 304, and the first upsampling network 304 is connected to the decoding network 305.

[0130] Among them, the encoding network 301 is used to extract the first latent feature 12 in a plurality of reference images 11. The first downsampling network 302 is used to downsample the first latent feature 12 (i.e., the image encoding feature) to obtain a latent feature 13 (i.e., the reference style feature). The latent features 13 of the plurality of reference images 11 can be stored in a storage stack, which is a temporary storage container rather than an actual network structure. The feature processing network 303 is used to balance the latent features 13 of the plurality of reference images 11, that is, to implement the processes of feature fusion and self-attention encoding. The feature extraction network 306 is used to extract the content feature 15 of the content information. The first upsampling network 304 is used to upsample the first noise feature 14, the content feature 15, and the latent feature output by the feature processing network 303 to obtain a second latent feature 16 (i.e., the target image feature). The decoding network 305 is used to decode the second latent feature 16 to obtain the target image 17.

[0131] Exemplarily, the encoding network 301 and the decoding network 305 are the encoding network and the decoding network in the VAE model architecture. VAE consists of two main parts: an encoder and a decoder. The encoder converts the image into a low-dimensional latent feature, which will be used as the input of the next network structure (in the embodiments of the present application, it refers to U-Net). The decoder converts the latent feature back into an image. During the training process, the latent feature of the input image in the forward diffusion process is obtained by using the encoder. In the inference process, the encoder of VAE is no longer needed, and a pure noise latent feature is input to U-Net, and the decoder of VAE will convert the latent feature output by U-Net back into an image.

[0132] Exemplarily, the first downsampling network 302 and the first upsampling network 304 are the downsampling network and the upsampling network in the U-Net model architecture. The U-Net consists of two main parts, an encoder (downsampling part) and a decoder (upsampling part), both of which are composed of multiple unet blocks. Each unet block is composed of a ResNet (residual network) and an attention (attention network). The encoder compresses the high-resolution image into low-resolution features, and the decoder decodes the low-resolution features back into a high-resolution image. To prevent the U-Net from losing important information during encoder downsampling, intermediate connections, such as residual connections (skip-connect), are usually added between the ResNet in the encoder and the upsampling ResNet in the decoder. In the Stable Diffusion model, the U-Net adds a cross-attention layer (cross attention) to adjust the latent features. The cross-attention layer is added between the ResNet in the encoder and the ResNet in the decoder. At the same time, a self-attention layer (self attention) is added between the ResNet and the cross-attention layer. Therefore, in the network structure of the U-Net, each unet block is divided into three parts, namely ResNet, self-attention layer, and cross-attention layer.

[0133] Exemplarily, the feature extraction network 306 can be an image encoder or a text encoder. Taking the text encoder as an example, the role of the text encoder is to convert the input text into embedding features that the U-Net can understand. The text encoder can be the text encoder module in the Clip model. The text encoder module is an encoder based on a transformer, which is used to map the text into latent text embedding features.

[0134] In a possible implementation, as Figure 3 shown, the first downsampling network 302 includes k first downsampling layers connected in sequence, and the first upsampling network 304 includes k first upsampling layers connected in sequence. Among them, the k first downsampling layers correspond to the k first upsampling layers one by one.

[0135] The image generation method provided by the embodiments of this application is implemented by the first image generation model shown above Figure 3 For the detailed process, refer to the following Figure 4 embodiments. Figure 4 is a flowchart of another image generation method provided by the embodiments of this application. The embodiments of this application are executed by a computer device. Refer to Figure 4 This method includes the following steps.

[0136] 401. The computer device obtains content information and multiple reference images. The multiple reference images have the same reference style, and the content information is used to indicate the content of the image to be generated.

[0137] The computer device obtains content information, which is used to indicate the content of the image to be generated. The content information can be any type of information. The computer device obtains multiple reference images, where the multiple reference images are different images, and the multiple reference images are used to indicate the reference style of the image to be generated. The embodiments of the present application do not limit the number of the multiple reference images.

[0138] In a possible implementation manner, the computer device displays an image generation interface, and the image generation interface includes a content upload entry and a style upload entry. Based on the content upload entry, the uploaded content information is obtained. Based on the style upload entry, the uploaded multiple reference images are obtained.

[0139] In a possible implementation manner, the computer device displays an image generation interface, and the image generation interface includes a content upload entry and a style selection area. Based on the content upload entry, the uploaded content information is obtained. Multiple style labels are displayed in the style selection area, and the multiple style labels are used to represent different styles. In response to a selection operation on any one of the style labels, multiple reference images with the style label are obtained. Among them, multiple reference images corresponding to any one of the style labels are stored in the computer device.

[0140] 402. The computer device extracts the content features corresponding to the content information through the feature extraction network of the first image generation model.

[0141] As Figure 3 shown, the first image generation model includes a feature extraction network. The computer device inputs the content information into the feature extraction network, and the feature extraction network performs a feature extraction operation on the content information and outputs the content features. The content features are used to represent the semantics of the content information and can reflect the content of the content information.

[0142] Exemplarily, the content information is a content image, and the feature extraction network is used to perform a feature extraction operation on any image. For example, the feature extraction network is an image encoder. The computer device inputs the content image into the image encoder, and the image encoder outputs the content features.

[0143] Exemplarily, the content information is a description text, and the feature extraction network is used to perform a feature extraction operation on any text. For example, the feature extraction network is a text encoder. The computer device inputs the description text into the text encoder, and the text encoder outputs the content features.

[0144] Exemplarily, the form of the content features can be a feature matrix or a feature vector, etc., and the embodiments of the present application do not limit this.

[0145] 403. The computer device performs an encoding operation on each reference image through the encoding network of the first image generation model to obtain image encoding features corresponding to each reference image respectively.

[0146] As Figure 3 shown, the first image generation model includes an encoding network. The computer device inputs each reference image in a plurality of reference images into the encoding network respectively, and the encoding network outputs image encoding features corresponding to each reference image respectively. This image encoding feature is also the first latent feature 12 output by the encoding network 301 in Figure 3 . Among them, the encoding processes of each reference image in the plurality of reference images are independent of each other and do not interfere with each other.

[0147] Exemplarily, the encoding network is an encoder in the VAE model architecture.

[0148] 404. The computer device performs a downsampling operation on the image encoding features corresponding to each reference image through the first downsampling network of the first image generation model to obtain reference style features corresponding to each reference image respectively.

[0149] As Figure 3 shown, the first image generation model includes a first downsampling network. For each reference image in a plurality of reference images, the computer device inputs the image encoding features corresponding to the reference image into the first downsampling network, and the first downsampling network outputs the reference style features corresponding to the reference image. Therefore, the computer device can obtain reference style features corresponding to each reference image respectively.

[0150] Among them, the reference style features of the reference images extracted by the first downsampling network not only contain the style semantics of the reference images, but also contain other semantics irrelevant to the style semantics. That is to say, the reference style features cannot simply reflect semantics only. Therefore, the following steps 405 and 406 are needed to purify the reference style features so that the purified features can improve the significance and prominence in the style dimension.

[0151] In this implementation manner, the encoding network extracts the latent features of the reference images, and the first downsampling network further extracts the style features on the basis of the latent features. By encoding and downsampling the reference images, the style features of the reference images can be effectively extracted, which is beneficial to improving the effect of style transfer in the image generation process.

[0152] In a possible implementation, the first downsampling network includes a first residual layer and a feature encoding layer. The first residual layer is connected to the feature encoding layer, and the output of the first residual layer is the input of the feature encoding layer. Among them, the first residual layer is used to perform a residual processing operation, and the residual processing operation is used to retain the features before the residual. The feature encoding layer is used to perform an attention encoding operation, and the attention encoding operation refers to the process of encoding input data using an attention mechanism. In this process, the model dynamically generates a set of encoding vectors according to the relevance of different parts of the input data to the current task. These encoding vectors not only contain the original information of the input data but also incorporate the model's understanding and attention to the input data.

[0153] Then, as Figure 5 shown, step 404 includes the following steps 4041 and 4042.

[0154] 4041. Through the first residual layer of the first downsampling network, map the image encoding features corresponding to each reference image to the first residual features corresponding to each reference image.

[0155] Among them, the first residual layer can be a ResNet network structure. The basic unit of ResNet is a residual block, which consists of two or more convolutional layers, and a skip connection is added between these layers. These skip connections directly add the input to the output, thus realizing the transmission of the residual.

[0156] For each reference image among multiple reference images, the computer device inputs the image encoding features corresponding to the reference image into the first residual layer, and the first residual layer outputs the first residual features.

[0157] 4042. Through the feature encoding layer of the first downsampling network, perform an attention encoding operation on the first residual features corresponding to each reference image to obtain the reference style features corresponding to each reference image.

[0158] For each reference image among multiple reference images, the computer device inputs the first residual features corresponding to the reference image into the feature encoding layer, and the feature encoding layer outputs the reference style features.

[0159] Optionally, the feature encoding layer includes a first self-attention layer and a second self-attention layer. The first self-attention layer is connected to the second self-attention layer, and the output of the first self-attention layer is the input of the second self-attention layer.

[0160] Then, step 4042 includes: performing self-attention encoding operations on the first residual features respectively corresponding to each reference image through a first self-attention layer to obtain first intermediate features respectively corresponding to each reference image; and performing self-attention encoding operations on the first intermediate features respectively corresponding to each reference image through a second self-attention layer to obtain reference style features respectively corresponding to each reference image.

[0161] Among them, the self-attention mechanism is a special attention mechanism that allows the model to directly consider the relationships between various positions of the input data when processing the input data without relying on external information. The core of this mechanism lies in determining the degree of attention between different positions by calculating the correlation (or attention weights) between each position in the input data and other positions.

[0162] Exemplarily, the network structure of the first self-attention layer is the same as that of the second self-attention layer, and the model parameters of the first self-attention layer are different from those of the second self-attention layer, or the model parameters of the first self-attention layer and the second self-attention layer can also be the same.

[0163] Exemplarily, taking the first self-attention layer as an example, the process of self-attention encoding includes: using a first mapping matrix to map the first residual features of the reference image into a first query feature, a first key feature, and a first value feature; performing a normalization operation on the product of the transpose of the first query feature, the first key feature, and a scaling factor to obtain a normalized feature, and multiplying the normalized feature by the first value feature to obtain a first intermediate feature. The key feature, value feature, and query feature in the embodiments of the present application belong to different feature spaces respectively.

[0164] Among them, the first mapping matrix is used to perform a spatial transformation operation on the first residual features. The first mapping matrix is the model parameter of the first self-attention layer and is obtained during the model training phase. The computer device multiplies the first residual features by the first mapping matrix to obtain a first spatial matrix, and splits out the first query feature, the first key feature, and the first value feature from the first spatial matrix. For example, the first mapping matrix is a 3-dimensional mapping matrix, the first spatial matrix obtained by multiplying the first residual features by the first mapping matrix is a 3-dimensional first spatial matrix, and the electronic device uses the features on each dimension in the first spatial matrix as the first query feature, the first key feature, and the first value feature respectively.

[0165] Among them, the process of performing self-attention encoding operations on the first intermediate features is the same as that of performing self-attention encoding operations on the image encoding features, and will not be elaborated here.

[0166] Figure 6It is a schematic structural diagram of a downsampling network provided by an embodiment of the present application. As shown in Figure 6 (1) in the figure, the first downsampling network 302 in the embodiment of the present application includes a plurality of first downsampling layers 31, and each first downsampling layer 31 is composed of a first residual layer, a first self-attention layer, and a second self-attention layer. Figure 6 (1) in the figure only takes the example that the first downsampling network 302 includes a plurality of first downsampling layers 31 for display. In fact, the first downsampling network 302 may also only include one first downsampling layer 31, or it can be understood that the first downsampling network 302 only includes a set of first residual layers, a first self-attention layer, and a second self-attention layer.

[0167] In this implementation manner, by introducing the residual layer, the problem of gradient disappearance in the deep network can be avoided, so as to extract deeper features. By introducing the attention encoding, it is possible to focus on the important parts of the style features, which is beneficial to extracting the key style features in the image, reducing the interference of the features irrelevant to the style in the image, and thus improving the effect of style transfer in the image generation process.

[0168] Moreover, by using the two-layer self-attention mechanism, the model can pay more attention to the key style features in the image more carefully, making the extracted style features more significant and prominent, ignoring the interference of redundant information, enhancing the detail expression ability of the image style, being beneficial to capturing the subtle differences of the style features in images of different styles, and thus improving the effect of style transfer in the image generation process.

[0169] 405. The computer device fuses the reference style features respectively corresponding to a plurality of reference images through the feature processing network of the first image generation model to obtain a first fused style feature.

[0170] As shown in Figure 3 the figure, the first image generation model includes a feature processing network. The computer device inputs the reference style features of a plurality of reference images into the feature processing network together, and the feature processing network fuses the reference style features of the plurality of reference images to obtain a first fused style feature.

[0171] In the embodiment of the present application, generating an image with a specified style mainly depends on the reference style features of a plurality of reference images with the specified style. Therefore, the first fused style features of the plurality of reference images are fused to balance the features of the plurality of reference images in the style dimension, and a first fused style feature that can represent the common style attributes of the plurality of reference images is obtained, so as to avoid bias towards the style features of a single reference image, which is beneficial to improving the accuracy of style transfer in the image generation task.

[0172] In a possible implementation, the reference style features of multiple reference images are averaged in the channel dimension to obtain a first fused style feature. The common attributes between multiple reference style features can be automatically extracted during the averaging process. Alternatively, the reference style features of multiple reference images are concatenated to obtain a first fused style feature.

[0173] Figure 7 It is a schematic structural diagram of an upsampling network and a feature processing network provided by an embodiment of the present application. As shown in (1) in Figure 7 As shown, the first image generation model in the embodiment of the present application includes a feature processing network 303. Taking the number of multiple reference images as 3 as an example, the computer device obtains the reference style feature a, the reference style feature b, and the reference style feature c of these 3 reference images. The reference style feature a, the reference style feature b, and the reference style feature c are input into the feature processing network 303, and in the feature processing network 303, the reference style feature a, the reference style feature b, and the reference style feature c are concatenated in the channel dimension, and then the average value is calculated in the channel dimension to obtain a first fused style feature.

[0174] In a possible implementation, generating an image with a specified style mainly depends on the reference style features of multiple reference images with the specified style. Once the degree of the reference style features of the multiple reference images in feature expression is different, it will cause the influence of the style features between different reference images. Therefore, when balancing the reference style features of multiple reference images, directly taking the average value may affect the final purification effect. Therefore, performing a weighted average operation on the reference style features of multiple reference images to obtain a first fused style feature can automatically weaken the role of the reference image with a poor style representation in purification, so as to more accurately transfer the style.

[0175] Among them, the weight of the reference style feature of the reference image can be determined according to the degree of style representation of the reference image. The higher the degree of style representation of the reference image, the greater the weight of the reference style feature of the reference image. The lower the degree of style representation of the reference image, the smaller the weight of the reference style feature of the reference image.

[0176] 406. The computer device performs a self-attention encoding operation on the first fused style feature through the feature processing network of the first image generation model to obtain a second fused style feature, and the second fused style feature is used to represent the reference style.

[0177] The feature processing network performs a self-attention encoding operation on the first fused style feature to obtain a second fused style feature. Among them, the process of performing the self-attention encoding operation on the first fused style feature is the same as the process of performing the self-attention encoding operation on the image encoding feature described above, and will not be elaborated here.

[0178] In the embodiment of the present application, the self-attention mechanism is adopted to perform a self-attention encoding operation on the first fused style feature to obtain a second fused style feature, so as to capture the key part in the first fused style feature, reduce the influence of features irrelevant to the style, and make the obtained second fused style feature pay more attention to the part related to the style.

[0179] In a possible implementation manner, the feature processing network includes a fourth self-attention layer. The first fused style feature is input into the fourth self-attention layer, and the second fused style feature output by the fourth self-attention layer is obtained. As Figure 7 shown, the feature processing network 303 includes a fourth self-attention layer 32.

[0180] 407. The computer device performs an upsampling operation on the content feature and the second fused style feature through the first upsampling network of the first image generation model to obtain a target image feature.

[0181] As Figure 3 shown, the first image generation model includes a first upsampling network. The computer device inputs the content feature and the second fused style feature into the first upsampling network, and the first upsampling network outputs a target image feature.

[0182] In a possible implementation manner, the first upsampling network includes a second residual layer and an attention layer. The second residual layer is connected to the attention layer, and the output of the second residual layer is the input of the attention layer. Among them, the second residual layer is used to perform a residual processing operation, and the residual processing operation is used to retain the features before the residual. The attention layer is used to perform an attention encoding operation. Attention encoding refers to the process of encoding input data using the attention mechanism. In this process, the model will dynamically generate a set of encoding vectors according to the relevance of different parts of the input data to the current task. These encoding vectors not only contain the original information of the input data, but also incorporate the model's understanding and attention degree of the input data.

[0183] Then, as Figure 8 shown, this step 407 includes the following steps 4071 and 4072.

[0184] 4071. Through the second residual layer of the first upsampling network, map the first noise feature to a second residual feature, and the first noise feature is used to represent random noise.

[0185] Among them, the second residual layer can be a ResNet network structure. The basic unit of ResNet is a residual block, which consists of two or more convolutional layers, and a skip connection is added between these layers. These skip connections directly add the input to the output, thus realizing the transmission of residuals.

[0186] The computer device obtains a preset or randomly generated first noise feature, inputs the first noise feature into the second residual layer, and the second residual layer outputs a second residual feature. Optionally, the first noise feature can be the noise feature of random Gaussian noise, etc.

[0187] 4072. Through the attention layer of the first upsampling network, perform an attention encoding operation on the content feature, the second fusion style feature, and the second residual feature to obtain a target image feature.

[0188] The computer device inputs the content feature, the second fusion style feature, and the second residual feature into the attention layer, and the attention layer outputs a target image feature.

[0189] Optionally, the attention layer includes a fourth self-attention layer, a first cross-attention layer, and a second cross-attention layer. The fourth self-attention layer is connected to the first cross-attention layer, and the output of the fourth self-attention layer is the input of the first cross-attention layer. The first cross-attention layer is connected to the second cross-attention layer, and the output of the first cross-attention layer is the input of the second cross-attention layer.

[0190] Then, step 4072 includes: through the fourth self-attention layer, perform a self-attention encoding operation on the second residual feature to obtain a third residual feature; through the first cross-attention layer, perform a cross-attention encoding operation on the third residual feature and the second fusion style feature to obtain a second intermediate feature; through the second cross-attention layer, perform a cross-attention encoding operation on the content feature and the second intermediate feature to obtain a target image feature.

[0191] Among them, the process of performing a self-attention encoding operation on the second residual feature is the same as the process of performing a self-attention encoding operation on the image encoding feature described above, and will not be elaborated here.

[0192] Among them, the cross-attention mechanism is a variant of the attention mechanism, which is used to establish an association between two different input data and calculate attention weights. This mechanism enables the model to dynamically focus on the part of one input data that is relevant to the other input data, thus effectively fusing the information of the two input data.

[0193] Exemplarily, the network structure of the first cross-attention layer is the same as that of the second cross-attention layer, and the model parameters of the first cross-attention layer are different from those of the second cross-attention layer, or the model parameters of the first cross-attention layer and the second cross-attention layer can also be the same.

[0194] Exemplarily, the process of the first cross-attention layer performing cross-attention encoding includes: using a second mapping matrix to map the third residual feature to a second query feature; using a third mapping matrix to map the second fused style feature to a second key feature and a second value feature; performing a normalization operation on the product of the transpose of the second query feature, the second key feature, and the scaling factor to obtain a normalized feature, and multiplying the normalized feature by the second value feature to obtain a second intermediate feature.

[0195] Exemplarily, the process of the second cross-attention layer performing cross-attention encoding includes: using a fourth mapping matrix to map the second intermediate feature to a third query feature; using a fifth mapping matrix to map the content feature to a third key feature and a third value feature; performing a normalization operation on the product of the transpose of the third query feature, the third key feature, and the scaling factor to obtain a normalized feature, and multiplying the normalized feature by the third value feature to obtain the target image feature.

[0196] As Figure 7 shown in (1) of , the first upsampling network 304 in the embodiment of the present application includes a plurality of first upsampling layers 33, and each first upsampling layer 33 is composed of a second residual layer, a fourth self-attention layer, a first cross-attention layer, and a second cross-attention layer. Figure 7 (1) only takes the first upsampling network 304 including a plurality of first upsampling layers 33 as an example to show. Actually, the first upsampling network 304 may also only include one first upsampling layer 33, or it can be understood that the first upsampling network 304 only includes a group of a second residual layer, a fourth self-attention layer, a first cross-attention layer, and a second cross-attention layer.

[0197] In this implementation manner, by introducing the residual layer, the problem of gradient disappearance in the deep network can be avoided, thereby increasing the expression ability of features. And by performing attention encoding on the content feature, style feature, and residual feature through the attention layer, the importance of each feature in the image generation process can be adaptively adjusted, thereby enhancing the expression ability of the target image feature and being beneficial to improving the accuracy of the generated target image.

[0198] Moreover, a cross-attention mechanism is used to perform cross-attention encoding on the style features and content features, effectively achieving a deep fusion between the content and style, enabling more refined collaborative processing of the content features and style features during the image generation process, and improving the adaptability of the image generation process based on style transfer.

[0199] 408. The computer device performs a decoding operation on the target image features through the decoding network of the first image generation model to obtain the target image.

[0200] As Figure 3 shown, the first image generation model includes a decoding network. The computer device inputs the target image features into the decoding network, and the decoding network outputs the target image. The target image features include features in the content dimension and features in the style dimension, which can guide the image generation process in the content dimension to ensure that the generated target image has the same content as the content information, and can guide the image generation process in the style dimension to ensure that the generated target image has the same style as multiple reference images.

[0201] Exemplarily, the decoding network is the decoder in the VAE model architecture.

[0202] In this implementation manner, by performing an upsampling operation on the content features and style features through the first upsampling network, the details of the image can be gradually restored during the image generation process, and the structural information of the image can be reconstructed. Moreover, since the target image features include the content features and style features, achieving the combination of content and style, performing a decoding operation on the target image features to obtain the target image can ensure that the content and style of the target image better meet the expected generation target.

[0203] The embodiment of the present application provides an image generation method, which performs a balance calculation on the style features of multiple parallel input reference images to implement a zero-shot based image style transfer scheme. By inputting an indefinite number of reference images with the same style, the style features of multiple reference images can be automatically extracted, and under the guidance of the style features, a specified content image with the same style as multiple reference images can be generated.

[0204] Figure 9 is a comparison diagram of a reference image and a target image provided by the embodiment of the present application. As Figure 9As shown, the first row is in the picture book style. The first three columns in the first row are reference images, and the fourth column is the target image generated based on the three reference images in the picture book style. The target image is in the picture book style. The second row is in the pixel style. The first three columns in the second row are reference images, and the fourth column is the target image generated based on the three reference images in the pixel style. The target image is in the pixel style. The third row is in the pixel style. The first three columns in the third row are reference images, and the fourth column is the target image generated based on the three reference images in the pixel style. The target image is in the pixel style. The fourth row is in the toy style. The first three columns in the fourth row are reference images, and the fourth column is the target image generated based on the three reference images in the toy style. The target image is in the toy style. The fifth row is in the stick figure style. The first three columns in the fifth row are reference images, and the fourth column is the target image generated based on the three reference images in the stick figure style. The target image is in the stick figure style.

[0205] The method provided by the embodiments of the present application, when an image of a specified style needs to be generated, provides multiple reference images of the specified style. First, the style features of multiple different reference images are separately extracted in parallel, and then the style features of multiple different reference images are fused and self-attention encoded, so as to capture the common and key style features in multiple different reference images, avoid being affected by the semantic features of the reference images themselves, and ensure the global relevance of the extracted style features. Furthermore, the target image is generated by using the content features of the content information and the extracted style features, so that it can be ensured that the content of the target image is the same as the content indicated by the content information, and the style of the target image is the same as the specified style. Therefore, the present application realizes the zero-shot style transfer image generation scheme. Only by providing multiple reference images of any specified style can an image of this specified style be generated, without being limited to only generating an image of a specific style, improving the flexibility and versatility of the image generation method. Moreover, under the guidance of the style features extracted from multiple reference images, the accuracy of the generated image can be ensured.

[0206] In addition, the present application can control the style of the generated target image by providing multiple reference images with the same style, without the need for additional style training of the model. The method of this invention can be mainly applied to the process of generating stylized pictures. For the style representations in all stylized pictures to be generated, they can be controlled by the input style pictures. At the same time, without additional style training, the fast style transfer function can be achieved.

[0207] Moreover, this solution constructs a processing mechanism for style transfer using multiple reference images. This processing mechanism allows an indefinite number of reference images to be received as the input of the first image generation model. This indefinite number of inputs enables a certain degree of scalability for the first image generation model. By capturing the style features of multiple reference images, the specified style can be accurately captured. The more reference images are provided, the more accurate the extracted style features will be, and the more accurate the generated target image will be.

[0208] In one embodiment, based on the above Figure 4 embodiment, as Figure 3 shown, the first downsampling network of the first image generation model includes multiple first downsampling layers, and the first upsampling network includes multiple first upsampling layers. Then, the image generation process can be referred to the following Figure 10 embodiment. Figure 10 is a flowchart of another image generation method provided by an embodiment of the present application. The embodiment of the present application is executed by a computer device. Refer to Figure 10 for details. This method includes the following steps.

[0209] 1001. The computer device obtains content information and multiple reference images. The multiple reference images have the same reference style, and the content information is used to indicate the content of the image to be generated.

[0210] The process of step 1001 for obtaining content information and multiple reference images is the same as that of step 401 for obtaining content information and multiple reference images above, and will not be elaborated here.

[0211] 1002. The computer device extracts the content features corresponding to the content information through the feature extraction network of the first image generation model.

[0212] The process of step 1002 for extracting content features is the same as that of step 402 for extracting content features above, and will not be elaborated here.

[0213] 1003. The computer device extracts the reference style features corresponding to each of the multiple reference images through the encoding network of the first image generation model.

[0214] The process of step 1003 for extracting image encoding features is the same as that of step 403 for extracting image encoding features above, and will not be elaborated here.

[0215] As Figure 3 shown, the first image generation model in the embodiment of the present application includes a first downsampling network. The first downsampling network includes k first downsampling layers connected in sequence, where k is an integer greater than 1. Then, the process of extracting reference style features through the first downsampling network is as follows in step 1004.

[0216] 1004. The computer device performs a downsampling operation on the image encoding features corresponding to each reference image through the first first downsampling layer among the k first downsampling layers, to obtain the first reference style feature corresponding to each reference image. The computer device performs a downsampling operation on the (m-1)th first reference style feature corresponding to each reference image through the mth first downsampling layer among the k first downsampling layers, to obtain the mth reference style feature corresponding to each reference image.

[0217] Where m is an integer greater than 1 and not greater than k.

[0218] Taking a certain reference image as an example, the computer device inputs the image encoding feature of the reference image into the first first downsampling layer to obtain the first reference style feature of the reference image. The computer device inputs the first reference style feature of the reference image into the second first downsampling layer to obtain the second reference style feature of the reference image. The computer device inputs the second reference style feature of the reference image into the third first downsampling layer to obtain the third reference style feature of the reference image. And so on, until the last first downsampling layer outputs the last reference style feature. That is, until the (k-1)th reference style feature of the reference image is input into the kth first downsampling layer to obtain the kth reference style feature of the reference image.

[0219] The computer device performs the above steps for each reference image, so as to obtain k reference style features of each reference image.

[0220] In this implementation manner, by introducing multiple first downsampling layers, the style features of the image can be gradually extracted at different scales, improving the diversity and richness of the style features, which is beneficial to capturing the details and global style in the image. And through layer-by-layer downsampling, the feature representation of the image is gradually compressed layer by layer, removing unimportant information and retaining the key features related to the style, improving the ability to capture style features and the processing efficiency, so as to generate more rich and fine style features in the image generation task based on style transfer.

[0221] In a possible implementation manner, each first downsampling layer includes a first residual layer and a feature encoding layer. The process of each first downsampling layer performing downsampling on the reference style feature to obtain the reference style feature is the same as the process of step 4041-step 4042 above. Exemplarily, the feature encoding layer includes a first self-attention layer and a second self-attention layer.

[0222] That is, as Figure 6 shown in (1) of, the first downsampling network 302 includes multiple first downsampling layers 31, and each first downsampling layer includes a first residual layer, a first self-attention layer and a second self-attention layer.

[0223] In a possible implementation, to make the model scalable and allow it to support an indefinite number of input reference images, a storage stack for reference style features is designed. As Figure 3 shown, the reference style feature 13 output by the first downsampling network 302 is stored in the storage stack. The first downsampling network 302 includes multiple first downsampling layers. According to this network structure, the storage stack stores the features in different stacks according to the differences of the multiple first downsampling layers. The reference style features 13 output by different first downsampling layers are respectively stored in different layers of the storage stack, which is convenient for subsequent one-to-one processing in the feature processing network and the first upsampling network according to the network layers.

[0224] Figure 11 FIG. is a schematic structural diagram of another first image generation model provided by an embodiment of the present application. As Figure 11 shown, taking the multiple reference images including reference image A, reference image B, and reference image C as an example. The first latent features 12 of the 3 reference images are respectively extracted by the encoding network 301. The first latent feature 12 is also the above-mentioned image encoding feature. The first latent features 12 of the 3 reference images are respectively input into the first downsampling network 302, and the first downsampling network 302 respectively outputs k reference style features 13 of each reference image. The reference style features 13 output by different first downsampling layers are respectively stored in different layers of the storage stack.

[0225] As Figure 3 shown, the first image generation model in the embodiment of the present application includes a feature processing network. The feature processing network includes k processing layers, and each processing layer includes a feature fusion layer and a third self-attention layer. It can be understood that the feature processing network includes k feature fusion layers and k third self-attention layers. The process of feature fusion and self-attention encoding through the feature processing network is as follows in steps 1004 and 1005.

[0226] 1005. The computer device fuses the first reference style features respectively corresponding to multiple reference images through the first feature fusion layer among the k feature fusion layers to obtain the first fused style feature of the first layer. Through the mth feature fusion layer among the k feature fusion layers, the mth reference style features respectively corresponding to multiple reference images are fused to obtain the mth first fused style feature.

[0227] Where m is an integer greater than 1 and not greater than k.

[0228] The computer device inputs the first reference style feature of multiple reference images into the first feature fusion layer to obtain the first first-fused style feature. It inputs the second reference style feature of multiple reference images into the second feature fusion layer to obtain the second first-fused style feature. It inputs the third reference style feature of multiple reference images into the third feature fusion layer to obtain the third first-fused style feature. And so on, until the last feature fusion layer outputs the last first-fused style feature. That is, until the k-th reference style feature of multiple reference images is input into the k-th feature fusion layer to obtain the k-th first-fused style feature.

[0229] In this implementation, fusing the style features of multiple reference images can more accurately capture the features of the same style corresponding to multiple reference images, and can avoid the problem of semantic interference of the image itself that may exist when using a single reference image, resulting in style deviation or distortion, and improving the accuracy of the extracted style features.

[0230] In a possible implementation, as in step 1004 above, the k reference style features of multiple reference images are respectively stored in different layers of the storage stack according to the network structure. The computer device sequentially reads the first reference style feature to the k-th reference style feature of the reference images in the storage stack according to the hierarchical structure of the storage stack.

[0231] In a possible implementation, the process of each feature fusion layer fusing the reference style features of multiple reference images is the same as that of step 405 above, and will not be elaborated here.

[0232] 1006. The computer device performs a self-attention encoding operation on the first first-fused style feature through the first third self-attention layer in the k third self-attention layers to obtain the first second-fused style feature, and performs a self-attention encoding operation on the m-th first-fused style feature through the m-th third self-attention layer in the k third self-attention layers to obtain the m-th second-fused style feature.

[0233] Where m is an integer greater than 1 and not greater than k.

[0234] The computer device inputs the first first-fused style feature into the first third self-attention layer to obtain the first second-fused style feature. It inputs the second first-fused style feature into the second third self-attention layer to obtain the second second-fused style feature. It inputs the first first-fused style feature into the first third self-attention layer to obtain the first second-fused style feature. And so on, until the last third self-attention layer outputs the last second-fused style feature. That is, until the k-th first-fused style feature is input into the k-th third self-attention layer to obtain the k-th second-fused style feature.

[0235] In this implementation manner, the self-attention mechanism is adopted to process the fused style features, which is beneficial to capturing the common and key style features in multiple different reference images, making the extracted style features more prominent and outstanding, avoiding being affected by the semantic features of the reference images themselves, and ensuring the accuracy of the extracted style features.

[0236] In a possible implementation manner, the process of each third self-attention layer performing self-attention encoding on the first fused style feature is the same as that in step 406 above, and will not be elaborated here.

[0237] In this implementation manner, feature balance purification processing based on the self-attention mechanism is added to each processing layer of the feature processing network, making the model pay more attention to the common style features in multiple reference images, which is beneficial to improving the model's ability to extract style features.

[0238] As Figure 3 shown, the first image generation model in the embodiment of the present application includes a first upsampling network, and the first upsampling network includes k first upsampling layers connected in sequence. The number of the second fused style features is k, and k is an integer greater than 1. Then the process of extracting the target image features through the first upsampling network is as follows in step 1007.

[0239] 1007. The computer device performs an upsampling operation on the first noise feature, the content feature, and the first second fused style feature through the first first upsampling layer among the k first upsampling layers to obtain the first target image feature, and performs an upsampling operation on the (m - 1)-th target image feature, the content feature, and the m-th second fused style feature through the m-th first upsampling layer among the k first upsampling layers to obtain the m-th target image feature.

[0240] Where m is an integer greater than 1 and not greater than k. The first second fused style feature is determined based on the feature output by the first first downsampling layer, and the m-th second fused style feature is determined based on the feature output by the m-th first downsampling layer. It can be understood that the k first downsampling layers and the k first upsampling layers are in one-to-one correspondence, and the feature extracted based on the first first downsampling layer is the input of the first first upsampling layer, and the feature extracted based on the m-th first downsampling layer is the input of the m-th first upsampling layer.

[0241] The computer device inputs the first noise feature, content feature, and the first second fusion style feature into the first first upsampling layer to obtain the first target image feature. The first target image feature, content feature, and the second second fusion style feature are input into the second first upsampling layer to obtain the second target image feature. The second target image feature, content feature, and the third second fusion style feature are input into the third first upsampling layer to obtain the third target image feature. And so on, until the last first upsampling layer outputs the last target image feature. That is, until the (k - 1)th target image feature, content feature, and the kth second fusion style feature are input into the kth first upsampling layer to obtain the kth target image feature.

[0242] In this implementation manner, introducing multiple first upsampling layers can restore the details of the image layer by layer, reconstruct the structural information of the image, and gradually restore the low-resolution features to high-resolution image features, which helps to improve the detail quality of the generated image. Moreover, each first upsampling layer can perform upsampling by combining the content feature with the corresponding second fusion style feature, which can ensure the gradual fusion of the content feature and the style feature at different scales and further improve the accuracy of the generated image.

[0243] In a possible implementation manner, each first upsampling layer includes a second residual layer and an attention layer. The process of each first downsampling layer extracting the target image feature is the same as the process of steps 4071 - 4072 above. Exemplarily, the attention layer includes a fourth self-attention layer, a first cross-attention layer, and a second cross-attention layer.

[0244] That is, as Figure 7 shown in (1) of, the first upsampling network 304 includes multiple first upsampling layers 33, and each first upsampling layer 33 includes a second residual layer, a fourth self-attention layer, a first cross-attention layer, and a second cross-attention layer.

[0245] 1008. The computer device performs a decoding operation on the kth target image feature through the decoding network of the first image generation model to obtain the target image.

[0246] The computer device inputs the kth target image feature into the decoding network, and the decoding network outputs the target image.

[0247] The method provided by the embodiments of the present application provides a mechanism for extracting multi-layer style features from multiple reference images. In multiple first downsampling layers and multiple first upsampling layers in the U-Net model architecture, the latent features of each layer are sequentially extracted. Through multi-layer parallel feature extraction, the model can capture more accurate style features at multiple levels, which is beneficial to improving the accuracy of the generated target image.

[0248] In one embodiment, for the construction process of the above first image generation model, refer to the following Figure 12 embodiment. Figure 12 FIG. is a flowchart of a model training method provided by an embodiment of the present application, which is executed by a computer device. Refer to Figure 12 and this method includes the following steps.

[0249] 1201. The computer device obtains sample content information and a sample label image, and the content of the sample label image matches the content indicated by the sample content information.

[0250] This sample content information is the same as the content information in the above Figure 4 embodiment and will not be elaborated here.

[0251] The content of the sample label image matches the content indicated by the sample content information. For example, if the sample content information is a running child, then the sample label image includes a running child. Among them, the sample label image can be an image of any style, and the present application embodiment does not limit the style of the sample label image.

[0252] 1202. The computer device extracts the sample content features corresponding to the sample content information through a second image generation model, extracts the first sample reference features corresponding to the sample label image, adds a second noise feature to the first sample reference features to obtain second sample reference features, and the second noise feature is used to represent random noise.

[0253] Figure 13 FIG. is a schematic structural diagram of a second image generation model provided by an embodiment of the present application. As Figure 13 shown, the second image generation model includes an encoding network 301, a feature extraction network 306, a second downsampling network 307, a connection network 308, a second upsampling network 309, and a decoding network 305.

[0254] Among them, the computer device inputs the sample content information into the feature extraction network 306 to obtain sample content features. The sample label image 41 and the second noise feature are input into the encoding network 301, and the encoding network 301 extracts the first sample reference features of the sample label image and adds the second noise feature to the first sample reference features to obtain the second sample reference features 42.

[0255] Among them, the encoding network and the feature extraction network of this second image generation model are the same as those of the first image generation model and will not be elaborated here.

[0256] 1203. The computer device generates a sample target image based on the sample content features under the guidance of the second sample reference features through the second image generation model.

[0257] As shown Figure 13 in the figure, the computer device inputs the sample content feature and the second sample reference feature 42 into the second downsampling network 307. The latent feature output by the second downsampling network 307 is input into the connection network 308, and the latent feature output by the connection network 308 is input into the second upsampling network 309. The second upsampling network 309 outputs the sample target image feature 43. The computer device inputs the sample target image feature 43 into the decoding network 305, and the decoding network 305 decodes the sample target image feature 43 to output the sample target image 44.

[0258] Among them, the second sample reference feature 42 is a feature after adding noise. The second sample reference feature 42 is processed through the second downsampling network 307, the connection network 308, and the second upsampling network 309 to perform noise reduction, so as to obtain a clean sample target image feature 43. The sample target image feature 43 contains both the semantics of the sample content information and the semantics of the sample label image.

[0259] 1204. The computer device trains the second image generation model based on the sample target image and the sample label image to obtain the trained second image generation model.

[0260] The sample label image is the real image indicated by the sample content information, while the sample target image is obtained based on the second image generation model. The smaller the difference between the sample target image generated by the second image generation model and the sample label image, the more accurate the second image generation model. Therefore, the computer device adjusts the model parameters of the second image generation model based on the sample target image and the sample label image to obtain a more accurate second image generation model. Among them, the training objective is to reduce the difference between the sample target image and the sample label image, so as to improve the image generation ability of the second image generation model.

[0261] In a possible implementation manner, the encoding network in the second image generation model is removed to obtain a third image generation model, which is used to generate a corresponding image based on any content information.

[0262] Among them, in the model training stage, it is necessary to use the encoding network to encode the sample label image, while in the model inference stage, there is no sample label image, and it is not necessary to encode the sample label image. Therefore, the encoding network is no longer needed.

[0263] Figure 14 is a schematic structural diagram of a third image generation model provided by an embodiment of the present application. As shown Figure 14As shown in the figure, the third image generation model includes a feature extraction network 306, a second downsampling network 307, a connection network 308, a second upsampling network 309, and a decoding network 305. In the task of generating an image that matches the content information, the computer device inputs the content information into the feature extraction network 306 to obtain content features. The computer device inputs the third noise feature 81 into the second downsampling network 307, and inputs the content features into the second downsampling network 307, the connection network 308, and the second upsampling network 309. After being processed by the second downsampling network 307, the connection network 308, and the second upsampling network 309, a denoised image feature 82 is obtained. The image feature 82 is input into the decoding network 305 to obtain a content image 83 that matches the content information.

[0264] 1205. The computer device constructs a first image generation model based on the trained second image generation model.

[0265] After obtaining the trained second image generation model, the computer device adjusts the model structure of the second image generation model to obtain a first image generation model suitable for style transfer.

[0266] In a possible implementation, the trained first image generation model includes an encoding network, a feature extraction network, a second downsampling network, a second upsampling network, and a decoding network connected in sequence. Among them, the second downsampling network includes a first residual layer, a first self-attention layer, and a third cross-attention layer connected in sequence, and the second upsampling network includes a second residual layer, a fourth self-attention layer, and a first cross-attention layer connected in sequence. Then, as Figure 15 shown, this step 1205 includes the following steps 1215 to 1245.

[0267] 1215. Duplicate the first self-attention layer to obtain a second self-attention layer.

[0268] 1225. Duplicate the first cross-attention layer to obtain a second cross-attention layer.

[0269] 1235. Duplicate the fourth self-attention layer to obtain a third self-attention layer, and construct a feature processing network based on the feature fusion layer and the third self-attention layer.

[0270] Among them, the feature fusion layer is used to perform an average operation or a weighted average operation. The feature fusion layer itself does not have model parameters and is only used to perform simple calculation operations, so it does not need to be trained. The computer device connects the third self-attention layer behind the feature fusion layer to obtain a feature processing network.

[0271] 1245. In the trained second image generation model, replace the third cross-attention layer of the second downsampling network with a second self-attention layer to obtain a first downsampling network, add a second cross-attention layer after the first cross-attention layer of the second upsampling network to obtain a first upsampling network, and add a feature processing network between the first downsampling network and the first upsampling network to obtain a first image generation model.

[0272] Among them, the computer device replaces the third cross-attention layer in the second downsampling network with a second self-attention layer to obtain a first downsampling network, and connects a second cross-attention layer behind the first cross-attention layer of the second upsampling network to obtain a first upsampling network. Add a feature processing network between the first downsampling network and the first upsampling network. The output of the first downsampling network is the input of the feature processing network, and the output of the feature processing network is the input of the first upsampling network. By adjusting the model structure of the second image generation model, the first image generation model can be obtained without retraining the first image generation model.

[0273] Exemplarily, as Figure 6 shown in (2) of, the second downsampling network 307 of the second image generation model includes k second downsampling layers 61, and each second downsampling layer 61 includes a first residual layer, a first self-attention layer, and a third cross-attention layer. For each second downsampling layer 61, copy the first self-attention layer of the second downsampling layer 61 to obtain a second self-attention layer, and replace the third cross-attention layer of the second downsampling layer 61 with the second self-attention layer to obtain a first downsampling layer 31. Perform the above adjustment operation for each second downsampling layer 61, and the first downsampling network 302 in the first image generation model shown in (1) of Figure 6 can be obtained.

[0274] Exemplarily, as Figure 7 shown in (2) of, the second upsampling network 309 of the second image generation model includes k second upsampling layers 71, and each second upsampling layer 71 includes a second residual layer, a fourth self-attention layer, and a first cross-attention layer. For each second upsampling layer 71, copy the first cross-attention layer of the second upsampling layer 71 to obtain a second cross-attention layer, and add a second cross-attention layer after the first cross-attention layer of the second upsampling layer 71 to obtain a first upsampling layer 33. Perform the above adjustment operation for each second upsampling layer 71, and the first upsampling network 304 in the first image generation model shown in (1) of Figure 7 can be obtained.

[0275] Exemplarily, the second downsampling network 307 of the second image generation model includes k second downsampling layers 61, and the second upsampling network 309 of the second image generation model includes k second upsampling layers 71. The feature extraction network includes k processing layers. Taking the m-th processing layer as an example, the computer device copies the fourth self-attention layer of the m-th second upsampling layer 71 to obtain the m-th third self-attention layer, and constructs the m-th processing layer based on the m-th feature fusion layer and the m-th third self-attention layer. That is, the m-th processing layer is composed of the m-th feature fusion layer and the m-th third self-attention layer. In the above manner, k processing layers can be constructed in sequence to obtain a feature processing network including k processing layers. Among them, the m-th processing layer is added between the m-th second downsampling layer and the m-th second upsampling layer.

[0276] In this implementation manner, by copying the network layers in the trained second image generation model and constructing a new first image generation model based on the trained second image generation model and the copied network layers, the time and resource costs required for training a new model are reduced, and the construction efficiency of the image generation model is improved.

[0277] In the embodiments of the present application, the second image generation model is trained using sample content information and sample label images to improve the image generation ability of the second image generation model. Furthermore, based on the trained second image generation model, a first image generation model is constructed, which is equivalent to reusing the image generation ability of the second image generation model to construct a new first image generation model. The first image generation model can be used to generate images of any specified style. Therefore, there is no need to train models for generating images of different styles separately for different styles, saving the time and resource costs consumed by model training and improving the processing efficiency and versatility of the model.

[0278] The image generation method provided in the embodiments of the present application can be applied to any scenario of generating images of a specified style. For example, in the scenario of generating portrait photos of a specified style, the image generation process is as detailed in the following Figure 16 embodiment. Figure 16 is a flowchart of another image generation method provided in the embodiments of the present application. The embodiments of the present application are executed by a computer device, as Figure 16 shown. The method includes the following steps.

[0279] 1601. Obtain a reference portrait photo and multiple reference style photos.

[0280] Among them, the reference portrait photo includes an image, and the image can be an image of a person or an animal, etc. The multiple reference style photos have the same reference style. For example, the reference style refers to an anime style, etc.

[0281] 1602. Extract the image features corresponding to the reference image photo through the first image generation model, and extract the image encoding features corresponding to each of the multiple reference style photos respectively.

[0282] 1603. Through the first image generation model, perform a downsampling operation on the image encoding features of each reference style photo to obtain the reference style features corresponding to each reference style photo respectively.

[0283] 1604. Through the first image generation model, fuse the reference style features corresponding to the multiple reference style photos respectively to obtain the first fused style feature.

[0284] 1605. Through the first image generation model, perform a self-attention encoding operation on the first fused style feature to obtain the second fused style feature.

[0285] 1606. Through the first image generation model, perform an upsampling operation on the second fused style feature and the image features to obtain the target image features.

[0286] 1607. Through the first image generation model, perform a decoding operation on the target image features to obtain the target image photo. The target image photo has the same image as the reference image photo, and the target image photo has the same style as the reference style photo.

[0287] In the embodiments of the present application, the user provides a reference image photo that needs to be stylized, as well as multiple reference style photos that can represent the desired style. The first image generation model can automatically extract the style features from the input multiple reference style photos, perform style transfer on the input reference image photo, and thus generate a stylized target image photo. This solution can be applied to the functional expansion part of the magic camera, facilitating users to customize and stylize image photos, and realizing the expansion of the photo shoot gameplay.

[0288] In one embodiment, the image generation method of the embodiments of the present application can also be applied to the scenario of making DIY game characters. The user provides a description of the game character that needs to be stylized, as well as multiple game painting style photos that can represent the desired style. The first image generation model can automatically extract the style features from the input multiple game painting style photos, and use the style features to guide the process of intelligently generating game characters based on the game character description, realizing the custom design of the painting style of game characters. In video games, by opening the DIY mode and allowing users to input pictures of the desired game painting style by themselves, users can design game characters by themselves, increasing the fun of the game, and thus providing great help for the production of the game.

[0289] In one embodiment, the image generation method of the embodiments of the present application can also be applied to the scenario of restoring the painting style of an art painter. For example, if the user provides multiple photos of paintings and a reference scene photo to be stylized, the first image generation model can automatically extract style features from the input multiple photos of paintings, and use the style features to guide the style transfer of the reference scene photo to obtain a target scene photo that conforms to the painting style of the multiple photos of paintings, which is conducive to quickly creating various paintings with this painting style.

[0290] In a related technology, if you want to generate an image of a specified style, you need to collect a batch of high-quality images of the specified style in advance as training sample data, and then input them into the Stable Diffusion model. Use the Stable Diffusion model of this batch of training sample data to inject the concept of this style into the Stable Diffusion model, and then use the Stable Diffusion model to generate an image of this specified style. However, this method usually requires a large number of images as training sample data, and for different styles, it is necessary to train StableDiffusion models for generating images of different styles separately, consuming a lot of time and processing resources and increasing the processing burden.

[0291] However, the first image generation model provided by the embodiments of the present application can implement a zero-shot style transfer image generation method. For any image of a specified style to be generated, only a few images of the specified style need to be provided, and there is no need to train different models for different styles additionally. By capturing and balancing the style features of multiple specified-style images, the first image generation model can accurately generate images of the specified style. Compared with the above related technology, it can greatly save time and processing resources and reduce the processing cost.

[0292] In another related technology, a style image is provided, encoded into image features, and input into the model together with the text features of the description text to control the style guidance of the image generation process. Although this method can transfer the style semantics in the style image to the model, because the style semantics are not purified during the encoding process, in addition to the style semantics, the image features input into the model also include other semantics (such as people, scenery, etc.), which affects the accurate generation of the stylized image by the model. And only using one style image for encoding is likely to cause inaccurate style capture, thus affecting the style accuracy of the generated image.

[0293] The first image generation model provided by the embodiments of the present application constructs a feature balancing mechanism for multiple style images, supports receiving an indefinite number of style images as the input of the model, and this indefinite number of inputs enables the model to have a certain degree of scalability. Moreover, it can accurately capture common style features through multiple style images. The more style images are input, the more accurate the extracted style features are, which is beneficial to guiding the style transfer process more accurately.

[0294] In another related technology, multiple style images are provided. One of the style images is used as the base style image, and the other auxiliary style images are respectively serially calculated with the base style image to extract the common style features among the multiple style images. However, this way of combining the base style image with the auxiliary style images depends to a large extent on the style representation of the base style image. If the style representation of the base style image is not strong, it will lead to inaccurate finally extracted style features. Moreover, this unbalanced serial extraction mechanism reduces the role of the auxiliary style images, thereby reducing the utilization rate of the style images.

[0295] The first image generation model provided by the embodiments of the present application performs balanced calculations on multiple parallel-input style images without discrimination, directly extracts common style features from multiple style images through the balance calculation mechanism, and purifies the style features of multiple style images in parallel. It can avoid being restricted by a single style image, is beneficial to accurately capturing the style features of multiple style images, improves the accuracy of the extracted style features, and further ensures the accuracy of image generation.

[0296] Figure 17 It is a schematic structural diagram of an image generation device provided by the embodiments of the present application. Refer to Figure 17 This device includes:

[0297] An acquisition module 1701, configured to acquire content information and multiple reference images. The multiple reference images have the same reference style, and the content information is used to indicate the content of the image to be generated;

[0298] A feature extraction module 1702, configured to extract the content features corresponding to the content information, and extract the reference style features respectively corresponding to each of the multiple reference images;

[0299] A feature fusion module 1703, configured to fuse the reference style features respectively corresponding to the multiple reference images to obtain a first fused style feature;

[0300] A self-attention encoding module 1704, configured to perform a self-attention encoding operation on the first fused style feature to obtain a second fused style feature, and the second fused style feature is used to represent the reference style;

[0301] The image generation module 1705 is configured to generate a target image based on content features, guided by a reference style characterized by second fused style features, where the content of the target image matches the content indicated by the content information.

[0302] The image generation device provided by the embodiments of this application, when an image of a specified style needs to be generated, provides multiple reference images of the specified style. First, the style features of multiple different reference images are respectively extracted in parallel, and then the style features of multiple different reference images are fused and self-attention encoded, so as to capture the common and key style features in multiple different reference images, avoiding being affected by the semantic features of the reference images themselves, and ensuring the global relevance of the extracted style features. Furthermore, using the content features of the content information and the extracted style features to generate the target image can ensure that the content of the target image is the same as the content indicated by the content information, and the style of the target image is the same as the specified style. Therefore, the present application realizes a zero-shot style transfer image generation solution. Only by providing multiple reference images of any specified style can an image of this specified style be generated, not limited to only generating an image of a specific style, improving the flexibility and versatility of the image generation method. Moreover, guided by the style features extracted from multiple reference images, the accuracy of the generated image can be ensured.

[0303] Optionally, referring to Figure 18 , the first image generation model includes an encoding network and a first downsampling network; the feature extraction module 1702 includes:

[0304] The encoding unit 1712 is configured to perform an encoding operation on each reference image through the encoding network to obtain the image encoding features respectively corresponding to each reference image;

[0305] The downsampling unit 1722 is configured to perform a downsampling operation on the image encoding features respectively corresponding to each reference image through the first downsampling network to obtain the reference style features respectively corresponding to each reference image.

[0306] Optionally, referring to Figure 18 , the first downsampling network includes a first residual layer and a feature encoding layer; the downsampling unit 1722 is configured to:

[0307] Map the image encoding features respectively corresponding to each reference image to the first residual features respectively corresponding to each reference image through the first residual layer;

[0308] Perform an attention encoding operation on the first residual features respectively corresponding to each reference image through the feature encoding layer to obtain the reference style features respectively corresponding to each reference image.

[0309] Optionally, referring to Figure 18, the feature encoding layer includes a first self-attention layer and a second self-attention layer; the downsampling unit 1722 is configured to:

[0310] Through the first self-attention layer, perform self-attention encoding operations on the first residual features respectively corresponding to each reference image to obtain the first intermediate features respectively corresponding to each reference image;

[0311] Through the second self-attention layer, perform self-attention encoding operations on the first intermediate features respectively corresponding to each reference image to obtain the reference style features respectively corresponding to each reference image.

[0312] Optionally, refer to Figure 18 , the first downsampling network includes k first downsampling layers connected in sequence, where k is an integer greater than 1; the downsampling unit 1722 is configured to:

[0313] Through the first of the k first downsampling layers, perform downsampling operations on the image encoding features respectively corresponding to each reference image to obtain the first reference style features respectively corresponding to each reference image;

[0314] Through the m-th first downsampling layer among the k first downsampling layers, perform downsampling operations on the (m - 1)-th first reference style features respectively corresponding to each reference image to obtain the m-th reference style features respectively corresponding to each reference image, where m is an integer greater than 1 and not greater than k.

[0315] Optionally, refer to Figure 18 , the first image generation model includes a feature processing network, and the feature processing network includes k feature fusion layers; the feature fusion module 1703 is configured to:

[0316] Through the first of the k feature fusion layers, fuse the first reference style features respectively corresponding to multiple reference images to obtain the first first fusion style feature;

[0317] Through the m-th feature fusion layer among the k feature fusion layers, fuse the m-th reference style features respectively corresponding to multiple reference images to obtain the m-th first fusion style feature.

[0318] Optionally, refer to Figure 18 , the feature processing network further includes k third self-attention layers; the self-attention encoding module 1704 is configured to:

[0319] Through the first of the k third self-attention layers, perform self-attention encoding operations on the first first fusion style feature to obtain the first second fusion style feature;

[0320] Through the m-th third self-attention layer among the k third self-attention layers, perform self-attention encoding operation on the m-th first fusion style feature to obtain the m-th second fusion style feature.

[0321] Optionally, refer to Figure 18 , the first image generation model includes a decoding network and a first upsampling network; the image generation module 1705 is used for:

[0322] The upsampling unit 1715 is used to perform upsampling operation on the content feature and the second fusion style feature through the first upsampling network to obtain the target image feature;

[0323] The decoding unit 1725 is used to perform decoding operation on the target image feature through the decoding network to obtain the target image.

[0324] Optionally, refer to Figure 18 , the first upsampling network includes a second residual layer and an attention layer; the upsampling unit 1715 is used for:

[0325] Map the first noise feature to the second residual feature through the second residual layer, and the first noise feature is used to represent random noise;

[0326] Perform attention encoding operation on the content feature, the second fusion style feature and the second residual feature through the attention layer to obtain the target image feature.

[0327] Optionally, refer to Figure 18 , the attention layer includes a fourth self-attention layer, a first cross-attention layer and a second cross-attention layer; the upsampling unit 1715 is used for:

[0328] Perform self-attention encoding operation on the second residual feature through the fourth self-attention layer to obtain the third residual feature;

[0329] Perform cross-attention encoding operation on the third residual feature and the second fusion style feature through the first cross-attention layer to obtain the second intermediate feature;

[0330] Perform cross-attention encoding operation on the content feature and the second intermediate feature through the second cross-attention layer to obtain the target image feature.

[0331] Optionally, refer to Figure 18 , the first image generation model further includes a first downsampling network, the first downsampling network includes k first downsampling layers connected in sequence, the first upsampling network includes k first upsampling layers connected in sequence, the number of second fusion style features is k, and k is an integer greater than 1; the upsampling unit 1715 is used for:

[0332] Through the first first upsampling layer among the k first upsampling layers, perform an upsampling operation on the first noise feature, the content feature, and the first second fusion style feature to obtain the first target image feature. The first second fusion style feature is determined based on the feature output by the first first downsampling layer, and the first noise feature is used to represent random noise;

[0333] Through the m-th first upsampling layer among the k first upsampling layers, perform an upsampling operation on the (m - 1)-th target image feature, the content feature, and the m-th second fusion style feature to obtain the m-th target image feature. The m-th second fusion style feature is determined based on the feature output by the m-th first downsampling layer, where m is an integer greater than 1 and not greater than k.

[0334] Optionally, refer to Figure 18 , the decoding unit 1725 is used for:

[0335] Perform a decoding operation on the k-th target image feature through the decoding network to obtain the target image.

[0336] Optionally, refer to Figure 18 , the device further includes a model training module 1706, which is used for:

[0337] Obtain sample content information and a sample label image, where the content of the sample label image matches the content indicated by the sample content information;

[0338] Extract the sample content feature corresponding to the sample content information through the second image generation model, extract the first sample reference feature corresponding to the sample label image, and add a second noise feature to the first sample reference feature to obtain the second sample reference feature. The second noise feature is used to represent random noise;

[0339] Generate a sample target image based on the sample content feature under the guidance of the second sample reference feature through the second image generation model;

[0340] Train the second image generation model based on the sample target image and the sample label image to obtain the trained second image generation model;

[0341] The device further includes a model construction module 1707, which is used for:

[0342] Construct the first image generation model based on the trained second image generation model.

[0343] Optionally, refer to Figure 18, the trained first image generation model includes an encoding network, a second downsampling network, a second upsampling network, and a decoding network connected in sequence. The second downsampling network includes a first residual layer, a first self-attention layer, and a third cross-attention layer connected in sequence. The second upsampling network includes a second residual layer, a fourth self-attention layer, and a first cross-attention layer connected in sequence;

[0344] The model construction module 1707 is used for:

[0345] Copy the first self-attention layer to obtain a second self-attention layer; copy the first cross-attention layer to obtain a second cross-attention layer;

[0346] Copy the fourth self-attention layer to obtain a third self-attention layer; based on the feature fusion layer and the third self-attention layer, construct a feature processing network;

[0347] In the trained second image generation model, replace the third cross-attention layer of the second downsampling network with the second self-attention layer to obtain a first downsampling network. Add the second cross-attention layer after the first cross-attention layer of the second upsampling network to obtain a first upsampling network. Add the feature processing network between the first downsampling network and the first upsampling network to obtain the first image generation model.

[0348] It should be noted that: for the image generation device provided in the above embodiment, only the above division of each functional module is used for illustration. In actual application, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the image generation device provided in the above embodiment and the image generation method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0349] An embodiment of the present application also provides a computer device, which includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the image generation method of the above embodiment.

[0350] Optionally, the computer device is provided as a terminal. Figure 19 The structural schematic diagram of a first terminal 101 provided by an exemplary embodiment of the present application is shown.

[0351] The first terminal 101 includes: a first processor 1901 and a first memory 1902.

[0352] The first processor 1901 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The first processor 1901 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). The first processor 1901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the first processor 1901 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the first processor 1901 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0353] The first memory 1902 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The first memory 1902 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the first memory 1902 is used to store at least one computer program, and the at least one computer program is used to be possessed by the first processor 1901 to implement the image generation method provided in the method embodiments of the present application.

[0354] In some embodiments, the first terminal 101 may further optionally include: a peripheral device interface 1903 and at least one peripheral device. The first processor 1901, the first memory 1902, and the peripheral device interface 1903 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1903 through a bus, signal lines, or a circuit board. Optionally, the peripheral devices include at least one of a radio frequency circuit 1904, a display screen 1905, a camera assembly 1906, an audio circuit 1907, and a power supply 1908.

[0355] The peripheral device interface 1903 can be used to connect at least one I / O (Input / Output) related peripheral device to the first processor 1901 and the first memory 1902. In some embodiments, the first processor 1901, the first memory 1902, and the peripheral device interface 1903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the first processor 1901, the first memory 1902, and the peripheral device interface 1903 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0356] The radio frequency circuit 1904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1904 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1904 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1904 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 1904 can communicate with other devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, each generation of mobile communication network (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1904 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0357] The display screen 1905 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1905 is a touch display screen, the display screen 1905 also has the ability to collect touch signals on or above the surface of the display screen 1905. The touch signals can be input as control signals to the first processor 1901 for processing. At this time, the display screen 1905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1905, which is provided on the front panel of the first terminal 101; in other embodiments, there may be at least two display screens 1905, which are respectively provided on different surfaces of the first terminal 101 or are in a foldable design; in other embodiments, the display screen 1905 may be a flexible display screen, which is provided on the curved surface or the folding surface of the first terminal 101. Even further, the display screen 1905 can also be set to an irregular non-rectangular shape, that is, an irregular-shaped screen. The display screen 1905 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0358] The camera module 1906 is used to capture images or videos. Optionally, the camera module 1906 includes a front camera and a rear camera. The front camera is provided on the front panel of the first terminal 101, and the rear camera is provided on the back of the first terminal 101. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera respectively, to implement functions such as the combination of the main camera and the depth camera to achieve the background blurring function, the combination of the main camera and the wide-angle camera to achieve panoramic shooting and VR (Virtual Reality) shooting functions, or other combined shooting functions. In some embodiments, the camera module 1906 may also include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. A dual-color-temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0359] The audio circuit 1907 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the first processor 1901 for processing, or input to the radio frequency circuit 1904 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the first terminal 101. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the first processor 1901 or the radio frequency circuit 1904 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1907 may further include a headphone jack.

[0360] The power supply 1908 is used to supply power to each component in the first terminal 101. The power supply 1908 may be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 1908 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0361] Those skilled in the art can understand that Figure 19 the structure shown in

[0362] does not constitute a limitation on the first terminal 101, and may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement. Figure 20

[0363] Embodiments of the present application further provide a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the image generation method in the above embodiments.

[0364] An embodiment of the present application further provides a computer program product, including a computer program, which is loaded and executed by a processor to implement the operations performed by the image generation method in the foregoing embodiment.

[0365] Those of ordinary skill in the art can understand that all or part of the steps of implementing the foregoing embodiment can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like.

[0366] The foregoing are only optional embodiments of the embodiments of the present application, and are not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the present application.

Claims

1. An image generation method, characterized in that: The method comprises: Acquire content information and a plurality of reference images, wherein the plurality of reference images have the same reference style, and the content information is used to indicate the content of the image to be generated; Extracting content features corresponding to the content information through a feature extraction network in a first image generation model, wherein the first image generation model further includes an encoding network, a first downsampling network, a feature processing network, a first upsampling network, and a decoding network; Performing a coding operation on each of the multiple reference images through the coding network to obtain image coding features corresponding to each of the reference images; Performing a downsampling operation on the image encoding features corresponding to each of the reference images through the first downsampling network to obtain reference style features corresponding to each of the reference images; By means of the feature processing network, reference style features corresponding to the plurality of reference images are fused to obtain a first fused style feature; by means of the feature processing network, a self-attention encoding operation is performed on the first fused style feature to obtain a second fused style feature, wherein the second fused style feature is used to characterize the reference style; Performing an upsampling operation on the content feature and the second fused style feature through the first upsampling network to obtain a target image feature; Through the decoding network, a decoding operation is performed on the target image features to obtain a target image, the content of the target image matches the content indicated by the content information, and the style of the target image is the same as the reference style.

2. The method according to claim 1, characterized in that The first downsampling network includes a first residual layer and a feature encoding layer; the first downsampling network performs a downsampling operation on the image encoding features corresponding to each reference image to obtain the reference style features corresponding to each reference image, including: Mapping the image coding features corresponding to each of the reference images to first residual features corresponding to each of the reference images through the first residual layer; Through the feature encoding layer, an attention encoding operation is performed on the first residual feature corresponding to each reference image to obtain a reference style feature corresponding to each reference image.

3. The method according to claim 2, characterized in that The feature coding layer includes a first self-attention layer and a second self-attention layer; the feature coding layer performs an attention coding operation on the first residual feature corresponding to each reference image to obtain a reference style feature corresponding to each reference image, including: Performing a self-attention encoding operation on the first residual feature corresponding to each reference image through the first self-attention layer to obtain a first intermediate feature corresponding to each reference image; Through the second self-attention layer, a self-attention encoding operation is performed on the first intermediate features corresponding to each reference image to obtain the reference style features corresponding to each reference image.

4. The method according to claim 1, characterized in that: The first downsampling network includes k first downsampling layers connected in sequence, where k is an integer greater than 1; performing a downsampling operation on the image coding features corresponding to each reference image through the first downsampling network to obtain the reference style features corresponding to each reference image, including: Performing a downsampling operation on the image coding features corresponding to each of the reference images through a first first downsampling layer among the k first downsampling layers, so as to obtain a first reference style feature corresponding to each of the reference images; A downsampling operation is performed on the m-1th first reference style feature corresponding to each reference image through the mth first downsampling layer among the k first downsampling layers to obtain the mth reference style feature corresponding to each reference image, where m is an integer greater than 1 and not greater than k.

5. The method according to claim 4, characterized in that The feature processing network includes k feature fusion layers; the reference style features corresponding to the multiple reference images are fused through the feature processing network to obtain a first fused style feature, including: fusing the first reference style features respectively corresponding to the multiple reference images through the first feature fusion layer among the k feature fusion layers to obtain a first first fused style feature; The mth reference style features respectively corresponding to the multiple reference images are fused through the mth feature fusion layer among the k feature fusion layers to obtain the mth first fused style feature.

6. The method according to claim 5, characterized in that The feature processing network further includes k third self-attention layers; performing a self-attention encoding operation on the first fused style feature through the feature processing network to obtain a second fused style feature includes: Performing a self-attention encoding operation on the first first fused style feature through a first third self-attention layer among the k third self-attention layers to obtain a first second fused style feature; A self-attention encoding operation is performed on the mth first fusion style feature through the mth third self-attention layer among the k third self-attention layers to obtain the mth second fusion style feature.

7. The method according to claim 1, characterized in that The first upsampling network includes a second residual layer and an attention layer; the first upsampling network performs an upsampling operation on the content feature and the second fused style feature to obtain a target image feature, including: Mapping the first noise feature to a second residual feature through the second residual layer, where the first noise feature is used to characterize random noise; Through the attention layer, an attention encoding operation is performed on the content feature, the second fusion style feature and the second residual feature to obtain the target image feature.

8. The method according to claim 7, characterized in that The attention layer includes a fourth self-attention layer, a first cross-attention layer, and a second cross-attention layer; the attention layer performs an attention encoding operation on the content feature, the second fusion style feature, and the second residual feature to obtain the target image feature, including: Performing a self-attention encoding operation on the second residual feature through the fourth self-attention layer to obtain a third residual feature; Performing a cross-attention encoding operation on the third residual feature and the second fused style feature through the first cross-attention layer to obtain a second intermediate feature; Through the second cross-attention layer, a cross-attention encoding operation is performed on the content feature and the second intermediate feature to obtain the target image feature.

9. The method according to claim 1, characterized in that: The first down-sampling network includes k first down-sampling layers connected in sequence, the first up-sampling network includes k first up-sampling layers connected in sequence, the number of the second fused style features is k, and k is an integer greater than 1; performing an up-sampling operation on the content features and the second fused style features through the first up-sampling network to obtain target image features includes: performing an upsampling operation on the first noise feature, the content feature, and the first second fused style feature through the first first upsampling layer among the k first upsampling layers to obtain a first target image feature, wherein the first second fused style feature is determined based on the feature output by the first first downsampling layer, and the first noise feature is used to characterize random noise; An upsampling operation is performed on the m-1th target image feature, the content feature and the mth second fused style feature through the mth first upsampling layer among the k first upsampling layers to obtain the mth target image feature, wherein the mth second fused style feature is determined based on the feature output by the mth first downsampling layer, and m is an integer greater than 1 and not greater than k.

10. The method according to claim 9, characterized in that The step of performing a decoding operation on the target image feature through the decoding network to obtain the target image includes: Through the decoding network, a decoding operation is performed on the kth target image feature to obtain the target image.

11. The method according to any one of claims 1 to 10, characterized in that: The training process of the first image generation model includes: Acquire sample content information and a sample label image, wherein the content of the sample label image matches the content indicated by the sample content information; Extracting, by means of a second image generation model, a sample content feature corresponding to the sample content information, extracting a first sample reference feature corresponding to the sample label image, adding a second noise feature to the first sample reference feature to obtain a second sample reference feature, wherein the second noise feature is used to characterize random noise; Generate a sample target image based on the sample content feature by using the second image generation model and guided by the second sample reference feature; Based on the sample target image and the sample label image, training the second image generation model to obtain a trained second image generation model; Based on the trained second image generation model, the first image generation model is constructed.

12. The method according to claim 11, characterized in that The trained first image generation model includes an encoding network, a second downsampling network, a second upsampling network and a decoding network connected in sequence, the second downsampling network includes a first residual layer, a first self-attention layer and a third cross-attention layer connected in sequence, and the second upsampling network includes a second residual layer, a fourth self-attention layer and a first cross-attention layer connected in sequence; The step of constructing the first image generation model based on the trained second image generation model includes: Copy the first self-attention layer to obtain a second self-attention layer; copy the first cross-attention layer to obtain a second cross-attention layer; Copy the fourth self-attention layer to obtain a third self-attention layer; construct a feature processing network based on the feature fusion layer and the third self-attention layer; In the trained second image generation model, the third cross-attention layer of the second down-sampling network is replaced with the second self-attention layer to obtain a first down-sampling network, the second cross-attention layer is added after the first cross-attention layer of the second up-sampling network to obtain a first up-sampling network, and the feature processing network is added between the first down-sampling network and the first up-sampling network to obtain the first image generation model.

13. An image generating device, characterized in that: The device comprises: An acquisition module, used to acquire content information and a plurality of reference images, wherein the plurality of reference images have the same reference style, and the content information is used to indicate the content of the image to be generated; A feature extraction module, configured to extract content features corresponding to the content information through a feature extraction network in a first image generation model, wherein the first image generation model further includes an encoding network, a first downsampling network, a feature processing network, a first upsampling network, and a decoding network; The feature extraction module is further configured to perform a coding operation on each of the multiple reference images through the coding network to obtain image coding features corresponding to each of the reference images; The feature extraction module is further configured to perform a downsampling operation on the image coding features corresponding to each reference image through the first downsampling network to obtain a reference style feature corresponding to each reference image; A feature fusion module, configured to fuse the reference style features respectively corresponding to the plurality of reference images through the feature processing network to obtain a first fused style feature; A self-attention encoding module, configured to perform a self-attention encoding operation on the first fused style feature through the feature processing network to obtain a second fused style feature, wherein the second fused style feature is used to characterize the reference style; An image generation module, configured to perform an upsampling operation on the content feature and the second fused style feature through the first upsampling network to obtain a target image feature; The image generation module is further used to perform a decoding operation on the target image features through the decoding network to obtain a target image, wherein the content of the target image matches the content indicated by the content information, and the style of the target image is the same as the reference style.

14. A computer device, characterized in that: The computer device includes a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the image generation method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the image generation method according to any one of claims 1 to 12.

16. A computer program product comprising a computer program, characterized in that The computer program is loaded and executed by a processor to implement the operations performed by the image generating method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment, storage medium and program product

    CN116958323A

  • Image generation method, model training method and corresponding device

    CN117593400A

  • Image processing method, device, equipment, medium and program product

    CN118711113A