Image generation method and device, computer equipment and storage medium
By extracting the initial viewing image and describing text features of the game resources, the attention calculation is carried out to generate the target viewing image, which solves the problems of large differences and low efficiency in the three-view generation in the prior art, and realizes efficient multi-view production.
Patent Information
- Application Number
- CN202410199309.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-22
- Publication Date
- 2025-08-22
AI Technical Summary
In the prior art, the three-view production of game resources depends on the text image generation function of Stable Diffusion, resulting in a large difference between the generated other views and the original images, increasing the workload of post-modification, and affecting production efficiency.
By obtaining the initial viewing image of the target object and the description text corresponding to the target viewing angle, the image and text features are extracted, and attention calculation is performed to generate the target viewing angle image from the target viewing angle.
It improves the production efficiency of multi-views, reduces the workload of manual modification, and improves the accuracy and efficiency of image generation.
Smart Images

Figure CN120526005A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image generation method, apparatus, computer equipment, and storage medium. Background Art
[0002] The creation of three-view drawings is a crucial aspect of game asset production. Since most game character designs require post-production 3D modeling, concept artists, after designing the front view, must also create three-view drawings of the character to allow 3D modelers to accurately construct the character. A three-view drawing depicts a character from three sides: front, side, and back. With the widespread use of generative image models, and to reduce labor consumption, the creation of three-view drawings for game assets is increasingly being done using these models.
[0003] In the prior art, the creation of three-dimensional images for game assets typically relies on Stable Diffusion's text-to-image feature. This involves fine-tuning a LoRa network in conjunction with OpenPose to generate a set of three images based on text descriptions. However, the additional images generated using prompts can differ significantly from the original image, requiring subsequent image modification, increasing workload and impacting image production efficiency. Summary of the Invention
[0004] The embodiments of the present application provide an image generation method, apparatus, computer device, and storage medium, which can improve the efficiency of multi-view production.
[0005] The present invention provides an image generation method, including:
[0006] Obtain the initial perspective image of the target object and the description text corresponding to the target perspective;
[0007] Extracting image features of the initial perspective image and text features of the description text;
[0008] Performing attention calculation on the image features and the text features to obtain weighted features;
[0009] A target perspective image of the target object at the target perspective is generated based on the weighted features.
[0010] Accordingly, an embodiment of the present application further provides an image generating device, comprising:
[0011] A first acquisition unit is used to acquire an initial perspective image of the target object and a description text corresponding to the target perspective;
[0012] an extraction unit, configured to extract image features of the initial perspective image and text features of the description text;
[0013] A first processing unit is configured to perform attention calculation on the image features and the text features to obtain weighted features;
[0014] A generating unit is configured to generate a target perspective image of the target object at the target perspective based on the weighted features.
[0015] Correspondingly, an embodiment of the present application also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes any image generation method provided in any embodiment of the present application.
[0016] Correspondingly, an embodiment of the present application further provides a storage medium, which stores a plurality of instructions suitable for loading by a processor to execute the above image generation method.
[0017] By acquiring the initial view image of the target object and the descriptive text corresponding to the target view, the embodiment of the present application can extract the image features of the initial view image and the text features of the descriptive text. Attention calculation is then performed on the image features and text features to obtain weighted features. Finally, based on the weighted features, the target view image of the target object at the target view can be generated. This can improve the efficiency of multi-view production. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 A flowchart of an image generation method provided in an embodiment of the present application.
[0020] Figure 2 A schematic diagram of an application scenario of an image generation method provided in an embodiment of the present application.
[0021] Figure 3 A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application.
[0022] Figure 4 A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application.
[0023] Figure 5A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application.
[0024] Figure 6 A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application.
[0025] Figure 7 A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application.
[0026] Figure 8 A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application.
[0027] Figure 9 This is a structural block diagram of an image generation device provided in an embodiment of the present application.
[0028] Figure 10 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of this application.
[0030] The embodiments of the present application provide an information recommendation method, apparatus, storage medium and computer equipment. Specifically, the information recommendation method of the embodiments of the present application can be executed by a computer device, wherein the computer device can be a terminal or server device. The terminal can be a terminal device such as a smart phone, a tablet computer, a laptop computer, a touch screen, a personal computer (PC, Personal Computer), a personal digital assistant (Personal Digital Assistant, PDA), etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network, content distribution network), and big data and artificial intelligence platforms.
[0031] For example, the computer device can be a server, which can obtain an initial perspective image of the target object and a descriptive text corresponding to the target perspective; extract image features of the initial perspective image and text features of the descriptive text; perform attention calculation on the image features and text features to obtain weighted features; and generate a target perspective image of the target object at the target perspective based on the weighted features.
[0032] Based on the above problems, the embodiments of the present application provide a first image generation method, apparatus, computer device and storage medium, which can improve the efficiency of multi-view production.
[0033] It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.
[0034] An embodiment of the present application provides an image generation method, which can be executed by a terminal or a server. The embodiment of the present application takes the image generation method executed by a server as an example for explanation.
[0035] See also Figure 1 , Figure 1 This is a flow chart of an image generation method provided in an embodiment of the present application. The specific flow of the image generation method may be as follows:
[0036] 101. Obtain an initial perspective image of the target object and a description text corresponding to the target perspective.
[0037] In the embodiment of the present application, the target object may include an object in a real scene or an object in a virtual scene. For example, the object in the real scene may include at least real people and objects, and the object in the virtual scene may include virtual people and objects.
[0038] In some embodiments, the target object may be a virtual character in a virtual scene, such as a virtual character in a game scene.
[0039] The target object may be a three-dimensional object and may correspond to multiple perspective images, that is, images of the target object captured from different perspectives.
[0040] For example, capturing the target object from the front side can obtain a front view of the target object; capturing the target object from the side side can obtain a side view of the target object; capturing the target object from the back side can obtain a back view of the target object, and so on.
[0041] For example, see Figure 2 , Figure 2 A schematic diagram of an application scenario of an image generation method provided in an embodiment of the present application. Figure 2It shows images of the same target object at different viewing angles, including front view, side view, and back view.
[0042] In the embodiment of the present application, the initial perspective image of the target object may be an image of the target object at the initial perspective, for example, the initial perspective may be a frontal perspective, etc.
[0043] In some embodiments, the target object is a virtual character in a game scene, and the initial perspective image may be a front view of the target object drawn by a designer.
[0044] The target viewing angle may be a viewing angle different from the initial viewing angle. For example, the initial viewing angle may be a frontal viewing angle, and the target viewing angle may be a side viewing angle or a rear viewing angle.
[0045] Among them, the description text corresponding to the target perspective refers to the prompt word corresponding to the target perspective. For example, the target perspective can be the back perspective, and the corresponding description text can be "back perspective", "rear perspective" and other keywords associated with the back perspective.
[0046] In some embodiments, the initial perspective image of the target object may be a photographed image of the target object at the initial perspective, or a drawn image of the target object at the initial perspective, or may be obtained from an image database, etc.
[0047] 102. Extract image features of the initial view image and text features of the description text.
[0048] Among them, image features refer to features extracted from the initial perspective image, which may include appearance features, morphological features, etc. of the target object in the initial perspective image; text features refer to features extracted from the descriptive text, which may include text semantic features, etc.
[0049] In some embodiments, after acquiring the initial perspective image of the target object, the initial perspective image may be preprocessed first, and then feature extraction may be performed on the preprocessed initial perspective image, which may improve the accuracy of feature extraction.
[0050] Among them, image preprocessing can include various processing methods, such as image cropping, image scaling, etc.
[0051] In some embodiments, the initial perspective image may include background objects other than the target object, etc. In order to prevent the background objects from affecting the features of the target object, the background objects may be removed.
[0052] For example, after acquiring the initial perspective image, image recognition can be performed on the initial perspective image to identify multiple objects in the initial perspective image, and then objects other than the target object can be eliminated to obtain the initial perspective image containing only the target object.
[0053] Furthermore, the initial perspective image may be cropped according to the position of the target object in the initial perspective image so that the length and width of the initial perspective image correspond to the size of the target object in the image.
[0054] For example, see Figure 3 , Figure 3 A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application. Figure 3 The left side is the initial perspective image of the target object. In this initial perspective image, the target object occupies a small part of the image and there is a large part of the blank area. The initial perspective image can be cropped to obtain Figure 3 The cropped perspective image on the right shows that the target object occupies most of the image in this cropped initial perspective image.
[0055] In some embodiments, the step of “extracting image features of the initial perspective image and text features describing the text” may include the following operations:
[0056] Input the initial view image and description text into the target generation model;
[0057] The initial view image is encoded through the target generation model to obtain image features;
[0058] The description text is encoded through the target generation model to obtain text features.
[0059] In an embodiment of the present application, the target generation model can be used to generate a perspective image corresponding to a perspective prompt word of an input object based on a perspective image of the object and a perspective prompt word different from the perspective image.
[0060] The target generation model may include an image encoding module and a text encoding module. The image encoding module may be used to extract features from images, and the text encoding module may be used to extract features from text.
[0061] Specifically, encoding the initial view image through the target generation model may include: inputting the initial view image into an image encoding module of the target generation model, and extracting features of the initial view image through the image encoding module to obtain image features.
[0062] The extracting of features of the initial perspective image may include extracting latent space features of the image.
[0063] Latent space features refer to the patterns and relationships hidden in data. They are crucial for machine learning and deep learning tasks. For example, in image recognition, latent space features can represent texture and shape information within an image. In natural language processing, latent space features can represent semantic information within text. Furthermore, latent space features can be used in fields such as speech recognition and recommendation systems.
[0064] To extract latent space features, various machine learning and deep learning algorithms can be used. For example, a convolutional neural network (CNN) can be used to extract features from images and capture local and global patterns in the image; or a recurrent neural network (RNN) can be used to model sequence data and capture temporal dependencies and contextual information; or an autoencoder can be used to compress and reconstruct data and extract its implicit structure.
[0065] In an embodiment of the present application, the image encoding module of the target generation model may be an autoencoder, which encodes the initial view image through the autoencoder to obtain the latent space features of the initial view image as image features. The autoencoder may be pre-trained and have fixed parameters.
[0066] The autoencoder can be a variational autoencoder (VAE). A VAE is a generative model whose core concept is an encoder and a decoder. The encoder compresses the input data into a vector in a latent space, also known as a latent vector, while the decoder maps the latent vector back to the data in the original space. The VAE also includes a loss function to optimize the model parameters so that the generated data is as close as possible to the real data distribution.
[0067] Specifically, a variational autoencoder (VAE) is a generative model based on neural networks and a variant of an autoencoder. Its structure consists of two parts: an encoder and a decoder. The encoder maps the input data into a latent space, while the decoder remaps the latent space vectors back to the original data output. In the latent space, the VAE samples a random vector z, which is called a "latent variable." The basic principle of the VAE is to introduce a latent space based on the autoencoder and transform it into a trainable model by introducing the concept of variational inference.
[0068] In a VAE, the latent space is modeled as a Gaussian distribution, meaning that latent vectors can be sampled from this distribution. The mean and standard deviation of this Gaussian distribution are output by the encoder network. Typically, the latent vector in a VAE is sampled from a standard normal distribution with mean 0 and variance 1. This allows the VAE to sample new data from the original data distribution, thereby generating new data. The inference process of a VAE can be divided into two steps: encoding and decoding. The encoder maps the input data x to latent variables z in the latent space. This mapping consists of two parts: a variational layer that calculates the mean and variance of the latent variables, and a noise vector ε sampled from a standard normal distribution. The variational layer generates the mean and variance based on the input data x and then calculates the latent vector z by sampling ε. This sampling process uses a reparameterization technique, which reparameterizes the connection between the sampling and variational layers so that gradients can be propagated through the sampling process. The decoder maps the latent variables z to the output of the original data. The input to the decoder is the latent vector z and some noise. The role of the noise is to make the decoder more robust so that it can generate continuous output in the space around the latent vector z.
[0069] Specifically, the decoder uses a feedforward neural network to take the latent vector z and noise as input and output the reconstructed data x1. The decoder output is scaled by a sigmoid or tanh layer to ensure that the output value is within the range of the original data. During inference, the input data x is first fed into the encoder to obtain the mean and variance of the latent vector z. Then, an ε vector is sampled from a standard normal distribution, and the latent vector z is calculated using the mean and variance of the latent vector and the ε vector. Finally, z is fed into the decoder to obtain the reconstructed data x1.
[0070] For example, the initial view image is fed into the VAE in the target generative model. The VAE maps the initial view image to a low-dimensional latent variable in a latent space, yielding image features. Using an autoencoder, an image can be mapped from pixel space to latent space (latent space), learning the implicit representation of the image.
[0071] Specifically, encoding the description text through the target generation model may include: inputting the description text into a text encoding module of the target generation model, and extracting features from the description text through the text encoding module to obtain text features.
[0072] In the embodiment of the present application, the image text encoding module of the target generation model can be a text encoder, which encodes the description text through the text encoder to obtain text features. The text encoder can be pre-trained and has fixed parameters.
[0073] Among them, a text encoder is a model used in natural language processing (NLP) that converts text into a numerical representation so that computers can better understand and process it. This numerical representation is often called an "embedding" or "encoding" and it captures the semantic and grammatical information in the text.
[0074] In some embodiments, the text encoder can be CLIP (Contrastive Language-Image Pre-Training). CLIP is a model that uses text as a supervisory signal for pre-training. Its characteristics are: a multimodal model involving text and images; through contrastive learning, the cosine similarity of text features and image features is calculated, allowing the model to learn the matching relationship between text and image, including: positive samples: text and image match; negative samples: text and image do not match; it includes two encoders, which encode text and images to a fixed length and then calculate cosine similarity; one encodes text, generally a Transformer-based model; the other encodes images, which can be a CNN (ResNet50-like) or ViT model.
[0075] For example, the description text can be input into the CLIP in the target generative model, and the CLIP encodes the description text to generate an embedding representation of size [B, K, E] as the text feature. K represents the maximum encoding length of the text, and E represents the size of the embedding.
[0076] In some embodiments, in order to improve the image generation accuracy of the target generation model, before the step of "inputting the initial perspective image and the description text into the target generation model", the following steps may be further included:
[0077] Get a sample set;
[0078] Based on the sample perspective image and perspective description text in the sample pair, a target generation model is constructed.
[0079] The sample set may include multiple sample pairs, and each sample pair may include a sample view image and a view description text of a sample object.
[0080] For example, the sample perspective image in a certain sample pair may be a front view of object A, and the perspective description text may be a perspective prompt word such as a side view or a back view, which is different from the front view.
[0081] In an embodiment of the present application, sample pairs in the sample set can be used to train a target generation model.
[0082] Acquiring a sample set may include: first, capturing perspective images of multiple objects; for each perspective image, generating perspective keywords that are different from the perspective corresponding to the perspective image as perspective description text; and constructing an image-text pair as a sample pair based on each perspective image and the perspective description text generated based on the perspective image. Thus, a sample set including multiple sample pairs may be obtained.
[0083] The perspective images of the multiple objects may be acquired from an existing character image database, or by drawing character images, etc. The perspective images of the multiple objects may include initial perspective images of multiple virtual characters.
[0084] The captured perspective image needs to include the entire body of the virtual character.
[0085] In some embodiments, to further improve the accuracy of the generated image, the step of "building a target generation model based on the sample view image and view description text in the sample pair" may include the following operations:
[0086] According to the image area corresponding to the sample object in the sample perspective image, the sample perspective image is cropped to obtain a processed sample perspective image;
[0087] Obtaining the image size of the processed sample view image, and generating a mask image of a corresponding size based on the image size;
[0088] Splicing the processed sample view image and the mask image to obtain a spliced image;
[0089] Based on the spliced image and perspective description text, a target generation model is constructed.
[0090] In an embodiment of the present application, in order to eliminate information in the image that is irrelevant to the sample object, restore useful real information, enhance the detectability of relevant information and simplify the data to the maximum extent, thereby improving the reliability of feature extraction, image segmentation, matching and recognition, the sample perspective image can be preprocessed.
[0091] Among them, image preprocessing can include various processing methods, such as image cropping, image scaling, etc.
[0092] In some embodiments, the sample view image may include background objects other than the sample object. In order to prevent the background objects from affecting the features of the sample object to be extracted, the background objects may be removed.
[0093] For example, image recognition can be performed on the sample view image to identify multiple objects in the sample view image. Objects other than the sample object can then be removed to obtain a sample view image containing only the sample object. Furthermore, the sample view image can be cropped based on the position of the sample object in the sample view image so that the length and width of the sample view image correspond to the size of the sample object in the image, thereby obtaining a processed sample view image.
[0094] Masking is a common technique in image processing and computer vision. It involves overlaying a mask on an image to hide or reveal specific image areas. A mask can be a binary image, where areas with a pixel value of 1 represent the areas to be hidden or protected, and areas with a pixel value of 0 represent the areas that can be revealed. In practical applications, mask images can be generated using image processing algorithms (such as edge detection and region filling) or manually drawn.
[0095] In the embodiment of the present application, the mask image may be a mask image drawn according to the processed sample perspective image, and the pixels in the mask image may be set to -1.
[0096] Acquiring the image size of the processed sample perspective image and generating a mask image of a corresponding size based on the image size may include: generating a mask image having the same size as the processed sample perspective image according to the image size of the sample perspective image.
[0097] For example, see Figure 4 , Figure 4 A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application. Figure 4 The left side is the processed sample view image of the sample object, and a mask image of the same image size is generated based on the processed sample view image, that is, Figure 4 The mask image shown on the right has the same image size as the processed sample view image, wherein the image pixel of the mask image can be -1.
[0098] Among them, stitching the processed sample perspective image and the mask image may include: stitching the processed sample perspective image and the mask image side by side, that is, stitching the mask image to the right of the processed sample perspective image, so as to obtain a stitched image.
[0099] In some embodiments, before the processed sample image and the mask image are spliced together, the processed sample image may be normalized to normalize the image pixels in the processed sample image to 0-1.
[0100] For example, see Figure 5 , Figure 5 A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application. Figure 5 The left side shows the processed sample view image of the sample object and the generated mask image. The spliced image is obtained by splicing the mask image to the right side of the processed sample view image.
[0101] In some embodiments, the step of “building a target generation model based on the spliced image and the perspective description text” may include the following operations:
[0102] Input the spliced image and the view description text into the preset generation model;
[0103] Encoding the processed sample perspective image in the spliced image through an image encoding module of a preset generation model to obtain sample image features corresponding to the processed sample perspective image;
[0104] The perspective description text is encoded by a text encoding module of a preset generation model to obtain sample text features corresponding to the perspective description text;
[0105] The preset generation model is trained based on the sample image features and sample text features to obtain the target generation model.
[0106] The preset generative model can be the Inpainting ControlNet model, which can be used for image restoration and editing. It trains a neural network to predict the texture and structure of missing or damaged parts of an image, thereby restoring the original image content.
[0107] Specifically, Inpainting ControlNet uses contextual information to estimate the texture and structure of the missing part. Specifically, the method first extracts information around the damaged part of the image, and then uses this information to generate a repair area that matches the rest of the original image. Inpainting ControlNet uses a convolutional neural network (CNN) that is trained to learn local and global features of image texture and structure. During training, Inpainting ControlNet uses an image with a missing part as input and calculates the loss by comparing the network's output with the corresponding part of the original image. The network parameters are then optimized through a backpropagation algorithm to minimize the loss.
[0108] ControlNet is a component of Stable Diffusion. The Stable Diffusion model uses a diffusion process to generate stable, realistic images. ControlNet is a key component in Stable Diffusion, responsible for controlling the diffusion process to ensure that the generated image meets the expected properties and requirements.
[0109] Specifically, ControlNet is a conditional generative network that receives a target image or target attribute as input and generates a control signal. This control signal guides the diffusion process toward generating an image that meets the target attribute. ControlNet is usually a convolutional neural network (CNN) that can learn the mapping relationship between the target attribute and the generated image and adjust the diffusion process based on the target attribute. During training, ControlNet generates a control signal based on the target attribute. This control signal is combined with the noise signal in the diffusion process to generate an intermediate image. This intermediate image is then fed into a loss function along with the original image and the target image to calculate the training error. By continuously adjusting the parameters of ControlNet, the generated image can be made closer and closer to the target image, thereby achieving the image generation task.
[0110] The preset generation model may include an image encoding module and a text encoding module. The image encoding module may be used to extract features from images, and the text encoding module may be used to extract features from text. For example, the image encoding module may be a VAE, and the text encoding module may be a CLIP.
[0111] Specifically, the spliced image is input into the image coding module in the preset model, and the processed sample image on the left side of the spliced image is encoded by the image coding module to obtain image coding, which can be used as the sample image feature corresponding to the processed sample perspective image.
[0112] Specifically, the viewpoint description text is input into the text encoding module of the preset generative model. The text encoding module encodes the viewpoint description text to obtain an image code, which can be used as the sample text features corresponding to the viewpoint description text. The preset generative model can then be trained based on the sample image features and sample text features to obtain the target generative model.
[0113] In some embodiments, the step of “training a preset generative model based on sample image features and sample text features to obtain a target generative model” may include the following operations:
[0114] The feature processing module of the preset generative model performs attention calculation on the sample image features and sample text features to obtain weighted sample features;
[0115] An image generation module of a preset generation model is used to generate an image in the mask image area of the spliced image based on the weighted sample features to obtain a generated image;
[0116] Obtain the actual image corresponding to the sample object that meets the perspective description text;
[0117] Determining the difference between the generated image and the actual image;
[0118] The model parameters of the preset generation model are adjusted based on the difference information until the preset generation model converges to obtain the target generation model.
[0119] In an embodiment of the present application, the preset generation model may further include a feature processing module, which may be used to perform attention calculation on image features and text features. The feature processing module may include a self-attention submodule and a cross-attention submodule.
[0120] Among them, the self-attention submodule can be used to perform self-attention calculation on the features of the same modality of the input; the cross-attention submodule can be used to perform cross-attention calculation on the features of different modalities of the input.
[0121] In some embodiments, to learn image features and text features, the step of "performing attention calculation on sample image features and sample text features by a feature processing module of a preset generative model to obtain weighted sample features" may include the following operations:
[0122] Obtain the first latent feature of the randomly initialized noise;
[0123] Based on the self-attention submodule, self-attention calculation is performed on the first latent feature and the sample image feature to obtain the processed image feature;
[0124] Based on the cross-attention submodule, cross-attention calculation is performed on the processed image features and sample text features to obtain weighted sample features.
[0125] The random initialization noise is the random noise initially generated by the preset generation model. The first potential feature is obtained by inputting the random initialization noise into the VAE of the preset generation model and extracting the features of the random initialization noise through the VAE.
[0126] Among them, the first latent feature and the sample image feature are both features extracted by VAE and belong to the same modality.
[0127] The self-attention submodule uses the self-attention mechanism, a method for calculating the correlation between different elements in a sequence. It helps the model better capture key information in the sequence. It learns different relationships in the sequence by calculating the similarity between different positions in the sequence and taking a weighted average of the sequences based on the weights of the similarities.
[0128] Specifically, in the self-attention mechanism, the input data is divided into multiple sequences, and the elements of each sequence are calculated to determine their relevance to other elements. These relevance scores are normalized and used to weight the average representation of each element. This way, the final representation of each element takes into account information about other related elements.
[0129] The Self-Attention calculation process can include: performing a linear transformation on the input embedding vector to obtain the query vector, key vector, and value vector, then using the query vector and key vector to calculate the attention score, and finally calculating the output result using the attention score and value vector. The calculation formula can be as follows:
[0130]
[0131] Among them, Attention refers to attention, Q represents the Query vector, K represents the Key vector, V represents the Value vector, and d represents the vector dimension.
[0132] The Cross Attention submodule uses the Cross Attention mechanism, which introduces additional input sequences based on self-attention to integrate information from multiple sources. In machine translation, for example, the source and target language sentences are treated as two different input sequences and interact with each other through the Cross Attention mechanism, thereby better capturing the dependencies between the two languages.
[0133] Specifically, CrossAttention refers to the cross-attention layer between the encoder and decoder. In this layer, the decoder adjusts its attention to the encoder's output to obtain information about the encoder relevant to the current decoding position. In the encoder-decoder architecture, the encoder encodes the input sequence into a series of feature vectors, while the decoder gradually generates the output sequence based on these feature vectors. The CrossAttention layer is introduced to enable the decoder to effectively model the context of the current generation position.
[0134] The CrossAttention calculation process can include the following steps: Encoder input (usually the output from the encoder): usually represented as enc_inputs, with a size of (batch_size, seq_len_enc, hidden_dim); Decoder input (generated partial sequence): they are usually represented as dec_inputs, with a size of (batch_size, seq_len_dec, hidden_dim). Each position of the decoder generates a query vector (query), which is used to calculate the attention weights at all positions of the encoder. All positions of the encoder generate a set of key vectors (keys) and value vectors (values). The query vector (query) and the key vector (keys) are dot-producted, and the attention weights are obtained through the softmax function. The attention weights are multiplied by the value vector, and the results are summed to obtain the output adjusted by the encoder.
[0135] The latent space features of the first potential feature and the latent space features of the sample image feature are obtained, and the two latent space features are input into the self-attention submodule. The self-attention submodule performs self-attention calculation on the two latent space features, and the calculation result can be used as the processed image feature. The sample text feature and the processed image feature are then input into the cross-attention submodule. The cross-attention submodule calculates the cross-attention of the sample text feature and the processed image feature, and the result can be used as the weighted sample feature.
[0136] For example, see Figure 6 , Figure 6 This is a schematic diagram of another application scenario of the image generation method provided in the embodiment of the present application. Figure 6 First, the latent space features of the first potential feature and the latent space features of the sample image features are input into the self-attention submodule, which performs self-attention calculations to obtain processed image features. Then, the processed image features and sample text features are input into the cross-attention submodule, which processes the processed image features and sample text features to obtain weighted sample features. Attention calculations can obtain more complete and reliable feature information from multiple features.
[0137] In some embodiments, the method may further include the following steps:
[0138] Obtain time information and encode the time information through the time encoding module of the preset generation model to obtain time features;
[0139] Among them, when the preset generative model is trained, a time step of 0-999 will be initialized. This time step will be mapped into a time-embedding (time embedding), that is, a time feature, through MLP (Multilayer Perceptron).
[0140] Temporal embedding is a method that converts temporal information into numerical representations. It is commonly used in sequence modeling tasks in natural language processing, such as machine translation, text summarization, and sentiment analysis. Temporal embedding maps temporal information into a fixed-length vector, making it easier for computers to process and understand temporal information. In natural language processing, temporal embedding is often used as an additional dimension to word vectors, improving the model's ability to understand and express temporal information.
[0141] In the embodiment of the present application, the reason for adding Time-embedding is that the image generation module U-net in the preset generation model is parameter-shared. Adding Time-embedding can generate different outputs according to different inputs. It is hoped that during reverse diffusion, that is, the image generation process, U-net can first generate some rough outlines, very rough images, which do not need to be very clear. As the reverse diffusion moves forward, when it is close to restoring the original image, it can learn the information of the edges and corners of the object, some high-frequency information, so that the output image can be more realistic.
[0142] In some embodiments, the step of “training a preset generation model based on sample image features and sample text features to obtain a target generation model” may include the following operations:
[0143] The preset generation model is trained based on time features, sample image features, and sample text features to obtain the target generation model.
[0144] Specifically, the time feature and the sample image feature can be spliced together to obtain the spliced feature, and then the preset generation model can be trained based on the spliced feature and the sample text feature to obtain the target generation model.
[0145] In an embodiment of the present application, the preset generation model may include an image generation module, and the image generation module may be used to generate an image based on the weighted features.
[0146] The image generation module may include a denoising diffusion submodule and an image decoding submodule. The denoising diffusion submodule may be a Denoising U-net, which may use a denoising diffusion probability model (DDPM) to denoise the weighted sample features to obtain denoised fused features. The image decoding submodule may be a decoder, which may be used to restore the denoised fused features into an image.
[0147] Among them, DDPM is a generative model based on the diffusion process. It gradually adds random noise to the data and then learns the inverse diffusion process to construct the required data samples from the noise. The main idea of DDPM is to use a random process to generate a series of noisy images from data samples under gradually increasing noise conditions, and then restore them to clear images through denoising. This process can be formally represented as a Markov chain and learned using the gradient backpropagation algorithm. DDPM uses a fixed program to learn, and the latent variables have the same high dimension as the original data. In addition, DPM only needs to calculate the negative log-likelihood as the loss function, avoiding the problem of strict restrictions on the model structure or dependence on proxy objectives.
[0148] Specifically, the weighted sample features are added to the output hidden state of each layer of the U-net through the denoising diffusion submodule to obtain the superimposed features, and then the superimposed features are input into the image decoding submodule. The denoised fusion features are restored to an image through the image decoding submodule to obtain the generated image.
[0149] The actual image refers to the image of the sample object at the viewing angle corresponding to the viewing angle description text. For example, the viewing angle corresponding to the viewing angle description text may be a side view, and the actual image refers to the actual side view of the sample object.
[0150] In some embodiments, the actual image can be obtained from an image database and manually calibrated to an actual perspective image.
[0151] Determining the difference information between the generated image and the actual image may include: calculating the difference between the image features of the generated image and the image features of the actual image, and obtaining the feature difference value between the generated image and the actual image, which may be used as the difference information.
[0152] In some embodiments, the difference between the image features of the generated image and the image features of the actual image can be calculated according to a preset loss function.
[0153] Furthermore, model parameters of the preset generation model may be adjusted according to the difference information until the preset generation model converges to obtain a target generation model.
[0154] The embodiment of the present application trains a preset generation model by collecting perspective images of multiple objects, extracting features of perspective images of different objects, and adding perspective description text as an auxiliary condition. This can guide the preset generation model to generate high-quality perspective images during the training process.
[0155] In some embodiments, for the target generation model obtained after training, high-quality perspective images can also be obtained from the sample set to further fine-tune the target generation model to improve the quality of images generated by the target generation model.
[0156] For example, see Figure 7 , Figure 7 A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application. Figure 7 The figure shows the process of training the preset generative model to obtain the target generative model. The preset generative model consists of two parts: the stable diffusion model and the control network.
[0157] Among them, the latent diffusion model can be a StableDiffusion model pre-trained on a large amount of two-view data. The Stable Diffusion model is used as the base model of the preset generative model. On the basis of StableDiffusion, the structure and parameters of the decoder part of U-net are copied, and zero convolution is added to construct ControlNet. The front view of the sample object and the mask image generated based on the front view are used as condition inputs to control the generation of three views during the denoising process of Stable Diffusion. During training, the StableDiffusion network parameters are released, and the spliced image of the front view and other views is used as the supervisory signal. At the same time, the ControlNet network parameters are fine-tuned to obtain the pre-trained Multiview ControlNet, which can be used as the target to generate images.
[0158] Among them, each encoding block can be used to encode the image to obtain image features, the intermediate block can be used to perform attention calculation on text features, image features, etc. to obtain weighted features; the decoding block can be used to decode the weighted features to generate an image.
[0159] Among them, the description text is the text corresponding to the perspective keyword. The text module can be used to encode the description text to obtain text features; the time encoding module can be used to encode time to obtain time features.
[0160] The output may be an image of the sample object from another perspective generated according to the description text and the front view of the sample object.
[0161] During the training process, the parameters of the latent diffusion model and the control network are adjusted simultaneously to improve the model training effect.
[0162] 103. Perform attention calculation on image features and text features to obtain weighted features.
[0163] In an embodiment of the present application, the target generation model may include a feature processing module, which includes a self-attention submodule and a cross-attention submodule.
[0164] Specifically, the image features can be input into the self-attention submodule of the target generation model, and the self-attention submodule performs self-attention calculation on the image features to obtain processed image features; then the processed image features and text features are input into the cross-attention submodule, and the cross-attention submodule performs cross-attention calculation on the processed image features and text features to obtain weighted features.
[0165] 104. Generate a target perspective image of the target object at a target perspective based on the weighted features.
[0166] In some embodiments, the step of “generating a target-viewing angle image of the target object at a target viewing angle based on the weighted features” may include the following operations:
[0167] The weighted features are superimposed on the output hidden state through the target generation model to obtain the superimposed features;
[0168] The superimposed features are decoded to generate the target perspective image.
[0169] In an embodiment of the present application, the target generation model may include an image generation module, which may include a denoising diffusion submodule and an image decoding submodule.
[0170] Among them, the output hidden state, also known as hidden states, is output by each layer of U-Net. The output hidden state refers to the intermediate calculation results of each layer, which contain feature information at different levels.
[0171] In each U-Net layer, the convolution operation outputs a new feature map, which is the hidden state of the current layer. These hidden states contain different feature information at different levels. High-level hidden states contain more abstract global information, while low-level hidden states contain more specific local information. To fully utilize this feature information, the U-Net decoder directly connects the hidden states of the encoder to the corresponding layers of the decoder. This allows low-level feature information to be directly passed to the output layer, thereby improving image segmentation accuracy.
[0172] Specifically, the weighted features can be input into the denoising diffusion submodule of the target generation model. The weighted features are added to the output hidden state of each layer in the U-Net through the denoising diffusion submodule to obtain the superimposed features. The superimposed features are then input into the image decoding submodule. The superimposed features are restored to an image through the image decoding submodule to obtain the target image.
[0173] The target image may be a perspective image of the target object at a target perspective.
[0174] For example, the target perspective may be a back perspective, and the target perspective image may be a back view of the target object.
[0175] For example, see Figure 8 , Figure 8 A schematic diagram of an application scenario of another image generation method provided in an embodiment of the present application. Figure 8 The dorsal view of the generated target object is shown.
[0176] The target generation model trained by this solution can quickly and automatically generate a back view or side view that is highly aligned with a given front view (or other views), which can effectively save the time of game resource production.
[0177] This embodiment of the application discloses an image generation method, which includes: obtaining an initial perspective image of a target object and descriptive text corresponding to the target perspective; extracting image features of the initial perspective image and text features of the descriptive text; performing attention calculation on the image features and text features to obtain weighted features; and generating a target perspective image of the target object at the target perspective based on the weighted features. This method can improve the efficiency of multi-view production.
[0178] To facilitate better implementation of the image generation method provided in the embodiment of the present application, the embodiment of the present application also provides an image generation device based on the above image generation method. The meanings of the terms herein are the same as those in the above image generation method, and the specific implementation details can be referred to the description in the method embodiment.
[0179] See also Figure 9 , Figure 9 This is a structural block diagram of an image generation device provided in an embodiment of the present application, the device comprising:
[0180] The first acquisition unit 301 is used to acquire an initial perspective image of a target object and a description text corresponding to the target perspective;
[0181] An extraction unit 302 is configured to extract image features of the initial perspective image and text features of the description text;
[0182] A first processing unit 303 is configured to perform attention calculation on the image feature and the text feature to obtain a weighted feature;
[0183] The generating unit 304 is configured to generate a target perspective image of the target object at the target perspective based on the weighted features.
[0184] In some embodiments, the extraction unit 302 may include:
[0185] An input subunit, configured to input the initial perspective image and the description text into a target generation model;
[0186] A first processing subunit is configured to perform encoding processing on the initial perspective image using the target generation model to obtain the image features;
[0187] The second processing subunit is used to encode the description text through the target generation model to obtain the text features.
[0188] In some embodiments, the generating unit 304 may include:
[0189] a third processing subunit, configured to superimpose the weighted features and the output hidden state through the target generation model to obtain superimposed features;
[0190] The fourth processing sub-unit is used to decode the superimposed features to generate the target perspective image.
[0191] In some embodiments, the apparatus may further comprise:
[0192] a second acquiring unit, configured to acquire a sample set, the sample set comprising a plurality of sample pairs, each sample pair comprising a sample perspective image of a sample object and a perspective description text;
[0193] A construction unit is used to construct the target generation model based on the sample perspective image and the perspective description text in the sample pair.
[0194] In some embodiments, a building block may include:
[0195] a cropping subunit, configured to crop the sample perspective image according to an image area corresponding to the sample object in the sample perspective image to obtain a processed sample perspective image;
[0196] an acquisition subunit, configured to acquire an image size of the processed sample perspective image and generate a mask image of a corresponding size based on the image size;
[0197] a stitching subunit, configured to stitch the processed sample view image and the mask image to obtain a stitched image;
[0198] A construction subunit is used to construct the target generation model based on the spliced image and the perspective description text.
[0199] In some embodiments, the construction subunit may be specifically used to:
[0200] Inputting the spliced image and the perspective description text into a preset generation model;
[0201] Encoding the processed sample perspective image in the spliced image by using the image encoding module of the preset generation model to obtain sample image features corresponding to the processed sample perspective image;
[0202] Encoding the perspective description text by the text encoding module of the preset generation model to obtain sample text features corresponding to the perspective description text;
[0203] The preset generation model is trained based on the sample image features and the sample text features to obtain the target generation model.
[0204] In some embodiments, the construction subunit may be specifically used to:
[0205] Inputting the spliced image and the perspective description text into a preset generation model;
[0206] Encoding the processed sample perspective image in the spliced image by using the image encoding module of the preset generation model to obtain sample image features corresponding to the processed sample perspective image;
[0207] Encoding the perspective description text by the text encoding module of the preset generation model to obtain sample text features corresponding to the perspective description text;
[0208] Performing attention calculation on the sample image features and the sample text features through the feature processing module of the preset generation model to obtain weighted sample features;
[0209] Performing image generation in the mask image region of the spliced image based on the weighted sample features by the image generation module of the preset generation model to obtain a generated image;
[0210] Acquire an actual image corresponding to the sample object that conforms to the perspective description text;
[0211] determining difference information between the generated image and the actual image;
[0212] The model parameters of the preset generation model are adjusted based on the difference information until the preset generation model converges to obtain the target generation model.
[0213] In some embodiments, the construction subunit may be specifically used to:
[0214] Inputting the spliced image and the perspective description text into a preset generation model;
[0215] Encoding the processed sample perspective image in the spliced image by using the image encoding module of the preset generation model to obtain sample image features corresponding to the processed sample perspective image;
[0216] Encoding the perspective description text by the text encoding module of the preset generation model to obtain sample text features corresponding to the perspective description text;
[0217] Obtain the first latent feature of the randomly initialized noise;
[0218] Performing self-attention calculation on the first latent feature and the sample image feature based on the self-attention submodule to obtain a processed image feature;
[0219] Performing cross-attention calculation on the processed image features and the sample text features based on the cross-attention submodule to obtain the weighted sample features;
[0220] Performing image generation in the mask image region of the spliced image based on the weighted sample features by the image generation module of the preset generation model to obtain a generated image;
[0221] Acquire an actual image corresponding to the sample object that conforms to the perspective description text;
[0222] determining difference information between the generated image and the actual image;
[0223] The model parameters of the preset generation model are adjusted based on the difference information until the preset generation model converges to obtain the target generation model.
[0224] In some embodiments, the apparatus may further comprise:
[0225] The third acquisition unit is used to acquire time information and encode the time information through the time encoding module of the preset generation model to obtain a time feature.
[0226] In some embodiments, the construction subunit may be specifically used to:
[0227] Inputting the spliced image and the perspective description text into a preset generation model;
[0228] Encoding the processed sample perspective image in the spliced image by using the image encoding module of the preset generation model to obtain sample image features corresponding to the processed sample perspective image;
[0229] Encoding the perspective description text by the text encoding module of the preset generation model to obtain sample text features corresponding to the perspective description text;
[0230] The preset generation model is trained based on the time feature, the sample image feature, and the sample text feature to obtain the target generation model.
[0231] This embodiment of the application discloses an image generation device. A first acquisition unit 301 acquires an initial view image of a target object and a description text corresponding to a target view. An extraction unit 302 extracts image features from the initial view image and text features from the description text. A first processing unit 303 performs attention calculation on the image features and text features to obtain weighted features. A generation unit 304 generates a target view image of the target object at the target view based on the weighted features. This improves the efficiency of multi-view production.
[0232] Accordingly, the embodiment of the present application also provides a computer device, which may be a server. Figure 10 As shown, Figure 10 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. The computer device 500 includes a processor 501 having one or more processing cores, a memory 502 having one or more computer-readable storage media, and a computer program stored in the memory 502 and executable on the processor. The processor 501 is electrically connected to the memory 502. Those skilled in the art will appreciate that the computer device structure shown in the figure does not constitute a limitation of the computer device, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0233] The processor 501 is the control center of the computer device 500. It uses various interfaces and lines to connect various parts of the entire computer device 500. By running or loading software programs and / or modules stored in the memory 502 and calling data stored in the memory 502, it executes various functions of the computer device 500 and processes data, thereby monitoring the computer device 500 as a whole.
[0234] In the embodiment of the present application, the processor 501 in the computer device 500 loads instructions corresponding to one or more application processes into the memory 502 according to the following steps, and the processor 501 runs the application stored in the memory 502 to implement various functions:
[0235] Obtain an initial perspective image of the target object and a descriptive text corresponding to the target perspective; extract image features of the initial perspective image and text features of the descriptive text; perform attention calculation on the image features and text features to obtain weighted features; based on the weighted features, generate a target perspective image of the target object at the target perspective.
[0236] By acquiring the initial view image of the target object and the descriptive text corresponding to the target view, the embodiment of the present application can extract the image features of the initial view image and the text features of the descriptive text. Attention calculation is then performed on the image features and text features to obtain weighted features. Finally, based on the weighted features, the target view image of the target object at the target view can be generated. This can improve the efficiency of multi-view production.
[0237] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0238] Optional, such as Figure 10 As shown, the computer device 500 further includes: a touch screen 503, a radio frequency circuit 504, an audio circuit 505, an input unit 506, and a power supply 507. Among them, the processor 501 is electrically connected to the touch screen 503, the radio frequency circuit 504, the audio circuit 505, the input unit 506, and the power supply 507 respectively. It can be understood by those skilled in the art that Figure 10 The computer device structure shown in the figure does not constitute a limitation to the computer device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0239] The touch display screen 503 can be used for displaying a graphical user interface and receiving the operation instructions generated by the user acting on the graphical user interface. The touch display screen 503 can include a display panel and a touch panel. Among them, the display panel can be used for displaying the information input by the user or the information provided to the user and various graphical user interfaces of the computer device, and these graphical user interfaces can be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD, Liquid Crystal Display), an organic light emitting diode (OLED, Organic Light-Emitting Diode) and the like. The touch panel can be used for collecting the touch operation of the user thereon or near it (such as the user uses any suitable object or accessory such as a finger, a stylus on the touch panel or near the touch panel), and generates corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch direction, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into the touch point coordinates, and then sends it to the processor 501, and can receive the command sent by the processor 501 and execute it. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 501 to determine the type of touch event, and then the processor 501 provides a corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 503 to realize input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize input and output functions. That is, the touch display screen 503 can also be used as part of the input unit 506 to realize the input function.
[0240] The radio frequency circuit 504 may be used to transmit and receive radio frequency signals, so as to establish wireless communication with a network device or other computer device through wireless communication, and to transmit and receive signals between the network device or other computer device.
[0241] Audio circuit 505 can be used to provide an audio interface between the user and the computer device through a speaker and microphone. Audio circuit 505 can convert received audio data into electrical signals and transmit them to the speaker, which then converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuit 505 and converted into audio data. The audio data is then output to processor 501 for processing, then transmitted via RF circuit 504 to, for example, another computer device, or to memory 502 for further processing. Audio circuit 505 may also include an earphone jack to allow communication between external headphones and the computer device.
[0242] The input unit 506 may be configured to receive input digital, character information, or user feature information (such as fingerprint, iris, or facial information), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.
[0243] Power supply 507 is used to supply power to various components of computer device 500. Optionally, power supply 507 can be logically connected to processor 501 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. Power supply 507 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0244] although Figure 10 Not shown in the figure, the computer device 500 may also include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.
[0245] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0246] As can be seen from the above, the computer device provided in this embodiment can obtain the initial perspective image of the target object and the descriptive text corresponding to the target perspective; extract the image features of the initial perspective image and the text features of the descriptive text; perform attention calculation on the image features and text features to obtain weighted features; and based on the weighted features, generate a target perspective image of the target object at the target perspective.
[0247] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0248] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of computer programs, which can be loaded by a processor to execute the steps of any of the image generation methods provided in the embodiments of the present application. For example, the computer program can execute the following steps:
[0249] Obtain the initial perspective image of the target object and the description text corresponding to the target perspective;
[0250] Extracting image features of the initial view image and text features of the description text;
[0251] Perform attention calculation on image features and text features to obtain weighted features;
[0252] Based on the weighted features, a target perspective image of the target object at the target perspective is generated.
[0253] By acquiring the initial view image of the target object and the descriptive text corresponding to the target view, the embodiment of the present application can extract the image features of the initial view image and the text features of the descriptive text. Attention calculation is then performed on the image features and text features to obtain weighted features. Finally, based on the weighted features, the target view image of the target object at the target view can be generated. This can improve the efficiency of multi-view production.
[0254] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0255] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0256] Since the computer program stored in the storage medium can execute the steps in any image generation method provided in the embodiments of the present application, the beneficial effects that can be achieved by any image generation method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0257] The above is a detailed introduction to an image generation method, device, storage medium and computer equipment provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. An image generation method, characterized in that: The method comprises: Obtain the initial perspective image of the target object and the description text corresponding to the target perspective; Extracting image features of the initial perspective image and text features of the description text; Performing attention calculation on the image features and the text features to obtain weighted features; A target perspective image of the target object at the target perspective is generated based on the weighted features.
2. The method according to claim 1, characterized in that The extracting of image features of the initial viewing angle image and text features of the description text includes: Inputting the initial perspective image and the description text into a target generation model; encoding the initial perspective image using the target generation model to obtain the image features; The description text is encoded using the target generation model to obtain the text features.
3. The method according to claim 2, characterized in that Generating a target perspective image of the target object at the target perspective based on the weighted features includes: Superimposing the weighted features and the output hidden state through the target generation model to obtain superimposed features; The superimposed features are decoded to generate the target perspective image.
4. The method according to claim 2, characterized in that Before inputting the initial perspective image and the description text into the target generation model, the method further includes: Acquire a sample set, the sample set comprising a plurality of sample pairs, each sample pair comprising a sample perspective image of a sample object and perspective description text; The target generation model is constructed based on the sample perspective image and the perspective description text in the sample pair.
5. The method according to claim 4, characterized in that The constructing the target generation model based on the sample perspective image and the perspective description text in the sample pair includes: performing cropping processing on the sample perspective image according to an image area corresponding to the sample object in the sample perspective image to obtain a processed sample perspective image; Acquiring the image size of the processed sample perspective image, and generating a mask image of a corresponding size based on the image size; splicing the processed sample view image and the mask image to obtain a spliced image; The target generation model is constructed based on the spliced image and the perspective description text.
6. The method according to claim 5, characterized in that The constructing the target generation model based on the spliced image and the perspective description text includes: Inputting the spliced image and the perspective description text into a preset generation model; Encoding the processed sample perspective image in the spliced image by using the image encoding module of the preset generation model to obtain sample image features corresponding to the processed sample perspective image; Encoding the perspective description text by the text encoding module of the preset generation model to obtain sample text features corresponding to the perspective description text; The preset generation model is trained based on the sample image features and the sample text features to obtain the target generation model.
7. The method according to claim 6, characterized in that The training of the preset generation model based on the sample image features and the sample text features to obtain the target generation model includes: Performing attention calculation on the sample image features and the sample text features through the feature processing module of the preset generation model to obtain weighted sample features; Performing image generation in the mask image region of the spliced image based on the weighted sample features by the image generation module of the preset generation model to obtain a generated image; Acquire an actual image corresponding to the sample object that conforms to the perspective description text; determining difference information between the generated image and the actual image; The model parameters of the preset generation model are adjusted based on the difference information until the preset generation model converges to obtain the target generation model.
8. The method according to claim 7, characterized in that The feature processing module includes a self-attention submodule and a cross-attention submodule; The feature processing module of the preset generation model performs attention calculation on the sample image features and the sample text features to obtain weighted sample features, including: Obtain the first latent feature of the randomly initialized noise; Performing self-attention calculation on the first latent feature and the sample image feature based on the self-attention submodule to obtain a processed image feature; Based on the cross-attention submodule, cross-attention calculation is performed on the processed image features and the sample text features to obtain the weighted sample features.
9. The method according to claim 6, characterized in that The method further comprises: Acquire time information, and encode the time information through the time encoding module of the preset generation model to obtain a time feature; The step of training the preset generation model based on the sample image features and the sample text features to obtain the target generation model includes: The preset generation model is trained based on the time feature, the sample image feature, and the sample text feature to obtain the target generation model.
10. An image generating device, characterized in that: The device comprises: A first acquisition unit is used to acquire an initial perspective image of the target object and a description text corresponding to the target perspective; an extraction unit, configured to extract image features of the initial perspective image and text features of the description text; A first processing unit is configured to perform attention calculation on the image features and the text features to obtain weighted features; A generating unit is configured to generate a target perspective image of the target object at the target perspective based on the weighted features.
11. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein: When the processor executes the program, the image generating method according to any one of claims 1 to 9 is implemented.
12. A storage medium, characterized in that: The storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the image generation method according to any one of claims 1 to 9.