Construction method of image generation model, image generation method and related device
By constructing an image generation model and using category labels to generate images, the problems of high cost and slow amplification of gallery in the prior art are solved, and images are generated quickly and controllably, meeting the needs of rapid amplification of merchants and rapid iteration of goods.
Patent Information
- Application Number
- CN202510153368.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology is difficult to quickly and efficiently expand the library of business platforms, especially when the number of merchants is rapidly amplified and the product iteration is rapid, the cost of manually purchasing images is high and the cycle is long, making it difficult to meet the demand for the product to be launched quickly.
By constructing an image generation model, determine the target object to generate an image, obtain the image samples of the target object, configure a category label for each target object, and input the category label to the preset image generation model, generate a predicted image based on the category label, and train it according to the error between the predicted image and the image samples corresponding to the category label to generate a target image.
Without manual purchase of images, the images required to expand the gallery are quickly generated through the trained image generation model, which reduces the amplification cost, meets the needs of rapid iteration of products and efficiently launches, and reduces the uncertainty of image generation, making the generation process more controllable.
Smart Images

Figure CN120070637A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this specification relate to the field of image generation, and in particular, to a method for constructing an image generation model, an image generation method, a computer program product, an electronic device, and a computer-readable storage medium. Background Art
[0002] Currently, the existing image library of the business platform provides a large number of images for merchants, so that merchants can directly select corresponding images for their products in the image library, and then the products with images can be launched on the business platform for sale. However, with the rapid expansion of the business system and the number of merchants, it has brought great challenges to the diversity of product images in the business platform's image library and the range of products covered by the images. If the image library is expanded by manually purchasing images from outside, due to the high cost and slow purchasing cycle of the manual purchasing method, it is difficult to meet the needs of rapid product iteration and efficient launch. Summary of the Invention
[0003] In view of the above technical problems, embodiments of this specification provide a method for constructing an image generation model, an image generation method, a computer program product, an electronic device, and a computer-readable storage medium. The technical solutions are as follows:
[0004] According to the first aspect of the embodiments of this specification, a method for constructing an image generation model is provided. The method includes:
[0005] Determine each target object for which an image needs to be generated;
[0006] Obtain an image sample corresponding to each target object and containing the target object;
[0007] Configure a category label for each target object, and determine the image sample corresponding to each category label;
[0008] Input the category label into a preset image generation model. The image generation model generates a predicted image based on the input category label and is trained according to the error between the predicted image and the image sample corresponding to the input category label. Among them, the trained image generation model is used to generate a target image based on the input target category label, and the target image contains the target object corresponding to the target category label.
[0009] According to the second aspect of the embodiments of this specification, an image generation method is provided. The method includes:
[0010] In a preset category label, determine the target category label corresponding to the target object for which an image needs to be generated;
[0011] Input the target category label into a preset image generation model. Based on the input target category label, the image generation model generates a target image, and the target image includes a target object corresponding to the target category label.
[0012] Wherein, the image generation model is trained by using the preset category label to learn an image sample containing a target object corresponding to the preset category label.
[0013] According to the third aspect of the embodiments of the present specification, there is provided a computer program product, which includes a computer program. When the computer program is executed by a processor, the method described in the first aspect or the second aspect is implemented.
[0014] According to the fourth aspect of the embodiments of the present specification, there is provided an electronic device, which includes:
[0015] A processor;
[0016] A memory for storing instructions executable by the processor;
[0017] Wherein, the processor is configured to implement the method described in the first aspect or the second aspect.
[0018] According to the fifth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method described in the first aspect or the second aspect are implemented.
[0019] The technical solutions provided by the embodiments of the present specification may include the following beneficial effects:
[0020] In the embodiments of the present specification, category labels are configured for each target object that needs to generate an image, and image samples containing the target object corresponding to each category label are determined. The category labels are directly input into a preset image generation model. The image generation model generates a predicted image based on the input category labels and is trained according to the error between the predicted image and the image sample corresponding to the input category label. The trained image generation model can generate a target image based on the input target category label, and the target image contains the target object corresponding to the target category label. On the one hand, there is no need to expand the images in the image library by manual procurement, and the trained image generation model can directly and quickly generate the images containing the target object required for expanding the image library, greatly reducing the cost of expanding the image library of the business platform and meeting the requirements of the rapid iteration and efficient online launch of the target object (such as goods). On the other hand, by directly inputting category labels into the image generation model to generate images, there is no need for complex text feature support, and the model does not need to understand complex text semantics first. This not only reduces the difficulty of model training during training, improves the stability of model training, but also reduces the uncertainty of image generation, making the generation of images containing the target object more controllable.
[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the embodiments of the present specification. Brief Description of the Drawings
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments recorded in the present specification. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings.
[0023] Figure 1 is a schematic diagram of the application scenario of the embodiments of the present specification;
[0024] Figure 2 is a schematic flowchart of the method for constructing an image generation model according to an embodiment of the present specification;
[0025] Figure 3 is a schematic diagram of the architecture of training and testing a diffusion model based on the Transformer architecture according to an embodiment of the present specification;
[0026] Figure 4 is a schematic diagram of the structure of an image generation model according to an embodiment of the present specification;
[0027] Figure 5 is a schematic diagram of the structure of an image generation model according to another embodiment of the present specification;
[0028] Figure 6 It is a schematic structural diagram of an image generation model according to another embodiment of this specification;
[0029] Figure 7 It is a schematic flowchart of an image generation method according to an embodiment of this specification;
[0030] Figure 8 It is a schematic structural diagram of an electronic device according to an embodiment of this specification. Detailed implementation manners
[0031] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0032] The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this specification. The singular forms "a", "the", and "said" used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0033] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0034] Please refer to Figure 1, first, an application scenario of the embodiments of this specification will be introduced. Taking the distribution scenario as an example of the application scenario, distribution services are widely used in scenarios such as online shopping, food delivery, and errand shopping. In the distribution service scenario, there are a business platform, merchants, distribution capacity, and users. Among them, the business platform builds a server and provides a client for users. Users can use the services provided by the business platform through the client. Different from the client for users, the business platform also provides a client for merchants. Merchants can use the services provided by the business platform through this client. Distribution capacity refers to the party with distribution capabilities, including but not limited to distribution personnel, such as the so-called riders. Distribution capacity can communicate with the server through the client used by the distribution capacity. In other examples, distribution capacity can also include unmanned distribution devices, such as drones, unmanned vehicles, and so on. Among them, each user can trade with merchants through the client and initiate a distribution order; the business party can allocate distribution capacity for this instant distribution order.
[0035] Exemplarily, after the user places an order, the business platform server (such as an e-commerce platform or other cooperating distribution platforms) needs to generate a corresponding distribution waybill for the distributed items and allocate this distribution waybill to the distribution capacity. Therefore, after receiving this distribution waybill, the distribution capacity will go to the pick-up location to pick up the goods and deliver them to the location specified by the user after successful pick-up.
[0036] For example, in the food delivery scenario, the user places an order with a certain physical store on the food delivery platform through the client. After the food delivery platform generates a corresponding food delivery order, it allocates the distribution waybill to the distribution capacity (such as the client device used by the rider). Then the rider goes to the physical store (i.e., the pick-up location of the distributed items) to pick up the food delivery and delivers it to the location specified by the user.
[0037] Currently, the existing picture library of the business platform provides a large number of images for merchants, so that merchants can directly select corresponding images for their products in the picture library, and then the products with images can be launched on the business platform for sale. However, with the rapid expansion of the business system and the number of merchants, it also brings great challenges to the diversity of product images in the picture library of the business platform and the product scope covered by the images. If the picture library is expanded by manually purchasing images externally, due to the high cost and slow purchasing cycle of the manual purchasing method, it is difficult to meet the needs of rapid product iteration and efficient launch.
[0038] Taking the business platform as the food delivery platform and the goods as the dishes as an example, on the merchant interface of the food delivery platform, different dishes and the corresponding images of the dishes are often displayed. When a merchant launches a certain dish on the food delivery platform, it often needs to launch the image of the dish at the same time for the food delivery users who browse the merchant interface to make an intuitive selection and order of the dishes. In order to launch the dishes more conveniently and efficiently, merchants often directly select the images of the dishes from the dish picture library provided by the food delivery platform. Due to the characteristics of the dishes, there are a wide variety of dish types that a single merchant can launch, and the dishes of different merchants also vary greatly. In addition, there is also a need for rapid iterative updates for dishes with seasonal timeliness. The current dish picture library has been difficult to meet the needs of rapid iterative and efficient launch of dishes.
[0039] In addition, in the related art, there is a way to generate images by using Artificial Intelligence Generated Content (AIGC). AIGC is a technology that uses artificial intelligence technology to automatically generate various forms of digital content, including text, images, audio, and video, etc. AIGC has made remarkable progress in the fields of natural language processing, computer vision, and speech recognition, and has been applied in many fields such as content creation, personalized recommendation, and intelligent customer service, effectively improving production efficiency, reducing costs, and enhancing the user experience. The application of AIGC is very extensive. It can not only generate personalized content but also provide customized services to meet the specific needs of users. For example, in the field of content creation such as advertising, games, and self-media, AIGC technology has been widely applied, and at the same time, it has also tried to expand the application scope in the fields of education, e-commerce, software development, and finance. Taking the application of AIGC technology in the field of image generation as an example, using the image generation model of AIGC technology, images can be quickly generated according to text prompts.
[0040] However, when training a model based on existing open-source base models, not only is it restricted in terms of the controllability of the model structure and the flexibility of training, but also in the technical route of generating images according to text prompts, first of all, the model needs to be able to accurately understand the specific semantic information expressed by the text prompts. For example, the semantic information expressed by the text describing the dishes. Compared with the text semantic understanding of general scenarios, there are a wide variety of dish types themselves, so the model has a greater difficulty in learning text semantics and has more uncertainties.
[0041] In view of the above problems, the embodiments of this specification provide a method for constructing an image generation model, which can meet the needs of rapid iteration and efficient launch of target objects (such as goods), reduce the difficulty of model training, improve the stability of model training, and also reduce the uncertainty of image generation, making the generation of images containing target objects more controllable. As Figure 2 shown, the method includes the following steps:
[0042] S201. Determine each target object for which an image needs to be generated.
[0043] S202. Obtain an image sample corresponding to each of the target objects and containing the target object.
[0044] S203. Configure a class label for each of the target objects and determine the image sample corresponding to each class label respectively.
[0045] S204. Input the class label into a preset image generation model.
[0046] The image generation model generates a predicted image based on the input class label and is trained according to the error between the predicted image and the image sample corresponding to the input class label; wherein, the trained image generation model is used to generate a target image based on the input target class label, and the target image contains the target object corresponding to the target class label.
[0047] In the embodiments of this specification, a class label is configured for each target object for which an image needs to be generated, and the image sample corresponding to each class label and containing the target object is determined. The class label is directly input into a preset image generation model. The image generation model generates a predicted image based on the input class label and is trained according to the error between the predicted image and the image sample corresponding to the input class label. The trained image generation model can generate a target image based on the input target class label, and the target image contains the target object corresponding to the target class label. On the one hand, there is no need to expand the images in the image library by means of manual procurement, and the trained image generation model can directly and quickly generate the images containing the target object required for expanding the image library, greatly reducing the cost of expanding the image library of the business platform and being able to meet the requirements of the rapid iteration and efficient online launch of the target object (such as a commodity). On the other hand, by directly inputting the class label into the image generation model to generate an image, there is no need for complex text feature support, and the model does not need to understand complex text semantics first. This not only reduces the difficulty of model training during training, improves the stability of model training, but also reduces the uncertainty of image generation, making the generation of images containing the target object more controllable.
[0048] Before building an image generation model, it is first necessary to determine each target object for which images are to be generated. The target object can have various specific implementations. For example, if the target object is a commodity, the commodity can include dishes, daily necessities, etc. There can be multiple ways to determine the target object. As an example, in an existing image database, an object that is not covered by the images in the image database can be determined as the target object, that is, if there is a lack of images of a certain object in the image database, then that object can be used as the target object. As another example, in an existing image database, an object whose associated number of images is less than a quantity threshold can be used as the target object. As another example, in an existing image database, an object whose associated image quality is lower than a quality threshold can be used as the target object.
[0049] As an example, the existing image database can be a database maintained by the business platform to which the target object needs to be launched, or a database maintained by other platforms, and there is no specific limitation in this regard. There are various sources of images in the existing image database. It can be images of the target object uploaded by merchants, or relevant images of the target object obtained in advance from other online platforms.
[0050] The following takes the target object as a dish as an example for illustration: In an existing dish image database, if all the images in the dish image database do not contain a certain dish, and this dish is expected to be launched on the business platform, then this dish can be used as the target dish, that is, the target object; if the number of images associated with a certain dish in the dish image database is small, such as the number of images associated with this dish is less than or equal to the quantity threshold, and this dish is expected to be launched on the business platform, then this dish can also be determined as the target object; if the quality of the images associated with a certain dish in the dish image database is poor, such as the quality of the images associated with this dish is lower than the quality threshold, and this dish is expected to be launched on the business platform, then this dish can also be determined as the target object. By determining as the target object those that are expected to be launched on the business platform and for which there are no associated eligible images in the existing image database, the need for rapid iteration and efficient launch of the target object (such as dishes) can be further met.
[0051] After determining each target object for which images are to be generated, it is necessary to obtain sample data for training the model accordingly. The sample data includes various types of data, including obtaining an image sample containing each target object corresponding to each target object. The image sample can be obtained through various methods. Taking the target object as a dish as an example, as an example, in an existing dish image dataset, multiple existing images associated with each target object can be determined, and based on the multiple existing images associated with each target object, an image sample containing each target object corresponding to each target object can be determined.
[0052] Considering that not all of the multiple existing images associated with each target object in the existing dish image dataset necessarily have sufficient quality. For example, the similarity between the existing image and the target object may be low, or the subject included in the existing image may be incomplete (i.e., the target object is not fully shown in the image), or the existing image contains information that should not exist, such as watermark information. If all the multiple existing images associated with each target object are directly used as image samples containing the target object, it may result in a relatively low overall quality of the image samples and poor model training effects.
[0053] To address the above problems, after determining the multiple existing images associated with each target object, data preprocessing can be performed first. For example, quality filtering and screening can be carried out among the multiple existing images associated with each target object to ensure that the existing images used as image samples themselves have sufficient quality.
[0054] As an example, a text feature vector of each target object can be obtained by a text-image relevance matching model (Contrastive Language-Image Pre-Training, CLIP), and an image feature vector of each existing image associated with the target object can be obtained. Based on the text feature vector of each target object and the image feature vector of each existing image associated with the target object, the similarity between each target object and the multiple existing images associated with the target object can be determined. Among the multiple existing images associated with each target object, image samples can be determined based on the existing images whose similarity to the target object is greater than or equal to the similarity threshold. As another example, the target object has a name, and a text feature vector of each target object can be obtained. Specifically, a text feature vector can be extracted from the name of the target object.
[0055] As another example, based on a preset subject integrity judgment network, it can be determined whether the subjects included in the multiple existing images associated with each target object are complete. After determining the existing images with complete subjects among the multiple existing images, image samples containing the target object corresponding to each target object can be determined based on the existing images with complete subjects, and / or, based on a preset watermark recognition network, it can be determined whether the multiple existing images associated with each target object contain watermark information. After determining the existing images without watermark information among the multiple existing images, image samples containing the target object corresponding to each target object can be determined based on the existing images without watermark information. It can be understood that the existing images determined as image samples can meet at least one of the above three screening conditions (similarity to the target object is greater than or equal to the similarity threshold, the subject is complete, does not contain watermark information), or all of the above three screening conditions need to be met, and no specific limitation is made in this regard.
[0056] It can be understood that through the above-mentioned CLIP, the similarity score between the name of the target object and the image can be calculated, and this score reflects whether the image meets the description of the text describing the name of the target object. The higher the score, the better the matching degree between the two. Taking the target object as a dish as an example, the similarity score between the name of the dish and the image can be calculated, and this score reflects whether the image meets the description of the dish name text. The higher the score, the better the matching degree between the two.
[0057] The above-mentioned preset main body integrity judgment network can have various specific implementations. As an example, this main body integrity judgment network can be a classification network based on Vision Transformer (ViT), or it can be other types of networks, and specific details are not limited in this regard.
[0058] The above-mentioned preset watermark recognition network can have various specific implementations. As an example, this watermark recognition network can be a classification network based on ViT, or it can be other types of networks, and specific details are not limited in this regard.
[0059] The above-mentioned watermark information can have various specific implementations. As an example, this watermark information can be text watermark, image watermark, or logo information. Therefore, specific implementations of the watermark information are not limited.
[0060] It should be noted that the above introduction to the acquisition method of sample data is only an exemplary display, and other acquisition methods are not excluded in actual applications, and specific details are not limited in this regard.
[0061] The sample data for training the model also includes class labels. After obtaining the image samples containing each target object corresponding to each target object, it is necessary to configure a class label for each target object and determine the image samples corresponding to each class label. There are various ways to configure class labels for each target object. As an example, the name of the target object can be a class. If the name of the target object is text information, the text (text) describing the class can be converted into a class label (label). The class label can be data without semantic information, for example, it can be a numerical value. Taking the target object as a dish, the name of the dish is a class. For example, scrambled eggs with tomatoes is a class, and scrambled eggs with peppers is another class. Different class labels are configured for different dishes. For example, the class label "1" is configured for scrambled eggs with tomatoes, and the class label "2" is configured for "scrambled eggs with peppers". If there are 1000 dishes, 1000 class labels "0 to 999" can be configured. By constructing the correspondence between the dish name and the class label and using the class label as the input of the model, the trained model can directly understand the relationship between the class label and the corresponding image without having to understand the complex semantics expressed by the text describing the dish name.
[0062] It can be understood that each target object can uniquely correspond to a class label. Then, the image sample containing each target object corresponding to each target object is the image sample corresponding to the class label uniquely corresponding to the target object.
[0063] After the above sample data for training the model is determined, an image generation model can be constructed. As an example, the image generation model can be a diffusion model based on the Transformer architecture. The diffusion model based on the Transformer architecture uses the Transformer architecture as the backbone network of the diffusion model to replace the traditional convolutional neural network (such as U-Net) used as the backbone network, enabling it to process serialized image patches (latent patches) and capture global dependencies through the self-attention mechanism. The diffusion model based on the Transformer architecture can handle the latent representation of images and usually operates in the latent space of the images. This means that the image is first encoded into a lower-dimensional latent representation, and then the diffusion model is trained in this latent space to reduce the computational cost required for direct training in the high-resolution pixel space. The core of the Transformer architecture lies in the self-attention mechanism, which enables the model to process all elements in the input discrete sequence in parallel, thus greatly improving the training efficiency and effectively capturing long-range dependencies in the sequence data. The self-attention mechanism allows the model to consider all positions in the sequence simultaneously when processing information, rather than processing step by step like a recurrent neural network (RNN) or a convolutional neural network (CNN). Moreover, the Transformer architecture adopts an encoder-decoder structure, where the encoder is responsible for encoding the input sequence into a vector representation, and the decoder generates the target sequence based on the output of the encoder.
[0064] Please refer to Figure 3 , and the following is an exemplary introduction to the architecture of training and testing (inference) of the diffusion model based on the Transformer architecture in the embodiments of the present application:
[0065] The framework of the diffusion model based on the above Transformer architecture follows the way of simulating the diffusion process (latent diffusion) in the latent space. During the training process, all images (img) can be encoded into latent representations latent by a variational autoencoder (VAE encoder), and then according to the diffusion process, noise is added to the latent at each step until the latent completely becomes noise, and then the above diffusion model based on the Transformer architecture (DiT) is used to predict the noise at each step, and then the original latent is restored. During the inference process, a random noise is generated, the noise is predicted and removed by the diffusion model based on the Transformer architecture, and finally a new image is generated after decoding.
[0066] The above image generation model may include multiple Blocks, namely Figure 3 DiT in , the Block, as the basic building unit of the diffusion model based on the Transformer architecture, is a sub-module reused in the model, responsible for processing the latent representation of the image. Especially in the process of simulating the step-by-step denoising of image data, the model is constructed by stacking multiple Blocks. Each Block processes the input features and gradually improves the quality of the image, thereby generating a clear image. In addition, stacking Blocks can ensure the scalability of the model, that is, the complexity of the model can be achieved by increasing the number or width of Blocks.
[0067] Among them, the input image of the model can be encoded into a latent representation through a VAE (Variational Autoencoder). The latent representation can be divided into multiple patches, and each patch is embedded as a vector. Each patch can be understood as the latent representation corresponding to each image block of the input image; thus, a Token sequence containing multiple Tokens can be obtained, which can be called the initial Token sequence.
[0068] For example, each token may correspond to an image block (patch_size * patch_size) in the image. Therefore, the number of tokens T is equal to the height and width of the image divided by the square of patch_size, that is, T = (H / patch_size) * (W / patch_size), where H and W are the height and width of the image respectively. The dimension of each token is d, which is a configuration parameter of the model, representing the length of the feature vector of each token. In the DiT model, this dimension d is a hyperparameter of the model and can be adjusted according to the design and requirements of the model. Therefore, the dimension of the input token sequence of the Block is T * d, where T and d represent the number of tokens and the dimension of each token respectively.
[0069] The initial Token sequence can be input into the first Block among multiple Blocks. Inside the Block, the input Token sequence can be processed through self-attention mechanism and feed-forward network, etc., and a new Token sequence is output for the next Block to use. Multiple Blocks are cascaded and stacked. The input Token sequence of each layer of Block is the new Token sequence generated and output by the Block of the previous layer. Through multi-layer processing, the model gradually extracts and fuses multi-level features of the image, and finally generates a high-quality image.
[0070] The DiT model follows the absolute position encoding method of the traditional Transformer, that is, for a fixed-resolution image, the position encoding of each patch sequence is fixed. Although this is computationally simple, the disadvantages are also very obvious. It can only generate images of a specified resolution because the change in image resolution will cause the number of image patch sequences to change, while the absolute position encoding does not support the change of the sequence. That is, at a fixed image resolution, the number of patches is fixed, so the position encoding is also fixed. When dealing with images of different resolutions, resulting in a change in the number of patches, the fixed absolute position encoding cannot adapt, especially when there is a longer token length during testing than during training because there is no pre-defined encoding for the new patch positions.
[0071] Thus, the existing image generation models can only process images of one aspect ratio; for example, most of the current image generation models process images with an aspect ratio of 1:1, such as images with a resolution of 512*512, and some are 1024*1024. That is, the resolution of the input image of the model is 512*512, and the resolution of the generated image of the model is 512*512.
[0072] In order to obtain images of different aspect ratios from one model, this embodiment adopts the implementation of relative position encoding. As an example, the image generation model includes a diffusion model based on the Transformer architecture. The diffusion model includes multiple Blocks, and the Blocks are used to: obtain an input Token sequence, encode relative position information for each Token in the input Token sequence, and then process the Token sequence encoded with relative position information to obtain a new Token sequence and output it to be passed to the next Block;
[0073] Among them, in the multiple Blocks, the Tokens in the input Token sequence of the first Block include: the latent representations corresponding to the respective image patches of the image sample;
[0074] The relative position information encoded for each Token in the input Token sequence of the first Block includes: the relative position information between the image patch corresponding to the Token and other image patches.
[0075] Taking the existing DiT processing of a 512*512 input image as an example, assuming the patch_size is 2, in the Token sequence (T*d) converted based on this input image, T is 256*256. In the existing absolute position encoding scheme, for each Token in the Token sequence, the absolute position information of each Token in the 512*512 input image is encoded.
[0076] In this embodiment, multiple different aspect ratios can be set. Taking 3:2 and 16:9 as examples, assume the following two resolutions are adopted: 750*500 and 800*450.
[0077] In this embodiment, the latent representation of the image is segmented into latent representations (i.e., Tokens) corresponding to multiple image patches. Specific segmentation methods such as the segmentation size and the number of segments can be configured according to actual needs.
[0078] Existing images with these two resolutions can be obtained as image samples. For example, for the 750*500 image sample, its Token sequence is still T*d. However, based on the resolution of the input image sample, T is no longer fixed at 256*256 but can vary based on the resolution of the input image sample. In this embodiment, each Token can encode the relative position information between image patches. For example, RoPE encodes the positional relationship relative to other Tokens by performing rotation operations on the query (Query) and key (Key) vectors, enabling the model to naturally consider the relative positions between image patches during self-attention calculation. Similarly, similar processing is performed on the existing 800*450 image. In this way, the model can generate images with a resolution of 750*500 and images with a resolution of 800*450.
[0079] During the training process, for multiple different aspect ratios, multiple batches of training data can be used for model training. For example, for each aspect ratio, multiple annotated training data can be used; each batch can contain multiple image samples and their corresponding class labels. The image samples within the same annotation have the same aspect ratio, and between different batches, the image samples can have different aspect ratios. As an example, the number of batches for each aspect ratio is the same, and the amount of training data is also the same, so that the model can evenly obtain the image generation ability for each aspect ratio. Optionally, the model can be trained in the order of the training data for each batch of each aspect ratio.
[0080] To prevent the model from forgetting its capabilities, training data with various aspect ratios can be used for alternating training. For example, assume there are 3 aspect ratios. Batch 1: First, train with aspect ratio A, then with aspect ratio B, and then with aspect ratio C; the amount of training data for the 3 aspect ratios within Batch 1 is the same. Then proceed to Batch 2. Similarly to Batch 1, first train with aspect ratio A, then with aspect ratio B, and then with aspect ratio C; the amount of training data for the 3 aspect ratios within Batch 2 is the same.
[0081] After training is completed, based on the various aspect ratios used for the training data during training, the model has the ability to generate images with these various aspect ratios. After training is completed, when using the model to generate images, the aspect ratio used for the latent space of the model can be set according to the required aspect ratio, that is, the VAE encoder in the model is set to set the aspect ratio of the latent representation generated by the VAE encoder, so that the VAE encoder of the model can generate images with the corresponding aspect ratio. For example, if the aspect ratio of the latent representation generated by the VAE encoder of the model is set to A, then the VAE encoder of the model can generate images with an aspect ratio of A.
[0082] As Figure 4 shown, where each Block can at least include a first Root Mean Square Layer Normalization (RMS Norm) network layer, a first Scale and Shift layer, a Grouped Query Attention (GQA) layer, a first Scale layer, a second Root Mean Square Layer Normalization network layer, a second Scale and Shift layer, a Feedforward network layer, and a second Scale layer connected in sequence.
[0083] As an example, the above-mentioned Grouped Query Attention layer can include: multiple pairs of Key-Value, each pair of Key-Value corresponds to a group of Query, each group of Query contains at least two Query, and each pair of Key-Value is shared by each Query in the corresponding group of Query. Since each pair of Key-Value is shared by each Query in the corresponding group of Query, it can reduce the cache required for model training and inference, relieve the pressure on bandwidth and machine communication, and thus improve the speed of model training and inference.
[0084] As an example, the above image generation model may further include an adaptive normalization network layer; multiple Blocks included in the image generation model may share the same adaptive normalization network layer, and the adaptive normalization network layer is independently arranged outside the multiple Blocks for converting an input class label into an input feature and inputting the input feature into each Block respectively. As another example, the above input feature at least includes a scaling parameter input to the first parameter scaling and translation layer, a translation parameter input to the first parameter scaling and translation layer, a scaling parameter input to the first parameter scaling layer, a scaling parameter input to the second parameter scaling and translation layer, a translation parameter input to the second parameter scaling and translation layer, and a scaling parameter input to the second parameter scaling layer. It can be understood that the above parameters are associated with the class label input to the adaptive normalization network layer. Since the adaptive normalization network layer is a fully connected layer with a large number of parameters, avoiding setting an adaptive normalization network layer in each Block can avoid redundant parameter quantities, thereby avoiding slowing down the operation speed. By setting the adaptive normalization network layer outside the multiple Blocks and sharing the same adaptive normalization network layer by the multiple Blocks, the parameter quantity of the model can be reduced.
[0085] Due to the above diffusion model based on the Transformer architecture, during the training process, each input image is encoded into a latent by the VAE encoder structure, which can be processed into serialized image patches (latent patches). Specifically, its input can be a discrete sequence (input token), and the input Token is the result of serializing and partitioning the above latent. After partitioning, each patch is a Token. If the size of the latent is I*I and the size of each serialized token is p, the number T of tokens after serialization is: T=(I / p) 2 , and the dimension of the serialized Token can be expressed as: Token = T*d.
[0086] Considering that in the traditional diffusion models of related technologies, layer normalization (Layer Normalization, Layernorm layer) is generally used for normalization, which uses the mean and variance for normalization and has a large amount of calculation.
[0087] For the above problems, as an example, the first root mean square normalization network layer is used to: normalize each Token by using the root mean square value of each Token input to the first root mean square normalization network layer; the second root mean square normalization network layer is used to: normalize each Token by using the root mean square value of each Token input to the second root mean square normalization network layer. RMS Norm removes the process of subtracting the mean and only needs to calculate one statistic, the root mean square, that is, square the activation value of each neuron, calculate the average value, then take the square root, and finally divide the activation value by this root mean square value for normalization, thereby reducing the computational complexity to improve the running efficiency of the model. It can reduce the internal covariate shift within the model, prevent gradient explosion, thereby accelerating model training and improving performance.
[0088] The above feed-forward network layer uses a fully connected layer for further feature transformation after the above grouped query attention network layer.
[0089] Considering that in the traditional diffusion models of related technologies, the feed-forward network generally uses the Relu (Rectified Linear Unit) activation function, and its formula is as follows:
[0090] F out =W 2 ·Relu(W 1 F input ) (1)
[0091] Among them, F input represents the input, F out represents the output, W 1 and W 2 represent two fully connected layers, and the Relu activation function is used between the two layers. However, although the Relu activation function has sparsity and can reduce overfitting to a certain extent, it is very likely to cause the neuron gradient to disappear.
[0092] For the above problems, as an example, the feed-forward network layer of the embodiment of the present application can still include two fully connected layers, and the two fully connected layers are connected by an activation function. This activation function does not use the Relu activation function, but is constructed by combining the Swish activation function and a gate structure. Among them, the gating signal of the gate structure is determined based on the Swish activation function.
[0093] As an example, the Swish activation function replaces the traditional Relu activation function, and the formula can be as follows:
[0094] F out =W 2·Swish(W 1 F input ) (2)
[0095] The curve of the Swish activation function is smooth, and the function is differentiable at all points. Its response to negative values is relatively small, overcoming the drawback of the Relu activation function where the output is always zero at some neurons, resulting in the vanishing gradient. Further, a gate structure (Glu, Gated Linear Units) is introduced, and the formula is as follows.
[0096]
[0097] where represents element-wise multiplication. At this time, Swish(W 1 F input ) represents the activation degree of each point on the feature, and it multiplies element-wise with the W d F input after the initial transformation, so as to screen out which features need to pass through. This mechanism can enable the network to learn useful representations more effectively, helping to improve the generalization ability of the model. Since both the Swish activation function and the Glu gate structure are adopted simultaneously, this structure can be called: SwiGlu.
[0098] Please refer to Figure 4 、 Figure 5 、 Figure 6 , and the specific structure of the image generation model of the embodiments of the present application will be introduced exemplarily as follows:
[0099] As Figure 4 shown, the image generation model may include N Dit Blocks. Taking the structure of one Block as an example for introduction, first, Input Tokens can be used as the input of the Block, and the dimension is T*d. The acquisition method of Input Tokens is as described above, and will not be elaborated here specifically.
[0100] Considering that before the Input Tokens are input into the Block, due to the diffusion model based on the Transformer architecture, all images are encoded into latent through the VAE encoder structure during training. The Input Tokens are the results of serializing and partitioning the above latent. After partitioning, each patch can be used as a Token. However, the Transformer architecture itself does not contain the position information of each Token. Therefore, additional position encoding is required to represent the position order of each input Token. In the related art, the absolute position encoding method is adopted for each Token, which strictly controls the size of the input latent. The number of image blocks (latent patches) after serializing and partitioning the latent is fixed. Then, the absolute position encoding method is used to add position information to each patch. That is, for an image with a fixed resolution (such as a latent of I*I), the position encoding of each patch sequence is fixed. Although the calculation is simple in this way, the disadvantages are also very obvious, that is, only image blocks (latent patches) of a specified resolution size can be generated. Because if the resolution of the image block changes, it will cause the change of the number of image patch sequences (that is, the change of the number of Tokens T), and the above absolute position encoding method does not support the change of the sequence number.
[0101] As Figure 4 shown, to address the above problem, as an example, the absolute position encoding can be replaced by the Rotary Position Embedding (RoPE) and applied to the image generation model in the embodiments of the present application. RoPE means encoding based on the relative positions between each patch (Token), adding position information to each patch after partitioning the latent. In this way, whether the number of sequences increases or decreases, as long as the relative position relationship between two sequences is measured. Through the position encoding method of RoPE, the change of the sequence number can be supported, that is, the number of Tokens T can support dynamic changes, enabling the image generation model to support the inference of dynamic sequences.
[0102] After the Input Tokens complete the position encoding, they can be input into the Block. First, they go through the normalization process of RMS Norm, and RMS Norm can prevent gradient explosion.
[0103] For the input class label (label, that is, Figure 4 "Conditioning" in Figure 5The conditional information in), which is processed by an Adaptive Layer Normalization (adaLN) layer (the adaLN layer is the Figure 4 Multilayer Perceptron (MLP) in), can convert the input class label into an input feature. Among them, N Blocks can share the same adaLN layer (MLP). This adaLN layer is independently set outside the N Blocks and is used to convert the input class label (conditional information Conditioning) into an input feature and input this input feature into the N Blocks respectively. Since the adaLN layer is a fully connected layer and has a large number of parameters, avoiding setting an adaLN layer in each Block can avoid redundant parameter quantities and thus avoid slowing down the operation speed. By setting the adaLN layer outside the N Blocks and having the N Blocks share the same adaLN layer, the parameter quantity of the model can be reduced.
[0104] The parameters included in the above input feature can be respectively input into multiple network layers included in each of the N Blocks (such as Figure 4 the Scale, Shift layer, Scale layer in), in order to Figure 4 Taking a single block in the N Blocks of as an example, the above input feature at least includes the scaling parameter γ1 input to the first parameter scaling and translation layer (i.e., the Scale, Shift layer before "GQA" in Figure 4 ) and the translation parameter β1, where γ1 is used for inputting Scale (scaling) and β1 is used for inputting Shift (translation). The first parameter scaling and translation layer (i.e., the Scale, Shift layer) means scaling first and then translating.
[0105] The above input feature also includes the scaling parameter α1 input to the first parameter scaling layer (i.e., the Scale layer after "GQA" in the figure), the scaling parameter γ2 and the translation parameter β2 input to the second parameter scaling and translation layer (i.e., the Scale, Shift layer before "SwiGlu Feedforward" in Figure 4 ), where γ2 is used for inputting Scale (scaling) and β2 is used for inputting Shift (translation). The second parameter scaling and translation layer (i.e., the Scale, Shift layer) also means scaling first and then translating.
[0106] The above input features also include the scaling parameter α2 input to the second parameter scaling layer (i.e., the Scale layer after "SwiGlu Feedforward" in the figure). Among them, γ1, β1, α1, γ2, β2, and α2 are all associated with the class labels of the input MLP. The role of these parameters is to transform the network features while inputting the class labels, enhancing the non-linear expression ability.
[0107] For any Scale, Shift layer, the formulas for its scaling and translation can be as follows:
[0108] F out =F in *γ + β (4)
[0109] Where γ is the scaling parameter of the input Scale, β is the translation parameter of the input Shift, F in is the feature input to the Scale, Shift layer, and F out is the feature output after passing through the Scale, Shift layer.
[0110] For any Scale layer, the formula for its scaling can be as follows:
[0111] F out =F in *α (5)
[0112] Where α is the scaling parameter, F in is the feature input to the Scale layer, and F out is the feature output after passing through the Scale layer.
[0113] After the Input Tokens are normalized by RMS Norm, they then go through the scaling and translation operations of the Scale, Shift layer and enter the GQA layer. The purpose of the GQA layer is to capture the relationships between multiple Tokens, model the relationships of Tokens globally, and focus on the key features among them. The GQA layer can include multiple pairs of key-value (Key-Value), such as Figure 6As shown, taking G = 2 as an example, each pair of Key-Value corresponds to a group of queries Query. Each group of queries Query contains two queries Query, and each pair of Key-Value is shared by the two queries Query in the corresponding group of queries Query. Each group of queries shares a Key and Value matrix (key-value pair). Since the features of Key and Value can be shared (instead of each query having a different Key, query, and Value matrix respectively, and the weights of all Key and Value matrices are not shared), it can reduce the cache required for model training and inference, relieve the pressure on bandwidth and machine communication, and thus achieve an improvement in operation speed.
[0114] After the Tokens are processed by the GQA layer, they go through the scaling operation of the Scale layer, then through the normalization process of RMS Norm, and then through the scaling and translation operations of the Scale and Shift layers to reach the SwiGlu layer. The purpose of the SwiGlu layer is to further extract and fuse the feature information of Tokens at different positions to enhance the non-linear ability and expression ability of the model. After being processed by the SwiGlu layer and then through the scaling operation of the Scale layer, the dimension of the input Tokens is still T*d.
[0115] It should be noted that the above introduction to the specific structure of the image generation model of the present application embodiment is only an exemplary display. In actual applications, other specific structures are not excluded, and specific details are not limited here.
[0116] Based on the image generation model construction method described in any of the above embodiments, that is, the preset image generation model, the embodiments of this specification also provide an image generation method, as Figure 7 shown, the method includes the following steps:
[0117] S701. In the preset category labels, determine the target category label corresponding to the target object for which the image is to be generated;
[0118] S702. Input the target category label into the preset image generation model, and the image generation model generates a target image based on the input target category label.
[0119] The target image includes the target object corresponding to the target category label.
[0120] Among them, the image generation model is trained by using the preset category labels to learn the image samples containing the target object corresponding to the preset category labels.
[0121] The construction method of the above image generation model can be referred to the construction method of the image generation model described in any of the foregoing embodiments, and will not be elaborated herein specifically.
[0122] The target object for which an image needs to be generated can be determined in various ways. As an example, before determining the target category label corresponding to the target object for which an image needs to be generated, description information input by the user can be received, and the target object for which an image needs to be generated can be determined based on the description information input by the user.
[0123] For example, multiple target objects can be presented to the user. For example, the names of each target object can be shown in the user interface, such as the names of different target objects like "scrambled eggs with tomatoes" and "scrambled eggs with peppers" as described above; a function can be provided on the user interface for the user to select the name of a certain target object. After obtaining the name of the target object selected by the user, the target category label corresponding to the name of the target object selected by the user can be determined. For example, after obtaining that the user selects "scrambled eggs with peppers", the category label of "scrambled eggs with peppers" is obtained as "2"; "2" can be input into the image generation model. Since the image generation model has learned the correspondence between the category label "2" and the image of "scrambled eggs with peppers", the image generation model can generate an image containing "scrambled eggs with peppers".
[0124] Optionally, multiple different aspect ratios can also be output to the user for selection. After obtaining the target aspect ratio selected by the user, after setting the aspect ratio of the latent representation of the VAE encoder of the model, the model with the settings is used to generate an image with the target aspect ratio.
[0125] It should be noted that the above introduction to the method of determining the target object for which an image needs to be generated is only an exemplary display, and other determination methods are not excluded in practical applications, and no specific limitation is made thereto.
[0126] Considering that some unreasonable results may be encountered during the inference process of the image generation model, such as the similarity between the generated target image and the target object may be low, or the main body included in the target image is incomplete (that is, the target object is not completely shown in the target image), or the target image contains information that should not exist, such as watermark information, etc.
[0127] For the above problems, after the image generation model generates the target image, it is necessary to further filter the target image to select more high-quality images. As an example, taking the target object as a dish, the text feature vector of the target object corresponding to the target category label is obtained by the text-image relevance matching model, and the image feature vector of the target image corresponding to the target category label is obtained. Based on the text feature vector of the target object corresponding to the target category label and the image feature vector of the target image corresponding to the target category label, the similarity between the target object and the target image is determined. Then, based on the target images with a similarity greater than or equal to the similarity threshold, the images to be stored are determined, and the images to be stored are used to be stored in a preset dish image database. As another example, it is also possible to judge whether the main body included in the generated target image is complete based on a preset main body integrity judgment network, and determine the images to be stored based on the target images with a complete main body, and the images to be stored are used to be stored in a preset dish image database, and / or judge whether the generated target image contains watermark information based on a preset watermark recognition network, and determine the images to be stored based on the target images without watermark information, and the images to be stored are used to be stored in a preset dish image database. It can be understood that the determined images to be stored can meet at least one of the above three screening conditions (the similarity to the target object is greater than or equal to the similarity threshold, the main body included is complete, and there is no watermark information), or all three screening conditions need to be met, and no specific limitation is made thereto.
[0128] It can be understood that through the above text-image relevance matching model, the similarity score between the name of the target object and the target image can be calculated, and this score reflects whether the target image meets the description of the text describing the name of the target object. The higher the score, the better the matching degree between the two. Taking the target object as a dish as an example, the similarity score between the name of the dish and the target image can be calculated, and this score reflects whether the image meets the description of the dish name text. The higher the score, the better the matching degree between the two.
[0129] There can be various specific implementations of the above preset main body integrity judgment network. As an example, the main body integrity judgment network can be a classification network based on Vision Transformer (ViT), or other types of networks, and no specific limitation is made thereto.
[0130] There can be various specific implementations of the above preset watermark recognition network. As an example, the watermark recognition network can be a classification network based on ViT, or other types of networks, and no specific limitation is made thereto.
[0131] The above watermark information can have various specific implementations. As an example, the watermark information can be text watermark, image watermark, or logo information. Therefore, the specific implementation of the watermark information is not limited.
[0132] It should be noted that the above introduction to the screening method of the target image is only an exemplary display. In actual applications, other screening methods are not excluded, and the specific details are not limited.
[0133] As an example, if the generated target images all meet the above three screening conditions (similarity to the target object is greater than or equal to the similarity threshold, the main body contained is complete, and there is no watermark information), then further, the target images that meet these three screening conditions can be submitted to manual evaluation. Specifically, the quality and aesthetics of the target images can be further manually reviewed, mainly evaluating factors such as main body deformation, background deformation, image color score, composition score, and image homogenization, and further screening out target images with no main body deformation, background deformation, high-quality color composition, and low repetition rate as the images to be stored and storing them in the preset dish image database.
[0134] It can be understood that the above preset dish image database can serve the actual business. For example, when a merchant needs to launch a dish, they can directly select the image corresponding to the dish from this dish image database.
[0135] The embodiments of this specification also provide a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the behavior prediction method or the construction method of the behavior prediction model described in any of the above embodiments.
[0136] The embodiments of this specification also provide an electronic device, as Figure 8 shown, the electronic device includes:
[0137] A processor 801;
[0138] A memory 802 for storing instructions executable by the processor;
[0139] Wherein, the processor 801 is configured to implement the image generation method or the construction method of the image generation model described in any of the above embodiments.
[0140] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the image generation method or the construction method of the image generation model described in any of the above embodiments.
[0141] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the objectives of the solutions in this specification. Those of ordinary skill in the art can understand and implement them without creative efforts.
[0142] The above embodiments can be applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. The hardware of the computer device includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0143] The computer device can be any electronic product that can interact with users. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet protocol television (IPTV), a smart wearable device, etc.
[0144] The computer device may also include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.
[0145] The network where the computer device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.
[0146] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0147] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this patent. Adding insignificant modifications to the algorithm or process or introducing insignificant designs, but without changing the core design of its algorithm and process, are all within the protection scope of this application.
[0148] Among them, the description of "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of this specification. In this specification, the schematic expression of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0149] After considering the specification and practicing the invention claimed herein, those skilled in the art will readily conceive of other embodiments of this specification. This specification is intended to cover any variations, uses, or adaptations of this specification, which follow the general principles of this specification and include the common general knowledge or conventional technical means in the technical field not claimed in this application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of this specification are pointed out by the following claims.
[0150] It should be understood that this specification is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this specification is only limited by the appended claims.
[0151] The above are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this specification shall be included within the protection scope of this specification.
Claims
1. A method for constructing an image generation model, comprising: Determine each target object for which images need to be generated; Acquire an image sample corresponding to each target object and containing the target object; Configuring a category label for each target object and determining image samples corresponding to each category label; The category label is input into a preset image generation model, which generates a predicted image based on the input category label and is trained according to the error between the predicted image and the image sample corresponding to the input category label; wherein the trained image generation model is used to generate a target image based on the input target category label, and the target image contains a target object corresponding to the target category label.
2. According to the method of claim 1, the image generation model comprises a diffusion model based on a Transformer architecture, the diffusion model comprises a plurality of blocks, and the blocks are used to: Obtain an input Token sequence, encode relative position information for each Token in the input Token sequence, process the Token sequence encoded with relative position information to obtain a new Token sequence, and then output it to pass it to the next Block; in, Among the multiple blocks, the token in the input token sequence of the first block includes: a potential representation corresponding to each image block of the image sample; The relative position information encoded by each Token in the input Token sequence of the first Block includes: the relative position information between the image block corresponding to the Token and other image blocks.
3. According to the method of claim 2, each Block at least includes a first RMS normalized network layer, a first parameter scaling and translation layer, a group query attention network layer, a first parameter scaling layer, a second RMS normalized network layer, a second parameter scaling and translation layer, a feedforward network layer, and a second parameter scaling layer connected in sequence; The group query attention network layer includes: There are multiple Key-Value pairs, each pair of Key-Value corresponds to a set of queries, each set of queries contains at least two queries, and each pair of Key-Value is shared by each query in the corresponding set of queries.
4. According to the method of claim 3, the image generation model further includes an adaptive normalization network layer, and the multiple blocks share the adaptive normalization network layer; The adaptive normalization network layer is independently arranged outside the plurality of blocks to convert the input category labels into input features, and input the input features into each of the blocks respectively.
5. The method according to claim 4, wherein the input features at least include: The zoom parameters input to the first parameter zoom and translation layer, the translation parameters input to the first parameter zoom and translation layer, the zoom parameters input to the first parameter zoom and translation layer, the zoom parameters input to the second parameter zoom and translation layer, the translation parameters input to the second parameter zoom and translation layer, and the zoom parameters input to the second parameter zoom and translation layer.
6. According to the method of claim 3, the first RMS normalized network layer is used to normalize each Token using the RMS value of each Token input into the first RMS normalized network layer; the second RMS normalized network layer is used to normalize each Token using the RMS value of each Token input into the second RMS normalized network layer.
7. A method for generating an image, comprising: Determine the target category label corresponding to the target object for which the image needs to be generated in the preset category labels; Inputting the target category label into a preset image generation model, and generating a target image based on the input target category label by the image generation model, wherein the target image includes a target object corresponding to the target category label; Wherein, the image generation model is constructed using the method described in any one of claims 1 to 6, and the preset category labels are obtained by learning and training image samples containing target objects corresponding to the preset category labels.
8. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
9. An electronic device, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Cited By
Image processing method and device, computer equipment, medium and program product
CN121937292A
Two-dimensional mixed position coding system and method, medium, terminal and program product
CN121982310A
Image generation model construction method, image generation method, and related apparatus
WO2026170701A1