Construction method of image generation model, image generation method and related device
Through the diffusion model and feature fusion network training image generation model based on Transformer architecture, the problem of high library expansion cost and insufficient controllability of generating uncommon dishes is solved, efficient and accurate image generation is achieved, and the rapid iteration needs of the catering field is met.
Patent Information
- Application Number
- CN202510395206.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-01
AI Technical Summary
In the prior art, the cost and slow speed of expanding the image library through manual purchasing is high and it is difficult to meet the needs of rapid iteration of products and efficient launch. The existing image generation model is insufficient in the catering field, especially when generating images of uncommon dishes.
The diffusion model based on the Transformer architecture is used as the initial model, combining feature fusion networks and feature extraction networks, and the image generation model is trained through the error of image samples and predicted images, and the target image is directly generated, reducing dependence on text semantics, and improving the controllability and accuracy of image generation.
It reduces the cost of gallery amplification, improves the controllability and accuracy of image generation, meets the needs of rapid iteration of products and efficient launch, and the generated images are of higher quality and can better reflect the visual characteristics of the target object.
Smart Images

Figure CN120236182A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this specification relate to the field of image generation, and in particular, to a method for constructing an image generation model, an image generation method, a computer program product, an electronic device, and a computer-readable storage medium. Background Art
[0002] Currently, the existing image library on the business platform provides a large number of images for merchants, so that merchants can directly select corresponding images for their products in the image library, and then the products with images can be launched on the business platform for sale. However, with the rapid expansion of the business system and the number of merchants, it has brought great challenges to the diversity of product images in the image library of the business platform and the range of products covered by the images. If the image library is expanded by manually purchasing images from outside, due to the high cost and slow purchasing cycle of the manual purchasing method, it is difficult to meet the needs of rapid product iteration and efficient launch. Summary of the Invention
[0003] In view of the above technical problems, embodiments of this specification provide a method for constructing an image generation model, an image generation method, a computer program product, an electronic device, and a computer-readable storage medium. The technical solutions are as follows:
[0004] According to a first aspect of the embodiments of this specification, a method for constructing an image generation model is provided. The method includes:
[0005] Obtain a trained initial model; the initial model is used to generate a target image based on an input target category label, and the target image includes a target object corresponding to the target category label; the initial model includes a diffusion model based on the Transformer architecture;
[0006] Add a feature fusion network and a feature extraction network based on the initial model to obtain an image generation model;
[0007] Input an image sample including the target object into the image generation model, so that after the feature extraction network extracts the first image feature of the image sample, the diffusion model extracts the second image feature of the image sample, uses the feature fusion network to obtain the fusion feature of the first image feature and the second image feature, generates a predicted image based on the fusion feature, and trains the image generation model based on the error between the image sample and the predicted image.
[0008] According to a second aspect of the embodiments of this specification, an image generation method is provided. The method includes:
[0009] Obtain an input image including a target object;
[0010] Input the input image into a preset image generation model, and the image generation model generates a target image based on the input image, where the target image includes the target object;
[0011] Wherein, the image generation model is constructed by using the method described in the first aspect.
[0012] According to the third aspect of the embodiments of the present specification, there is provided a computer program product, which includes a computer program, and when the computer program is executed by a processor, the method described in the first aspect or the second aspect is implemented.
[0013] According to the fourth aspect of the embodiments of the present specification, there is provided an electronic device, which includes:
[0014] A processor;
[0015] A memory for storing instructions executable by the processor;
[0016] Wherein, the processor is configured to implement the method described in the first aspect or the second aspect.
[0017] According to the fifth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the method described in the first aspect or the second aspect are implemented.
[0018] The technical solutions provided by the embodiments of the present specification may include the following beneficial effects:
[0019] In the embodiments of the present specification, an input image including a target object is obtained, the input image is input into a preset image generation model, and the image generation model generates a target image based on the input image, where the target image includes the target object. The image generation model is obtained by adding a feature fusion network and a feature extraction network to an initial model that has been trained. The initial model includes a diffusion model based on the Transformer architecture, which is used to generate a target image based on the input target class label, and the target image includes the target object corresponding to the target class label. The image generation model generates a predicted image based on the fused features and is trained based on the error between the image sample and the predicted image. Among them, the fused features are obtained by fusing the first image feature and the second image feature using the feature fusion network. The first image feature is extracted from the image sample by the feature extraction network, and the second image feature is extracted from the image sample by the diffusion model.
[0020] On the one hand, instead of expanding the images in the image library through manual procurement, images containing target objects required for expanding the image library can be directly and quickly generated by a trained image generation model, greatly reducing the cost of expanding the image library of the business platform and meeting the requirements of the rapid iteration and efficient online launch of target objects (such as commodities).
[0021] On the other hand, by using a pre-trained diffusion model based on the Transformer architecture as the initial model, it is possible to initially classify the target objects at a coarse granularity, enabling the model to first understand the pixel semantic relationships between and within large categories from a macroscopic perspective, initially summarize the commonalities, and at the same time initially learn the differences between the images of different categories of target objects. The initial model provides the basic network parameters for the training of the subsequent image generation model, reducing the training difficulty of the image generation model during the secondary training, improving the stability of model training, and also reducing the uncertainty of image generation, making the generation of images containing target objects more controllable.
[0022] On the other hand, compared with the limitations of text descriptions, images contain richer visual features and the information presented is more intuitive. Using images as descriptions can better control the content and details of the target images generated by the image generation model. Therefore, using only images (image prompt) containing target objects as the input to the image generation model instead of text (text-to-image) as the input can improve the quality of the generated target images containing target objects.
[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the embodiments of this specification. Brief Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the embodiments of this specification or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0025] Figure 1 is a schematic diagram of the application scenario of the embodiments of this specification;
[0026] Figure 2 is a schematic flowchart of the method for constructing an image generation model according to an embodiment of this specification;
[0027] Figure 3 is a schematic diagram of the architecture of the training and testing of a diffusion model based on the Transformer architecture according to an embodiment of this specification;
[0028] Figure 4 It is a schematic structural diagram of an initial model according to an embodiment of this specification;
[0029] Figure 5 It is a schematic structural diagram of an image generation model according to an embodiment of this specification;
[0030] Figure 6 It is a schematic structural diagram of an image generation model according to an embodiment of this specification;
[0031] Figure 7 It is a schematic flowchart of an image generation method according to an embodiment of this specification;
[0032] Figure 8 It is a schematic structural diagram of an electronic device according to an embodiment of this specification. Detailed implementation manners
[0033] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0034] The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms "a", "the", and "said" used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0035] It should be understood that although the terms first, second, third, etc. may be used in this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0036] Please refer to Figure 1, first, an application scenario of the embodiments of this specification will be introduced. Taking the delivery scenario as an example of the application scenario, the delivery service is widely used in scenarios such as online shopping, food delivery, and errand shopping. In the delivery service scenario, there are a business platform, merchants, delivery capacity, and users. Among them, the business platform builds a server and provides a client for users. Users can use the services provided by the business platform through the client. Different from the client for users, the business platform also provides a client for merchants. Merchants can use the services provided by the business platform through this client. The delivery capacity refers to the party with delivery capabilities, including but not limited to delivery personnel, such as the so-called riders. The delivery capacity can communicate with the server through the client used by the delivery capacity. In other examples, the delivery capacity can also include unmanned delivery devices, such as drones, autonomous vehicles, and so on. Among them, each user can trade with merchants through the client and initiate a delivery order; the business side can allocate delivery capacity for this instant delivery order.
[0037] Exemplarily, after the user places an order, the business platform server (such as an e-commerce platform or other cooperating delivery platforms) needs to generate a corresponding delivery waybill for the delivered item and allocate this delivery waybill to the delivery capacity. Therefore, after receiving this delivery waybill, the delivery capacity will go to the pick-up location to pick up the goods and deliver them to the location designated by the user after successful pick-up.
[0038] For example, in the food delivery scenario, the user places an order with a certain physical store on the food delivery platform through the client. After the food delivery platform generates a corresponding food delivery order, it allocates the delivery waybill to the delivery capacity (such as the client device used by the rider), so that the rider can go to the physical store (i.e., the pick-up location of the delivered item) to pick up the food delivery and deliver it to the location designated by the user.
[0039] Currently, the existing image library of the business platform provides a large number of images for merchants, so that merchants can directly select corresponding images for their products in the image library, and then the products with images can be launched on the business platform for sale. However, with the rapid expansion of the business system and the number of merchants, it also brings great challenges to the diversity of product images in the business platform's image library and the scope of products covered by the images. If the image library is expanded by manually purchasing images externally, due to the high cost and slow purchasing cycle of the manual purchasing method, it is difficult to meet the needs of rapid product iteration and efficient launch.
[0040] Taking the business platform as a food delivery platform and the goods as dishes as an example, the merchant interface of the food delivery platform often displays different dishes and the corresponding images of the dishes. When a merchant launches a certain dish on the food delivery platform, it often needs to launch the image of the dish at the same time for food delivery users browsing the merchant interface to make an intuitive selection and order of the dishes. In order to launch dishes more conveniently and efficiently, merchants often directly select the images of the dishes from the dish picture library provided by the food delivery platform. Due to the characteristics of the dishes, there are a wide variety of dish types that a single merchant can launch, and there are also significant differences in the dishes of different merchants. In addition, dishes with seasonal timeliness also have the need for rapid iterative updates. The current dish picture library has been difficult to meet the demand for rapid iterative and efficient launch of dishes.
[0041] In addition, in the related technologies, there is a way to generate images by using Artificial Intelligence Generated Content (AIGC). AIGC is a technology that uses artificial intelligence technology to automatically generate various forms of digital content, including text, images, audio, and video, etc. AIGC has made remarkable progress in the fields of natural language processing, computer vision, and speech recognition, and has been applied in many fields such as content creation, personalized recommendation, and intelligent customer service, effectively improving production efficiency, reducing costs, and enhancing the user experience. The application of AIGC is very extensive. It can not only generate personalized content but also provide customized services to meet the specific needs of users. For example, in content creation fields such as advertising, games, and self-media, AIGC technology has been widely applied, and at the same time, it has also tried to expand the application scope in fields such as education, e-commerce, software development, and finance. Taking the application of AIGC technology in the field of image generation as an example, in the related technologies, the image generation model using AIGC technology can quickly generate images according to text prompts.
[0042] However, training a model based on existing open-source base models is not only restricted in terms of the controllability of the model structure and the flexibility of training, but also in terms of the technical route of generating images based on text prompts (text to image). First of all, the model needs to be able to accurately understand the specific semantic information expressed by the text prompt. For example, for the semantic information expressed by the text describing a dish, compared with the text semantic understanding of general scenes, there is a large variety of dish types, so it is more difficult for the model to learn text semantics and there is more uncertainty. Secondly, taking the food and beverage field as an example, currently in the food and beverage field, dish images are often generated by inputting text prompts of dishes. However, the food and beverage industry belongs to a typical vertical domain. The rich ingredients and various cooking methods make the styles and details of dish presentations extremely complex and diverse. Therefore, it is very difficult to comprehensively and meticulously describe fine-grained dish images relying solely on text, resulting in low controllability of the images generated by the model, and often the problem that the generated images do not match the required image content.
[0043] For example, for relatively common dishes such as scrambled eggs with tomatoes or shredded potatoes, although the method of generating images using text prompts and an image generation model can also generate images that match the dishes. However, for some uncommon dishes and dishes with rare names, such as stir-fried bacon with bamboo shoots, pounded pepper and rice, etc., their ingredient compositions and cooking methods may be relatively complex, and the corresponding images may also be relatively complex. In this case, if the method of still using text prompts and an image generation model to generate images is adopted, on the one hand, there is less data available for the model to learn, and on the other hand, the frequency of uncommon dishes and dishes with rare names appearing in the collected dish data is also relatively low. Eventually, it will lead to insufficient model learning, the generated images not matching the dishes, and not meeting the requirements of users.
[0044] In view of the above problems, the embodiments of this specification provide a method for constructing an image generation model, which can reduce the cost of expanding the image library of the business platform, meet the requirements of rapid iteration and efficient online launch of target objects (such as goods), reduce the uncertainty of image generation, make the generation of images containing the target object more controllable, and improve the quality of the generated target images containing the target object. As Figure 2 shown, the method includes the following steps:
[0045] S201. Obtain a trained initial model.
[0046] The initial model is used to generate a target image based on the input target category label, and the target image contains the target object corresponding to the target category label; the initial model includes a diffusion model based on the Transformer architecture.
[0047] S202. Add a feature fusion network and a feature extraction network based on the initial model to obtain an image generation model.
[0048] S203. Input an image sample containing the target object into the image generation model, so that after the feature extraction network extracts the first image feature of the image sample, the diffusion model extracts the second image feature of the image sample. Use the feature fusion network to obtain the fusion feature of the first image feature and the second image feature, generate a prediction image based on the fusion feature, and train the image generation model based on the error between the image sample and the prediction image.
[0049] The technical solution provided in the embodiments of this specification obtains an input image containing a target object, inputs the input image into a preset image generation model, and the image generation model generates a target image based on the input image, where the target image contains the target object. The image generation model is obtained by adding a feature fusion network and a feature extraction network based on a trained initial model. The initial model includes a diffusion model based on the Transformer architecture, which is used to generate a target image based on the input target category label, and the target image contains the target object corresponding to the target category label. The image generation model generates a prediction image based on the fusion feature and is trained based on the error between the image sample and the prediction image. Among them, the fusion feature is obtained by using the feature fusion network to fuse the first image feature and the second image feature. The first image feature is extracted from the image sample by the feature extraction network, and the second image feature is extracted from the image sample by the diffusion model.
[0050] On the one hand, there is no need to expand the images in the image library by manual procurement. Instead, the trained image generation model can be directly used to quickly generate the images containing the target object required for expanding the image library, greatly reducing the cost of expanding the image library of the business platform and meeting the requirements of the rapid iteration and efficient online launch of the target object (such as goods).
[0051] On the other hand, using a pre-trained diffusion model based on the Transformer architecture as the initial model can first roughly classify the target object at a coarse-grained level, enabling the model to first understand the pixel semantic relationships between large categories and within categories from a macroscopic perspective, initially summarize the commonalities, and at the same time initially learn the differences between the images of different category target objects. This reduces the training difficulty of the image generation model during training, improves the stability of model training, and also reduces the uncertainty of image generation, making the generation of images containing the target object more controllable.
[0052] On the other hand, compared with the limitations of text descriptions, images contain richer visual features and the information presented is more intuitive. Using images as descriptions can better capture the content and details of the target images generated by the image generation model. Therefore, using only the image (image prompt) containing the target object as the input to the image generation model instead of text (text-to-image) can improve the quality of the generated target images containing the target object.
[0053] For ease of description, the training process that the initial model undergoes is hereinafter referred to as the pre-training process, and the training process of the image generation model is referred to as the secondary training process.
[0054] There can be multiple specific implementations of the pre-training that the initial model undergoes. As an example, the initial model can be trained as follows: Determine each initial target object of the image to be generated; Obtain the initial image samples containing each initial target object respectively corresponding to each initial target object; Configure a class label for each initial target object and determine the initial image samples respectively corresponding to each class label; Input the class label into the initial model so that the initial model generates an initial predicted image based on the input class label, and train the initial model according to the error between the initial predicted image and the initial image sample corresponding to the input class label.
[0055] As another example, a trained initial model can be used to generate a target image based on the input target class label, and the target image contains the target object corresponding to the target class label. The trained initial model can already generate images by directly inputting class labels without the support of complex text features. Then, the image generation model obtained by adding a feature fusion network and a feature extraction network based on the trained initial model does not need to understand complex text semantics, which not only reduces the model training difficulty and improves the model training stability during training, but also further reduces the uncertainty of image generation, making the generation of images containing the target object more controllable.
[0056] Before pre-training the initial model, it is first necessary to determine each initial target object for which images are to be generated. The initial target object can have various specific implementations. For example, if the initial target object is a commodity, this commodity can include food items, daily necessities, etc. There can be multiple ways to determine the initial target object. As an example, in an existing image database, an object that is not covered by the images in the image database can be determined as the initial target object, that is, if there are no images of a certain object in the image database, then this object can be used as the initial target object. As another example, in an existing image database, an object with the number of associated images less than a quantity threshold can be used as the initial target object. As another example, in an existing image database, an object with the quality of the associated images lower than a quality threshold can be used as the initial target object.
[0057] As an example, the existing image database can be a database maintained by the business platform to which the initial target object is to be launched, or it can be a database maintained by other platforms, and no specific restrictions are imposed on this. There are various sources of images in the existing image database. They can be images of the initial target object uploaded by merchants, or relevant images of the initial target object obtained in advance from other online platforms.
[0058] The following takes the initial target object being a food item as an example for illustration: In an existing food item image database, if all the images in the food item image database do not contain a certain food item, and this food item is expected to be launched on the business platform, then this food item can be used as the target food item, that is, the initial target object; if the number of images associated with a certain food item in the food item image database is small, such as the number of images associated with this food item being less than or equal to the quantity threshold, and this food item is expected to be launched on the business platform, then this food item can also be determined as the initial target object; if the quality of the images associated with a certain food item in the food item image database is poor, such as the quality of the images associated with this food item being lower than the quality threshold, and this food item is expected to be launched on the business platform, then this food item can also be determined as the initial target object. By determining as the initial target object an object that is expected to be launched on the business platform and for which there are no associated images meeting the criteria in the existing image database, the need for rapid iteration and efficient launch of the initial target object (such as food items) can be further met.
[0059] After determining all the initial target objects for which images need to be generated, it is necessary to obtain the initial sample data for training the initial model accordingly. The initial sample data includes various types of data, among which are the initial image samples that contain each initial target object corresponding to each initial target object. The initial image samples can be obtained through various means. Taking the initial target object as a dish as an example, as an illustration, in an existing dataset of dish images, multiple existing images associated with each initial target object can be determined, and based on the multiple existing images associated with each initial target object, the initial image sample that contains the initial target object corresponding to each initial target object can be determined.
[0060] The initial sample data for pre-training the initial model also includes class labels. After obtaining the initial image sample that contains each initial target object corresponding to each initial target object, it is necessary to configure a class label for each initial target object and determine the initial image sample corresponding to each class label respectively. There are various ways to configure class labels for each initial target object. As an example, the name of the initial target object can be a class. The name of the initial target object is text information, and the text (text) describing this class can be converted into a class label (label). The class label can be data without semantic information, for example, it can be a numerical value. Taking the initial target object as a dish as an example, the name of the dish is a class. For example, scrambled eggs with tomatoes is a class, and scrambled eggs with peppers is another class. Different class labels are configured for different dishes. For example, the class label "1" is configured for scrambled eggs with tomatoes, and the class label "2" is configured for "scrambled eggs with peppers". If there are 1000 dishes, 1000 class labels "0 to 999" can be configured. By constructing the correspondence between the dish name and the class label and using the class label as the input of the initial model, the trained initial model can directly understand the relationship between the class label and the corresponding image without having to understand the complex semantics expressed by the text describing the dish name.
[0061] It can be understood that each initial target object can uniquely correspond to a class label. Then, the initial image sample that contains the initial target object corresponding to each initial target object is the initial image sample corresponding to the class label uniquely corresponding to this initial target object.
[0062] It should be noted that the above introduction to the acquisition method of the initial sample data is only an exemplary display, and other acquisition methods are not excluded in actual applications, and specific limitations are not made in this regard.
[0063] After the initial sample data for training the initial model is determined, it can be used to train the initial model. As an example, the initial model includes a diffusion model based on the Transformer architecture. The diffusion model based on the Transformer architecture uses the Transformer architecture as the backbone network of the diffusion model to replace the traditional convolutional neural network (such as U-Net) as the backbone network, enabling it to process serialized image patches (latent patches) and capture global dependencies through the self-attention mechanism. The diffusion model based on the Transformer architecture can process the latent representation of images and usually operates in the latent space of images. This means that the image is first encoded into a lower-dimensional latent representation, and then the diffusion model is trained in this latent space to reduce the computational amount required for direct training in the high-resolution pixel space. The core of the Transformer architecture lies in the self-attention mechanism, which enables the model to process all elements in the input discrete sequence in parallel, thus greatly improving the training efficiency and effectively capturing long-range dependencies in the sequence data. The self-attention mechanism allows the model to consider all positions in the sequence simultaneously when processing information, rather than processing step by step like a recurrent neural network (RNN) or a convolutional neural network (CNN). Moreover, the Transformer architecture adopts an encoder-decoder structure, where the encoder is responsible for encoding the input sequence into a vector representation, and the decoder generates the target sequence based on the output of the encoder.
[0064] Please refer to Figure 3 , and an exemplary introduction to the architecture of training and testing (inference) of the diffusion model based on the Transformer architecture in the embodiments of this specification is as follows:
[0065] The framework of the diffusion model based on the above Transformer architecture follows the method of simulating the diffusion process (latent diffusion) in the latent space. During the training process, all images (img) can be encoded into latent representations by a Variational Autoencoder Encoder (VAE encoder). Then, according to the diffusion process, noise is added to the latent at each step until the latent completely becomes noise. After that, the above Diffusion models with transformer (DiT) are used to predict the noise at each step, and then the original latent is restored. During the inference process, a random noise is generated, the noise is predicted and removed by the diffusion model based on the Transformer architecture, and finally a new image is generated after decoding.
[0066] The above diffusion model may include multiple Blocks, that is Figure 3 the DiT in []. As the basic building unit of the diffusion model based on the Transformer architecture, the Block is a sub-module that is reused in the diffusion model and is responsible for processing the latent representation of the image. Especially in the process of gradually denoising the image data simulation, the diffusion model is constructed by stacking multiple Blocks. Each Block processes the input features and gradually improves the quality of the image, thereby generating a clear image. In addition, stacking Blocks can ensure the scalability of the diffusion model, that is, the complexity of the diffusion model can be achieved by increasing the number or width of the Blocks.
[0067] As an example, the number of Blocks included in the diffusion model can be adjusted to 28 (N = 28), or it can be adjusted to other numbers, which are not specifically limited.
[0068] Among them, the input image of the diffusion model can be encoded into a latent representation by a VAE (Variational Autoencoder). The latent representation can be divided into multiple patches, each patch is embedded as a vector, and each patch can be understood as the latent representation corresponding to each image block of the input image. In this way, a Token sequence containing multiple Tokens can be obtained, which can be called the initial Token sequence.
[0069] For example, each token can correspond to an image patch (patch_size * patch_size) in the image. Therefore, the number of tokens T is equal to the height and width of the image divided by the square of patch_size, i.e., T = (H / patch_size) * (W / patch_size), where H and W are the height and width of the image respectively. The dimension of each token is d, which is a configuration parameter of the model and represents the length of the feature vector of each token. In the DiT model, this dimension d is a hyperparameter of the model and can be adjusted according to the model design and requirements. Therefore, the dimension of the input token sequence of the Block is T * d, where T and d represent the number of tokens and the dimension of each token respectively.
[0070] The initial Token sequence can be input into the first Block among multiple Blocks. Inside the Block, the input Token sequence can be processed through self-attention mechanism, feed-forward network, etc., and a new Token sequence is output for the next Block to use. Multiple Blocks are cascaded and stacked, and the input Token sequence of each layer of Block is the new Token sequence generated and output by the Block of the previous layer. Through multi-layer processing, the model gradually extracts and fuses multi-level features of the image, and finally generates high-quality images.
[0071] In related technologies, the DiT model uses the absolute position encoding method of traditional transformers, that is, for a fixed-resolution image, the position encoding of each patch sequence is fixed. Although this is computationally simple, the disadvantages are also very obvious, making it only possible to generate images of a specified resolution size. Because the change in image resolution will cause the change in the number of image patch sequences, while the absolute position encoding does not support the change of the sequence. That is to say, at a fixed image resolution, the number of Patches is fixed, so the position encoding is also fixed. When dealing with images of different resolutions, resulting in a change in the number of Patches, the fixed absolute position encoding cannot adapt, especially when the token length is longer during testing than during training, because there is no pre-defined encoding for the new Patch positions.
[0072] Thus, the existing DiT models can only process images of one aspect ratio; for example, currently, most DiT models process images with an aspect ratio of 1:1, such as images with a resolution of 512 * 512, and some are 1024 * 1024. That is, the resolution of the input image of the model is 512 * 512, and the resolution of the generated image of the model is 512 * 512.
[0073] In order to obtain images with different aspect ratios from a model, relative position encoding is implemented in this embodiment. As an example, the Block of the DiT model in this embodiment is used to: obtain an input Token sequence, encode relative position information for each Token in the input Token sequence, and then process the Token sequence encoded with relative position information to obtain a new Token sequence for output, so as to be passed to the next Block;
[0074] Among them, in the multiple Blocks, the Tokens in the input Token sequence of the first Block include: the latent representations corresponding to the respective image patches of the image sample;
[0075] The relative position information encoded by each Token in the input Token sequence of the first Block includes: the relative position information between the image patch corresponding to the Token and other image patches.
[0076] Taking the existing DiT for processing a 512*512 input image as an example, assuming the patch_size is 2, in the Token sequence (T*d) converted based on this input image, T is 256*256. In the existing absolute position encoding scheme, absolute position information of each Token in the 512*512 input image is encoded for each Token in the Token sequence.
[0077] In this embodiment, multiple different aspect ratios can be set. Taking 3:2 and 16:9 as two examples, assume the following two resolutions are adopted: 750*500 and 800*450.
[0078] In this embodiment, the latent representation of the image is segmented into latent representations (i.e., Tokens) corresponding to multiple image patches. Specific segmentation methods such as the segmentation size and the number of segments can be configured according to actual needs.
[0079] Existing images with these two resolutions can be obtained as initial image samples. For example, for an initial image sample of 750*500, its Token sequence is still T*d. However, based on the resolution of the input initial image sample, T is no longer fixed at 256*256, but can vary based on the resolution of the input initial image sample. In this embodiment, each Token can encode the relative position information between image patches. For example, RoPE encodes the positional relationship relative to other Tokens by performing rotation operations on the query and key vectors, enabling the model to naturally consider the relative positions between image patches during self-attention calculation. Similarly, similar processing is also performed on the existing image of 800*450. In this way, the DiT model can generate images with a resolution of 750*500 and images of 800*450.
[0080] During the pre-training process of the initial model, multiple batches of training data can be used for model training for multiple different aspect ratios. For example, multiple annotated training data can be used for each aspect ratio. Each batch can contain multiple initial image samples and their corresponding class labels. The initial image samples within the same batch have the same aspect ratio, and between different batches, the initial image samples can have different aspect ratios. As an example, the number of batches for each aspect ratio is the same, and the amount of training data is also the same, so that the model can evenly obtain the image generation ability for each aspect ratio. Optionally, the initial model can be trained in the order of the training data for each batch of each aspect ratio.
[0081] To prevent the initial model from experiencing capacity forgetting, training data with various aspect ratios can be used for alternating training. For example, assume there are 3 aspect ratios. Batch 1: First, train with aspect ratio A, then with aspect ratio B, and then with aspect ratio C. The amount of training data for the 3 aspect ratios within Batch 1 is the same. Then proceed to Batch 2. Similarly to Batch 1, first train with aspect ratio A, then with aspect ratio B, and then with aspect ratio C. The amount of training data for the 3 aspect ratios within Batch 2 is the same.
[0082] After the initial model pre-training is completed, based on the various aspect ratios used for the training data during training, the initial model has the ability to generate images with these various aspect ratios. After the pre-training is completed, when using the initial model to generate images, the aspect ratio used for the latent space of the model can be set according to the required aspect ratio, that is, the VAE encoder in the initial model is set to set the aspect ratio of the latent representation generated by the VAE encoder, so that the VAE encoder of the initial model can generate images with the corresponding aspect ratio. For example, if the aspect ratio of the latent representation generated by the VAE encoder of the model is set to A, then the VAE encoder of the model can generate images with an aspect ratio of A.
[0083] As Figure 4 shown, each Block of the initial model can at least include a first Root Mean Square Layer Normalization (RMS Norm) network layer, a first Scale and Shift layer, a Multi-Head Self-Attention (MHSA) layer, a first Scale layer, a second Root Mean Square Layer Normalization network layer, a second Scale and Shift layer, a Feedforward network layer, and a second Scale layer connected in sequence.
[0084] As an example, the above diffusion model can also include a conditional injection layer. The multiple Blocks included in the diffusion model can share the same conditional injection layer. The conditional injection layer can be independently set outside the multiple Blocks to convert the conditional information input to the conditional injection layer into conditional features and input the conditional features into each Block respectively; where the conditional information is timestamp information (Timestep).
[0085] As another example, the timestamp information can represent the current time step in the model diffusion process. In the diffusion model, the image generation process starts from noise, gradually denoises, and finally generates a clear image. The timestamp information tells the model which stage it is in the diffusion process, thus helping the model adjust its behavior.
[0086] It can be understood that during the training stage (pre-training process) of the initial model (the above diffusion model), the conditional information of the conditional injection layer can include not only the timestamp information but also the above category labels. During the training process of the initial model, the conditional injection layer can convert the category labels input to the initial model into input features and input the input features into each Block respectively.
[0087] As another example, the above input features at least include the scaling parameter input to the first parameter scaling and translation layer, the translation parameter input to the first parameter scaling and translation layer, the scaling parameter input to the first parameter scaling layer, the scaling parameter input to the second parameter scaling and translation layer, the translation parameter input to the second parameter scaling and translation layer, and the scaling parameter input to the second parameter scaling layer.
[0088] It can be understood that the above parameters are associated with the class labels of the input conditional injection layer. Since the conditional injection layer is a fully connected layer with a large number of parameters, avoiding setting a conditional injection layer in each Block can avoid redundant parameter quantities and thus avoid slowing down the operation speed. By setting the conditional injection layer outside multiple Blocks and sharing the same conditional injection layer among multiple Blocks, the parameter quantity of the model can be reduced. After the initial model training is completed and a feature fusion network and a feature extraction network are added based on the trained initial model to obtain an image generation model, since the input data of the image generation model is no longer the class label and no longer requires class information for guidance, the conditional information input to the conditional injection layer no longer includes the class label at this time, and the conditional information is the timestamp information at this time. This can further reduce the parameter quantity of the model during the training process (the secondary training process after the pre-training is completed) of the obtained image generation model.
[0089] Before the above diffusion model completes pre-training, that is, during the pre-training process, each input image is encoded into a latent by the VAE encoder structure, which can be processed into serialized image patches (latent patches). Specifically, its input can be a discrete sequence (input token), and the input Token is the result of serializing and dividing the above latent. After division, each patch is a Token. If the size of the latent is I*I and the size of each serialized token is p, the number of serialized tokens T is: T=(I / p) 2 , and the dimension of the serialized Token can be expressed as: Token = T*d.
[0090] Considering that in the traditional diffusion models of related technologies, layer normalization (Layer Normalization, Layernorm layer) is generally used for normalization, which uses the mean and variance for normalization and has a large amount of calculation.
[0091] In view of the above problems, as an example, the above first root mean square normalization network layer is used to: normalize each Token by using the root mean square value of each Token input to the first root mean square normalization network layer; the second root mean square normalization network layer is used to: normalize each Token by using the root mean square value of each Token input to the second root mean square normalization network layer. RMS Norm removes the process of subtracting the mean, and only needs to calculate one statistic, the root mean square, that is, square the activation value of each neuron, calculate the average value, then take the square root, and finally normalize the activation value by dividing it by this root mean square value, thereby reducing the computational complexity to improve the running efficiency of the model. It can reduce the internal covariate shift within the model, prevent gradient explosion, thereby accelerating model training and improving performance.
[0092] The above feed-forward network layer uses a fully connected layer for further feature transformation. Considering that in the traditional diffusion model of the related art, the feed-forward network generally uses the Relu (Rectified Linear Unit) activation function, and its formula is as follows:
[0093] F out =W2·Relu(W1F input ) (1)
[0094] where, F input represents the input, F out represents the output, W1 and W2 represent two fully connected layers, and the Relu activation function is used between the two layers. However, although the Relu activation function has sparsity and can reduce overfitting to a certain extent, it is very likely to cause the neuron gradient to disappear.
[0095] In view of the above problems, as an example, the feed-forward network layer of the embodiment of the present application can still include two fully connected layers, and the two fully connected layers are connected by an activation function. The activation function does not adopt the Relu activation function, but is constructed by combining the Swish activation function and a gate structure. Among them, the gating signal of the gate structure is determined based on the Swish activation function.
[0096] As an example, the Swish activation function replaces the traditional Relu activation function, and the formula can be as follows:
[0097] F out =W2·Swish(W1F input ) (2)
[0098] The curve of the Swish activation function is smooth, and the function is differentiable at all points. Its response to negative values is relatively small, overcoming the drawback of the vanishing gradient caused by the fact that the output of the Relu activation function is always zero for some neurons. Further, a gate structure (Glu, Gated Linear Units) is introduced, and the formula is as follows.
[0099]
[0100] Among them, denotes element-wise multiplication. At this time, Swish(W1F input ) represents the activation degree of each point on the feature, and is multiplied element-wise with W d F input after the initial transformation, so as to screen which features need to pass through. This mechanism can enable the network to learn useful representations more effectively, helping to improve the generalization ability of the model. Since both the Swish activation function and the Glu gate structure are adopted simultaneously, this structure can be called: SwiGlu (such as Figure 4 shown as "SiGLU").
[0101] It can be understood that the image samples input into the image generation model during the secondary training process can come from the initial image samples of the initial target object corresponding to the class labels input into the initial model during the pre-training process, or can be image samples re-obtained by other means before the secondary training, and no specific limitation is made on this.
[0102] After the pre-training process is completed, a trained initial model is obtained. And there are multiple ways to obtain an image generation model by adding a feature fusion network and a feature extraction network based on the trained initial model. As an example, the diffusion model based on the Transformer architecture included in the initial model contains multiple Blocks, and each Block includes at least a multi-head self-attention layer and a feed-forward network layer. And the feature fusion network includes a multi-head cross-attention layer, and this multi-head cross-attention layer can be set between the multi-head self-attention layer and the feed-forward network layer to fuse the first image feature and the second image feature to obtain the above-mentioned fusion feature.
[0103] As another example, the multi-head cross-attention layer can include an output projection layer, and the output of this output projection layer is initially set to zero. To facilitate the image generation model to use the network weights obtained by the initial model after the pre-training process, the output projection layer of the multi-head cross-attention layer can be initialized to zero, effectively acting as an identity mapping and retaining the input of the subsequent layers, so as to ensure that there is no large fluctuation during the model training process and ensure the stability of the model training.
[0104] The feature extraction network can extract the above first image features in various ways. As an example, the first image features may include a first Token sequence. The feature extraction network may include an image encoding layer and a projection layer. The image encoding layer is used to convert the image samples input to the image generation model into image feature vectors, and the projection layer is used to project the image feature vectors into a first Token sequence. Among them, the dimension of the above second image features is the same as the dimension of each Token in the first Token sequence.
[0105] As another example, the above projection layer may include a linear layer and a layer normalization layer. Among them, the linear layer can be used to adjust the dimension of the input image feature training to adjust it to the target dimension. The layer normalization layer is used to normalize the features output by the linear layer to prevent the problem of gradient disappearance / explosion.
[0106] The above image encoding layer can have various specific implementations. As an example, when the target object includes dishes, the image encoding layer (image encoder) can be constructed based on the Contrastive Language-Image Pre-Training (CLIP) model, and the Contrastive Language-Image Pre-Training model can be trained based on the existing dish image dataset. The features extracted by the image encoding layer constructed using the CLIP model can effectively represent the content and high-level semantic features of the images in the catering field.
[0107] As another example, since the image encoding layer (image encoder) can only be used to convert the image samples input to the image generation model into image feature vectors, during the training process (secondary training process) of the image generation model, the parameters of the image encoding layer constructed based on the CLIP model can be in a frozen state and do not participate in the training. Freezing the parameters of the image encoding layer can retain the features learned by the image encoding layer during the training process that has passed (such as the above CLIP model trained based on the existing dish image dataset), retain the existing knowledge. At the same time, since the image encoding layer does not participate in the training, it can also save computing resources, shorten the training time of the image generation model, and improve the training efficiency.
[0108] To effectively process the image feature vectors converted from the image samples input to the image generation model by the image encoding layer, the above projection layer (Project layer) can be constructed to project the image feature vectors into a first Token sequence whose dimension of each Token is the same as the dimension of the above second image features. By converting the first image features into a Token sequence with the same dimension as the second image features, it can ensure the efficient fusion of the first image features and the second image features in the above multi-head cross-attention layer, and improve the model training efficiency.
[0109] As another example, the projection layer can project the image feature vector into a first token sequence with a specific length. For example, the specific length can be 4 or other lengths, and no specific limitation is made in this regard. As another example, the projection layer is trainable.
[0110] The diffusion model can extract the above-mentioned second image features in various ways. As an example, the second image features of the image samples extracted by the diffusion model can be the features output by the multi-head self-attention layer of the diffusion model when the image samples are processed by the diffusion model during the secondary training process.
[0111] As another example, the features output by the multi-head self-attention layer of the diffusion model are fused with the first image features in the multi-head cross-attention layer to obtain the above-mentioned fusion features, and the multi-head cross-attention layer outputs the fusion features, and the fusion features are further processed by the feed-forward network layer of the diffusion model to finally generate the predicted image. As another example, the features output by the multi-head self-attention layer of the diffusion model can be first processed by the above-mentioned first parameter scaling layer and then enter the multi-head cross-attention layer for feature fusion with the first image features to obtain the above-mentioned fusion features.
[0112] Considering that not all of the multiple existing images associated with each target object in the existing dish image dataset necessarily have sufficient quality. For example, the similarity between the existing image and the target object may be low, or the main body contained in the existing image is incomplete (i.e., the target object is not fully displayed in the image), or the existing image contains information that should not exist, such as watermark information. If all the multiple existing images associated with each target object are directly used as the image samples containing the target object and input into the image generation model for training, it may lead to a low overall quality of the image samples and a poor training effect of the image generation model.
[0113] To address the above problems, after determining the multiple existing images associated with each target object, data preprocessing can be performed first. For example, quality filtering and screening can be performed on the multiple existing images associated with each target object to ensure that the existing images used as the image samples input into the image generation model themselves have sufficient quality. As an example, when the target object includes a dish, the image samples containing the target object can be obtained in the following way: in the existing dish image dataset, determine the multiple existing images associated with each target object; determine the existing images whose quality meets the preset conditions among the multiple existing images, and based on the existing images whose quality meets the preset conditions, determine the image samples that can be input into the image generation model.
[0114] There are multiple ways to determine the existing images with quality meeting the preset conditions among multiple existing images. As an example, a text feature vector of each target object can be obtained by a Contrastive Language-Image Pre-Training (CLIP) model, and an image feature vector of each existing image associated with the target object can be obtained. Based on the text feature vector of each target object and the image feature vector of each existing image associated with the target object, the similarity between each target object and the multiple existing images associated with the target object is determined. Among the multiple existing images associated with each target object, the above-mentioned image samples are determined based on the existing images with a similarity to the target object greater than or equal to the similarity threshold.
[0115] As another example, the target object has a name, and a text feature vector of each target object is obtained. Specifically, a text feature vector can be extracted from the name of the target object.
[0116] As another example, based on a preset body integrity judgment network, it can be judged whether the main body included in the multiple existing images associated with each target object is complete. After determining the existing images with a complete main body among the multiple existing images, each target object corresponding to the image sample including the target object can be determined based on the existing images with a complete main body.
[0117] As another example, based on a preset watermark recognition network, it can be judged whether the multiple existing images associated with each target object contain watermark information. After determining the existing images without watermark information among the multiple existing images, each target object corresponding to the image sample including the target object can be determined based on the existing images without watermark information. As another example, based on a preset clarity calculation network, the clarity corresponding to each of the multiple existing images associated with each target object can be calculated, and the image samples can be determined based on the existing images with a clarity greater than or equal to the clarity threshold.
[0118] It can be understood that the existing images determined as image samples can meet at least one of the above four screening conditions (similarity to the target object greater than or equal to the similarity threshold, complete main body included, no watermark information, clarity greater than or equal to the clarity threshold), or all four of the above screening conditions need to be met, and specific limitations are not made in this regard.
[0119] It is understandable that the above four screening conditions (the similarity to the target object is greater than or equal to the similarity threshold, the contained subject is complete, there is no watermark information, and the clarity is greater than or equal to the clarity threshold) can be used not only for screening the image samples input to the image generation model during the secondary training process, but also for screening the initial image samples corresponding to the category labels input to the initial model during the pre-training process.
[0120] It is understandable that through the above CLIP model, the similarity score between the name of the target object and the image can be calculated, and this score reflects whether the image meets the description of the text describing the name of the target object. The higher the score, the better the matching degree between the two. Taking the target object as a dish as an example, the similarity score between the name of the dish and the image can be calculated, and this score reflects whether the image meets the description of the dish name text. The higher the score, the better the matching degree between the two.
[0121] There can be various specific implementations of the above-mentioned preset subject integrity judgment network. As an example, this subject integrity judgment network can be a classification network based on Vision Transformer (ViT), or other types of networks, and specific details are not limited in this regard.
[0122] There can be various specific implementations of the above-mentioned preset watermark recognition network. As an example, this watermark recognition network can be a classification network based on ViT, or other types of networks, and specific details are not limited in this regard.
[0123] There can be various specific implementations of the above-mentioned watermark information. As an example, this watermark information can be text watermark, image watermark, or logo information. Therefore, specific implementations of the watermark information are not limited.
[0124] Please refer to Figure 5 、 Figure 6 , and below, a specific structure of the image generation model in the embodiments of this specification will be introduced exemplarily:
[0125] As Figure 5 shown, the image generation model can include N Dit Blocks. Taking the structure of one Block as an example for exemplary introduction, first, Input Tokens can be used as the input of the Block, with a dimension of T*d, and the acquisition method of Input Tokens is as described above, and specific details will not be elaborated here.
[0126] Considering that before the Input Tokens enter the Block, due to the diffusion model based on the Transformer architecture, all images are encoded into latent through the VAE encoder structure during training. The Input Tokens are the result of serializing and partitioning the above latent. Each patch after partitioning can be used as a Token. Since the Transformer architecture itself does not contain the position information of each Token, additional position encoding is required to represent the position order of each input Token. In related technologies, an absolute position encoding method is used for each Token, which strictly controls the size of the input latent. The number of image blocks (latent patches) after serializing and partitioning the latent is fixed. Then, the absolute position encoding method is used to add position information to each patch. That is, for an image with a fixed resolution (such as a latent of I*I), the position encoding of each patch sequence is fixed. Although this is computationally simple, the disadvantage is also very obvious, that is, only image blocks (latent patches) of a specified resolution size can be generated. Because if the resolution of the image block changes, it will cause the number of image patch sequences to change (i.e., the number of Tokens T changes), and the above absolute position encoding method does not support the change of the sequence number.
[0127] As Figure 5 shown, to address the above problem, as an example, the Rotary Position Embedding (RoPE) can be used to replace the absolute position encoding and applied to the image generation model in the embodiments of the present application. RoPE means encoding based on the relative positions between each patch (Token), adding position information to each patch after partitioning the latent. In this way, whether the number of sequences increases or decreases, as long as the relative position relationship between two sequences is measured. Through the position encoding method of RoPE, the change of the sequence number can be supported, that is, the number of Tokens T can support dynamic changes, enabling the image generation model to support the inference of dynamic sequences.
[0128] After the Input Tokens complete the position encoding, they can be input into the Block. First, they go through the normalization process of RMS Norm, which can prevent gradient explosion.
[0129] As Figure 6 shown, for the conditional information (i.e., Figure 5 "Conditioning" inFigure 5 The Multilayer Perceptron (MLP) therein can convert conditional information into conditional features. Among them, N Blocks can share the same adaLN layer (MLP), and this adaLN layer is independently set outside the N Blocks and is used to convert the input conditional information into conditional features and input the conditional features into the N Blocks respectively. Since the adaLN layer is a fully connected layer with a large number of parameters, avoiding setting an adaLN layer in each Block can avoid redundant parameter quantities and thus avoid slowing down the operation speed. By setting the adaLN layer outside the N Blocks and sharing the same adaLN layer among the N Blocks, the parameter quantity of the model can be reduced.
[0130] The parameters included in the above conditional features can be respectively input into multiple network layers (such as Figure 5 the Scale, Shift layer and Scale layer in Figure 5 Taking a single block among the N Blocks as an example, the above conditional features at least include the scaling parameter γ1 and the translation parameter β1 input into the first parameter scaling and translation layer (i.e., the Scale, Shift layer before "MHSA" in Figure 5 ), where γ1 is used for inputting Scale (scaling) and β1 is used for inputting Shift (translation). The first parameter scaling and translation layer (i.e., the Scale, Shift layer) means scaling first and then translating.
[0131] The above conditional features also include the scaling parameter α1 input into the first parameter scaling layer (i.e., the Scale layer after "MHSA" in the figure), the scaling parameter γ2 and the translation parameter β2 input into the second parameter scaling and translation layer (i.e., the Scale, Shift layer before "SiGlu Feedforward" in Figure 5 ), where γ2 is used for inputting Scale (scaling) and β2 is used for inputting Shift (translation). The second parameter scaling and translation layer (i.e., the Scale, Shift layer) also means scaling first and then translating.
[0132] The above conditional features also include the scaling parameter α2 input into the second parameter scaling layer (i.e., the Scale layer after "SiGlu Feedforward" in the figure). Among them, γ1, β1, α1, γ2, β2 and α2 are all associated with the conditional information input into the MLP. The role of these parameters is to transform the network features while inputting the conditional information and enhance the non-linear expression ability.
[0133] For any Scale, Shift layer, its scaling and translation formulas can be as follows:
[0134] F out = F in *γ + β (4)
[0135] where γ is the scaling parameter of the input Scale, β is the translation parameter of the input Shift, and F in is the feature input to the Scale and Shift layers, and F out is the feature output after passing through the Scale and Shift layers.
[0136] For any Scale layer, its scaling formula can be as follows:
[0137] F out = F in *α (5)
[0138] where α is the scaling parameter, F in is the feature input to the Scale layer, and F out is the feature output after passing through the Scale layer.
[0139] After the Input Tokens are normalized by RMS Norm, they then undergo scaling and translation operations of the Scale and Shift layers and enter the MHSA layer. The purpose of the MHSA layer is to capture the relationships between multiple Tokens, perform global Token relationship modeling, and focus on the key features among them.
[0140] After the Tokens are processed by the MHSA layer and then undergo the scaling operation of the Scale layer, (as the second image feature) they enter the feature fusion network (i.e., Figure 5 the Multi-Head Cross-Attention (MHCA) layer in Figure 5 ), and are fused with the first image feature input to the feature fusion network by the feature extraction network (i.e., Figure 5 "Image Features" in Figure 5 ) to obtain the fused feature; where, as
[0141] shown, the feature extraction network includes an image encoding layer (imageencorder) and a projection layer (Project layer, i.e., Figure 5In the case of “SiGLU”), the purpose of the SwiGlu layer is to further extract and fuse the feature information of Tokens at different positions to enhance the non-linear and expressive capabilities of the model. After being processed by the SwiGlu layer and then through the scaling operation of the Scale layer, the dimension of the input Token remains T*d.
[0142] It should be noted that the above introduction to the specific structure of the image generation model of the present application is only an exemplary display, and other specific structures are not excluded in actual applications, and specific details are not limited here.
[0143] Based on the image generation model constructed by the construction method of the image generation model described in any of the above embodiments, the embodiments of the present specification also provide an image generation method, as Figure 7 shown, the method includes the following steps:
[0144] S701. Obtain an input image including a target object;
[0145] S702. Input the input image into a preset image generation model, and the image generation model generates a target image based on the input image.
[0146] The target image includes the target object.
[0147] The construction method of this image generation model can refer to the construction method of the image generation model described in any of the above embodiments, and specific details are not elaborated here.
[0148] The target object can have multiple specific implementations. As an example, the target object can be a dish, and the target image is an image including the dish; the target object can also be other types of objects, and specific details are not limited here.
[0149] As an example, the image generation method described in any of the above embodiments can be applied to a client. As another example, the input image including the target object can be input by the user through the client.
[0150] Optionally, multiple different aspect ratios can also be output to the user for selection. After obtaining the target aspect ratio selected by the user, set the aspect ratio of the latent representation of the VAEencoder of the image generation model, and then use the set image generation model to generate an image with the target aspect ratio.
[0151] Considering that some unreasonable results may be encountered during the inference process of the image generation model, such as the similarity between the generated target image and the target object may be low, or the main body included in the target image is incomplete (that is, the target object is not completely displayed in the target image), or the target image includes information that should not exist, such as watermark information, etc.
[0152] In view of the above problems, after the image generation model generates the target image, the target image can be further filtered to screen out higher-quality images. As an example, taking the target object as a dish, the text feature vector of the target object is obtained by the text-image relevance matching model, and the image feature vector of the generated target image containing the target object is obtained. Based on the text feature vector and the image feature vector, the similarity between the target object and the target image is determined. Then, based on the target images whose similarity is greater than or equal to the similarity threshold, the images to be stored are determined, and the images to be stored are used to be stored in a preset dish image database.
[0153] As another example, it is also possible to judge whether the main body contained in the generated target image is complete based on a preset main body integrity judgment network, and determine the images to be stored based on the target images whose contained main bodies are complete. The images to be stored are used to be stored in a preset dish image database, and / or judge whether the generated target image contains watermark information based on a preset watermark recognition network, and determine the images to be stored based on the target images that do not contain watermark information. The images to be stored are used to be stored in a preset dish image database.
[0154] It can be understood that the determined images to be stored can meet at least one of the above three screening conditions (the similarity to the target object is greater than or equal to the similarity threshold, the contained main body is complete, and no watermark information is included), or all three of the above screening conditions need to be met, and no specific limitation is made in this regard.
[0155] It can be understood that through the above text-image relevance matching model, the similarity score between the name of the target object and the target image can be calculated, and this score reflects whether the target image meets the description of the text describing the name of the target object. The higher the score, the better the matching degree between the two. Taking the target object as a dish as an example, the similarity score between the name of the dish and the target image can be calculated, and this score reflects whether the image meets the description of the dish name text. The higher the score, the better the matching degree between the two.
[0156] There can be various specific implementations of the above-mentioned preset main body integrity judgment network. As an example, the main body integrity judgment network can be a classification network based on Vision Transformer (ViT), or other types of networks, and no specific limitation is made in this regard.
[0157] There can be various specific implementations of the above-mentioned preset watermark recognition network. As an example, the watermark recognition network can be a classification network based on ViT, or other types of networks, and no specific limitation is made in this regard.
[0158] The above watermark information can have multiple specific implementations. As an example, the watermark information can be text watermark, image watermark, or logo information. Therefore, the specific implementation of the watermark information is not limited.
[0159] It should be noted that the above introduction to the screening method of the target image is only an exemplary display. In actual applications, other screening methods are not excluded, and the specific details are not limited.
[0160] As an example, if for the above three screening conditions (similarity to the target object is greater than or equal to the similarity threshold, the main body contained is complete, and no watermark information is included), the generated target images all meet the conditions, then further, the target images that meet these three screening conditions can be submitted to manual evaluation. Specifically, the quality and aesthetics of the target images can be further manually reviewed, mainly evaluating factors such as main body deformation, background deformation, image color score, composition score, and image homogenization, etc., to further screen out target images with no main body deformation, background deformation, high-quality color composition, and low repetition rate as the images to be stored and stored in the preset dish image database.
[0161] It can be understood that the above preset dish image database can serve actual business. For example, when a merchant needs to launch a dish, they can directly select the image corresponding to the dish in the dish image database through the above client.
[0162] The embodiments of this specification also provide a computer program product, including a computer program, which when executed by a processor implements the image generation method or the construction method of the image generation model described in any of the above embodiments.
[0163] The embodiments of this specification also provide an electronic device, as Figure 8 shown, the electronic device includes:
[0164] A processor 801;
[0165] A memory 802 for storing instructions executable by the processor;
[0166] Wherein, the processor 801 is configured to implement the image generation method or the construction method of the image generation model described in any of the above embodiments.
[0167] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored, which when executed by a processor implements the image generation method or the construction method of the image generation model described in any of the above embodiments.
[0168] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions in the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the objectives of the solutions in this specification. A person of ordinary skill in the art can understand and implement them without creative efforts.
[0169] The above embodiments can be applied to one or more computer devices. The computer device is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. The hardware of the computer device includes, but is not limited to, a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0170] The computer device can be any electronic product that can interact with users, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.
[0171] The computer device may also include a network device and / or a user device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing.
[0172] The network where the computer device is located includes, but is not limited to, the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.
[0173] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0174] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this patent. Making insignificant modifications to the algorithm or process or introducing insignificant designs, but without changing the core design of its algorithm and process, are all within the protection scope of this application.
[0175] Among them, the description of "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of this specification. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0176] After considering the specification and practicing the invention claimed herein, those skilled in the art will readily conceive of other embodiments of this specification. This specification is intended to cover any variations, uses, or adaptations of this specification, which follow the general principles of this specification and include the common general knowledge or conventional technical means in the technical field not claimed in this application. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of this specification are pointed out by the following claims.
[0177] It should be understood that this specification is not limited to the precise structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of this specification is only limited by the appended claims.
[0178] The above are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this specification shall be included within the protection scope of this specification.
Claims
1. A method for constructing an image generation model, comprising: Get the trained initial model; The initial model is used to generate a target image based on an input target category label, wherein the target image contains a target object corresponding to the target category label; the initial model includes a diffusion model based on a Transformer architecture; Based on the initial model, a feature fusion network and a feature extraction network are added to obtain an image generation model; An image sample containing the target object is input into the image generation model so that the feature extraction network extracts a first image feature of the image sample, and then the diffusion model extracts a second image feature of the image sample. A fusion feature of the first image feature and the second image feature is obtained using a feature fusion network, a predicted image is generated based on the fusion feature, and the image generation model is trained based on the error between the image sample and the predicted image.
2. According to the method described in claim 1, the diffusion model comprises a plurality of blocks, each block comprises at least a multi-head self-attention layer and a feedforward network layer, the feature fusion network comprises a multi-head cross-attention layer, and the multi-head cross-attention layer is arranged between the multi-head self-attention layer and the feedforward network layer to fuse the first image feature and the second image feature to obtain the fused feature.
3. According to the method of claim 1, the first image feature comprises a first Token sequence; the feature extraction network comprises an image encoding layer and a projection layer, the image encoding layer is used to convert the image sample into an image feature vector; the projection layer is used to project the image feature vector into the first Token sequence; in, The dimension of the second image feature is the same as the dimension of each Token in the first Token sequence.
4. According to the method of claim 3, during the training process of the image generation model, the parameters of the image coding layer are in a frozen state.
5. According to the method of claim 3, the target object includes dishes; the image encoding layer is constructed based on a picture-text correlation matching model, and the picture-text correlation matching model is trained based on an existing dish image data set.
6. The method according to claim 1, wherein the target object comprises a dish; and the image sample containing the target object is obtained by: In an existing dish image dataset, determining a plurality of existing images associated with each of the target objects; An existing image whose quality meets a preset condition is determined among the multiple existing images, and the image sample is determined based on the existing image whose quality meets the preset condition.
7. A method for generating an image, comprising: Get an input image containing the target object; Inputting the input image into a preset image generation model, and generating a target image based on the input image by the image generation model, wherein the target image includes the target object; Wherein, the image generation model is constructed using the method described in any one of claims 1 to 6.
8. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
9. An electronic device, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the steps of the method according to any one of claims 1 to 6 when executed by a processor.
Citation Information
Cited By
Diffusion Transform reasoning acceleration method and device based on distributed perception sparsity
CN121787559A
A method and apparatus for accelerating inference based on distribution-aware sparse Diffusion Transformer
CN121787559B