Image generation method and device, electronic equipment and medium
By acquiring image prompt text and object box data, determining instance semantic features and target box mask features, and denoising processing with diffusion model, the problem of insufficient accuracy and detailed performance of multi-instance image generation in the prior art is solved, and high-quality multi-instance image generation is achieved.
Patent Information
- Application Number
- CN202412000116.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
When generating multiple instance images, the prior art is difficult to accurately convey user intentions, resulting in a gap between the generated image and the desired result, and it is insufficient in generating complex fine-grained features, limiting its application in scenarios requiring high precision.
By obtaining the image prompt text and its corresponding object box data, the instance semantic features are determined, and the target box mask features are generated based on the object box data, and the noise graph is denoised in combination with the diffusion model to generate high-quality multi-instance images.
It improves the positioning accuracy of each instance in the generated image, enhances the detailed performance and visual effects of the image, and meets the user's needs for details and complexity.
Smart Images

Figure CN119941892A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image generation technology, and in particular to an image generation method, device, electronic device and medium. Background Art
[0002] With the rapid development of Internet technology, diffusion models have made remarkable progress in the field of text-to-image (T2I) synthesis. These models excel in generating high-quality images of a single instance and can effectively transform natural language descriptions into visual content. However, when it comes to generating images containing multiple instances, these models face a series of challenges. First, the ambiguity in natural language descriptions often hinders the accurate transmission of user intent. This ambiguity leads to imprecise instance localization, making it difficult to ensure that the generated image can accurately reflect the user's needs.
[0003] For example, when a user wants to generate a scene containing multiple people or objects, ambiguity in the description may cause the generated image to be different from the expected result. Secondly, current diffusion models usually rely on global cues to describe the features of multiple instances, which makes it difficult to correctly associate these features with the corresponding instances. The global description often fails to provide enough details, resulting in the lack of clear distinction and positioning between instances in the generated images. In addition, although precise control of instance-level positions is achieved through bounding boxes as spatial signals in layout-to-image (L2I) tasks, most L2I methods are still insufficient in generating complex and fine-grained features. This deficiency limits its application in scenarios that require high precision, such as in art creation, product design, and virtual reality. Therefore, in response to the above challenges, a new method is urgently needed to improve the accuracy and detail expression of multi-instance image generation to meet the needs of users in diverse scenarios. Summary of the invention
[0004] In view of this, the embodiments of the present disclosure provide an image generation method, device, electronic device, and computer-readable storage medium to solve the technical problem that most L2I methods in the prior art are still insufficient in generating complex fine-grained features, which limits their application in scenes requiring high precision.
[0005] According to a first aspect of an embodiment of the present disclosure, an image generation method is provided, including: obtaining at least one image prompt text and object frame data corresponding to each image prompt text; determining at least one instance semantic feature according to at least one image prompt text and the object frame data corresponding to each image prompt text; generating a target frame mask feature corresponding to each object frame data based on the object frame data corresponding to each image prompt text; and denoising a noise map based on at least one instance semantic feature and the target frame mask feature corresponding to each object frame data by a diffusion model to obtain a target image.
[0006] According to a second aspect of an embodiment of the present disclosure, an image generating device is provided, comprising: an acquiring module for acquiring at least one image prompt text and object frame data corresponding to each image prompt text; a determining module for determining at least one instance semantic feature based on at least one image prompt text and object frame data corresponding to each image prompt text; a generating module for generating target frame mask features corresponding to each object frame data based on the object frame data corresponding to each image prompt text; and a denoising module for denoising a noise map based on at least one instance semantic feature and target frame mask features corresponding to each object frame data through a diffusion model to obtain a target image.
[0007] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0008] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0009] Compared with the prior art, the disclosed embodiment has the following beneficial effects: the disclosed embodiment can convey the user's intention more accurately by acquiring at least one image prompt text and its corresponding object frame data, thereby effectively improving the positioning accuracy of each instance in the generated image. Clear object frame data enables the model to better understand the location and characteristics of each instance. By determining at least one instance semantic feature, the generated image can better reflect the details and features of each instance. The enhancement of this semantic feature helps to generate higher quality multi-instance images to meet the user's needs for details and complexity. The target frame mask feature is generated based on the object frame data, so that the model can clearly distinguish the areas of different instances. This mask feature provides important spatial information for subsequent image generation, ensuring the clarity and independence between instances. The noise map is denoised by a diffusion model, and a high-quality target image can be generated by combining the instance semantic features and the target frame mask features. The denoising process effectively reduces the noise in the generated image and improves the visual effect and detail performance of the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0011] Figure 1 A schematic diagram showing an exemplary system architecture to which the technical solution of an embodiment of the present invention can be applied;
[0012] Figure 2 is a flowchart of an image generation method provided by an embodiment of the present disclosure;
[0013] Figure 3 is a schematic diagram of a framework of a semantic feature generation module provided in an embodiment of the present disclosure;
[0014] Figure 4 is a schematic diagram of a framework of a diffusion module provided in an embodiment of the present disclosure;
[0015] Figure 5 is a block diagram of an image generating device provided by an embodiment of the present disclosure;
[0016] Figure 6 It is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0017] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present disclosure. However, it should be clear to those skilled in the art that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present disclosure with unnecessary details.
[0018] It should be noted that the user information (including but not limited to terminal device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0019] Figure 1 A schematic diagram showing an exemplary system architecture to which the technical solution of an embodiment of the present invention can be applied.
[0020] like Figure 1 As shown, the system architecture 100 may include one or more of a first terminal device 101, a second terminal device 102, and a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0021] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. According to the implementation requirements, there may be any number of terminal devices, networks and servers. For example, the server 105 may be a server cluster composed of multiple servers.
[0022] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send data, etc. The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with display screens, including but not limited to smart phones, tablet computers, portable computers, desktop computers, and the like.
[0023] The server 105 may be a server that provides various services. For example, the server 105 may obtain at least one image prompt text and object frame data corresponding to each image prompt text from the first terminal device 103 (or the second terminal device 102 or the third terminal device 103); determine at least one instance semantic feature based on at least one image prompt text and object frame data corresponding to each image prompt text; generate target frame mask features corresponding to each object frame data based on the object frame data corresponding to each image prompt text; perform denoising on the noise map based on at least one instance semantic feature and target frame mask features corresponding to each object frame data through a diffusion model to obtain a target image. In the process of generating multiple instance images, the present disclosure significantly improves the accuracy and detail performance of the generated image through precise positioning, enhanced semantic expression and high-quality denoising, providing users with a more satisfactory image synthesis experience.
[0024] In some embodiments, the image generation method provided in the embodiments of the present invention is generally executed by the server 105, and accordingly, the image generation device is generally disposed in the server 105. In other embodiments, some terminal devices may have functions similar to those of the server to execute the method. Therefore, the image generation method provided in the embodiments of the present invention is not limited to being executed on the server side.
[0025] The image generation method and device according to the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0026] Figure 2 is a flowchart of an image generation method provided by an embodiment of the present disclosure. The method provided by an embodiment of the present disclosure can be executed by any electronic device with computer processing capabilities, for example, the electronic device can be Figure 1 The server is shown.
[0027] like Figure 2 As shown, the image generating method includes steps S210 to S240.
[0028] In step S210, at least one image prompt text and object frame data corresponding to each image prompt text are obtained.
[0029] Step S220: determining at least one instance semantic feature according to at least one image prompt text and object frame data corresponding to each image prompt text.
[0030] In step S230, based on the object frame data corresponding to each image prompt text, a target frame mask feature corresponding to each object frame data is generated.
[0031] In step S240 , based on the object frame data corresponding to each image prompt text, a target frame mask feature corresponding to each object frame data is generated.
[0032] The method can convey the user's intention more accurately by obtaining at least one image prompt text and its corresponding object frame data, thereby effectively improving the positioning accuracy of each instance in the generated image. The clear object frame data enables the model to better understand the location and characteristics of each instance. By determining at least one instance semantic feature, the generated image can better reflect the details and characteristics of each instance. The enhancement of this semantic feature helps to generate higher quality multi-instance images to meet the user's needs for details and complexity. The target frame mask feature is generated based on the object frame data, so that the model can clearly distinguish the areas of different instances. This mask feature provides important spatial information for subsequent image generation and ensures the clarity and independence between instances. The noise map is denoised by the diffusion model, and the instance semantic features and the target frame mask features are combined to generate a high-quality target image. The denoising process effectively reduces the noise in the generated image and improves the visual effect and detail performance of the image.
[0033] In some embodiments of the present disclosure, at least one image hint text and the object frame data corresponding to each image hint text are key components for achieving high-quality generation of at least one instance image. The image hint text is a description content written by the user according to his actual needs, which is intended to clearly indicate the characteristics, attributes and relationships of each object in the generated image. It can be a sentence or phrase in natural language, describing the type, color, shape, action, emotion, etc. of the object. For example, the hint text entered by the user includes "a girl in a red skirt playing in the park", "a yellow cat sitting on the windowsill", "a table with fruits and vases", etc. These descriptions not only cover the basic information of the object, but also can include descriptions about the relationship between objects, such as "the girl is next to the cat" or "the fruit is on the left side of the table". Through the image hint text, users can clearly express their creative intentions and expectations, helping the model understand the details that need to be paid attention to when generating images. This method also allows users to flexibly adjust according to different needs, so as to generate images that meet specific scenes or themes. The object frame data is the annotation information provided by the user for the position layout of each object in the generated image. It is usually expressed in the form of a bounding box, which defines the specific position and size of each object in the image. The object box data provides important spatial information for the model, so that the generated image can accurately reflect the object layout expected by the user. This precise location information not only helps to identify and distinguish instances, but also enhances the overall structure and logic of the image. By combining the image prompt text and object box data provided by the user, the user's intentions and needs can be effectively captured, thereby achieving higher accuracy and detail when generating multi-instance images. This approach not only improves the user experience, but also provides richer and clearer input information for the image generation model.
[0034] In some embodiments of the present disclosure, determining at least one instance semantic feature based on at least one image prompt text and object frame data corresponding to each image prompt text includes: processing at least one image prompt text through a CLIP text encoder to obtain a first semantic feature corresponding to each image prompt text; processing at least one image prompt text through a T5 text encoder to obtain a second semantic feature corresponding to each image prompt text; fusing the first semantic feature corresponding to each image prompt text, the second semantic feature corresponding to each image prompt text, and a preset learnable query feature to obtain a target semantic feature corresponding to each image prompt text; processing the object frame data corresponding to each image prompt text to obtain a target frame feature of each object frame data; and splicing the target semantic feature corresponding to each image prompt text and the target frame feature of each object frame data to obtain at least one instance semantic feature.
[0035] At least one image prompt text is processed using a CLIP (Contrastive Language-Image Pretraining) text encoder. The CLIP text encoder can convert text information into a high-dimensional vector representation and capture the semantic features of the text, that is, obtain the first semantic features corresponding to each image prompt text. These features can reflect the keywords, context and overall intention in the text description. Similarly, at least one image prompt text is processed using a T5 (Text-to-Text Transfer Transformer) text encoder. The T5 text encoder takes text as input and outputs corresponding text features, which are suitable for a variety of natural language processing tasks. The second semantic features corresponding to each image prompt text are obtained through the T5 text encoder. These features usually emphasize the structure and contextual relationship of the text. The first semantic features, the second semantic features and the preset learnable query features corresponding to each image prompt text are fused. This process can be achieved by a variety of methods, such as weighted averaging, splicing or feature fusion through a neural network. The goal is to obtain the target semantic features corresponding to each image prompt text, which integrate the information from different encoders and can more comprehensively express the semantics of the text. The object box data corresponding to each image prompt text is processed to extract the target box features of each object box data. This process can encode the geometric features of the frame (such as position, size, shape), or use a convolutional neural network (CNN) to analyze the content in the object frame. The target frame features provide spatial information and specific details of the object for subsequent image generation. The target semantic features corresponding to each image prompt text are spliced with the target frame features of each object frame data. The splicing method can be a simple vector connection to form a higher-dimensional feature representation. This splicing process enables the generation model to simultaneously consider the semantic information of the text and the spatial layout of the object, thereby forming at least one instance semantic feature. Through the above steps, the present invention can extract rich instance semantic features from the image prompt text and object frame data provided by the user. These features not only contain the semantic information of the text, but also combine the spatial features of the object, providing a solid foundation for subsequent image generation. This method effectively improves the accuracy and detail performance of the generated image, so that the generated image can better meet the expectations and needs of the user.
[0036] In some embodiments of the present disclosure, the first semantic feature corresponding to each image prompt text, the second semantic feature corresponding to each image prompt text, and the preset learnable query feature are fused to obtain the target semantic feature corresponding to each image prompt text, including: fusing the second semantic feature corresponding to each image prompt text and the preset learnable query feature through a cross-attention mechanism to obtain the first fused feature; fusing the first semantic feature corresponding to each image prompt text and the first fused feature through a cross-attention mechanism to obtain the target semantic feature corresponding to each image prompt text. For example, through an effective feature fusion mechanism, semantic features from different sources are integrated into a unified target semantic feature to improve the quality and accuracy of image generation. Feature fusion is to combine features from different models or different processing stages to obtain richer and more comprehensive information in subsequent tasks. In this embodiment, feature fusion mainly involves the following three parts: the first semantic feature corresponding to each image prompt text, the second semantic feature corresponding to each image prompt text, and the preset learnable query feature. The cross-attention mechanism is used to fuse the second semantic feature corresponding to each image prompt text and the preset learnable query feature. The cross-attention mechanism allows the model to focus on relevant information of other inputs while processing a specific input. In this process, the second semantic feature is used as the "value" and the query feature is used as the "query".
[0037] (Query), the cross-attention mechanism calculates the correlation between them and generates a weighted representation, namely the first fused feature. The first fused feature can effectively integrate the information in the second semantic feature and the contextual information in the query feature, thereby providing a more refined representation for subsequent feature fusion. Next, the first semantic feature and the first fused feature corresponding to each image prompt text are further fused using the cross-attention mechanism. In this stage, the first semantic feature is used as the "value" and the first fused feature is used as the "query". By calculating the attention weights between them, the model can extract the information most relevant to the target generation task. The target semantic features corresponding to each image prompt text obtained through this process contain rich information from the first and second semantic features, as well as the contextual guidance provided by the query feature. This feature representation not only considers the specific content of the text, but also combines the requirements of the generation task. Through the two applications of the cross-attention mechanism, semantic features from different sources can be effectively fused to generate a comprehensive target semantic feature. This method makes the generation model more accurate in understanding user intent and can better capture the details and contextual relationships required in image generation, thereby improving the quality and conformity of the generated image.
[0038] In some embodiments of the present disclosure, processing the object frame data corresponding to each image prompt text to obtain the target frame features of each object frame data includes: processing the object frame data corresponding to each image prompt text by Fourier algorithm to obtain the spatial information of each object frame; processing the spatial information of each object frame by multi-layer perceptron to obtain the target frame features of each object frame data. For example, the object frame data generally refers to the position of the target object in the image and its related information. Each object frame includes the following information: the position information is generally represented by coordinates (such as the coordinates of the upper left corner and the lower right corner). The size information is the width and height of the object frame. The category information refers to the category of the object contained in the frame (such as people, animals, objects, etc.). The object frame data corresponding to each image prompt text is processed by Fourier algorithm to extract the spatial information of the object frame. Fourier transform is a mathematical tool that can convert time domain signals into frequency domain signals to analyze the frequency components of the signal. In the context of image processing, Fourier transform can help identify the shape features, edge information and position features of the object frame. By converting the object frame data into frequency domain representation, the model can capture the distribution characteristics of the object in space. The processed spatial information usually includes frequency domain features of the object box, which can reflect information such as the geometric shape, position change and relative size of the object. The extracted spatial information is input into a multi-layer perceptron for further processing. The multi-layer perceptron is a feedforward neural network that can learn complex feature representations through multiple layers of nonlinear transformations. In this process, the MLP maps the spatial information to a new feature space and generates target box features for each object box data. Through training, the MLP can learn how to transform spatial information into features that are useful for subsequent tasks (such as image generation). The final target box features will contain the spatial features, geometric information, and category features of the object box. These features provide the necessary spatial context for subsequent image generation, ensuring that the generated image can accurately reflect the actual position and morphology of the object. Through the combination of Fourier algorithm and multi-layer perceptron, rich target box features can be extracted from the object box data. These features not only capture the spatial information of the object, but also provide important contextual support for the image generation process, thereby improving the quality and accuracy of the generated image.
[0039] In some embodiments of the present disclosure, the noise map is denoised based on at least one instance semantic feature and a target frame mask feature corresponding to each object frame data by a diffusion model to obtain a target image, including: fusing at least one instance semantic feature, a target frame mask feature corresponding to each object frame data, and the noise map to obtain a global feature of at least one instance; performing gated semantic fusion processing on at least one instance global feature to obtain a gated semantic fusion feature of at least one instance; and generating a target image based on the noise map and the gated semantic fusion feature of at least one instance. For example, a diffusion model is a generative model that generates data from random noise by stepwise denoising. In image generation tasks, the diffusion model can effectively capture the distribution characteristics of data, thereby generating images with high quality and high consistency. The instance semantic feature represents the semantic information of each instance in the image, and generally includes the category, attribute, and context information of the object. The target frame mask feature is a mask feature generated based on the object frame data, indicating the position and shape of the object in the image. The initial random noise map of the noise map is used as the starting input of the diffusion model. The semantic feature of at least one instance, the target frame mask feature corresponding to each object frame data, and the noise map are fused. The fusion method can be simple concatenation, weighted summation, or processing through a more complex neural network structure. The fusion process aims to combine the advantages of different features to generate a comprehensive global feature representation that can better reflect the overall information and details of the image. The obtained global feature of at least one instance is subjected to gated semantic fusion processing. The gating mechanism selectively retains or suppresses certain features by setting weights, thereby enhancing the expression of important information. In specific implementation, a gating unit (such as the gating structure in LSTM) can be used to dynamically adjust the contribution of features so that the model can flexibly respond to different feature combinations in different situations. The features after gating are called gated semantic fusion features, which can more effectively capture the key information of the instance and provide stronger semantic support for subsequent image generation. Based on the noise map and the gated semantic fusion features of at least one instance, the diffusion model gradually generates the target image through multiple iterative denoising processes. In each iteration, the gated semantic fusion features are used to guide the denoising process, thereby ensuring that the generated image not only conforms to the structural characteristics of the noise map, but also maintains the semantic information of the instance. The target image finally generated will reflect the consistency with the input instance semantic features and target frame mask features, with high-quality details and accurate object positions, and can effectively meet the needs of users. Through the denoising process of combining the diffusion model with the instance semantic features and the target frame mask features, the present invention can generate high-quality target images. The application of feature fusion and gating mechanism not only improves the model's ability to capture key information, but also enhances the semantic consistency and detail expression of the generated image.
[0040] In some embodiments of the present disclosure, fusing at least one instance semantic feature, target frame mask features corresponding to each object frame data, and noise map to obtain the global feature of at least one instance includes: performing self-attention processing on the noise map to obtain the target feature of the noise map; performing mask cross-attention processing on the target feature of the noise map, the target frame mask features corresponding to each object frame data, and at least one instance semantic feature through a mask cross-attention mechanism to obtain the global feature of at least one instance. For example, the self-attention mechanism is a mechanism that enables a model to pay attention to different parts of the input when processing input data. In image generation, self-attention can help the model capture the relationship between different regions in the image. The purpose of performing self-attention processing on the noise map is to extract the target features in the noise map, that is, to identify potential structures and patterns from random noise. In the self-attention processing, the noise map is first mapped to the query, key, and value space. By calculating the similarity between the query and the key, it is possible to determine which parts of the information are important and weight the corresponding values accordingly. After this process, the target features of the obtained noise map will contain potential structural information in the image, providing a basis for subsequent feature fusion. The masked cross-attention mechanism is an extension of the self-attention mechanism, which aims to combine feature information from different sources (such as noise map target features, target box mask features, and instance semantic features). In this process, the target box mask features are used to indicate which areas are important, thereby guiding the model to focus on these areas when calculating attention. The target features of the noise map, the target box mask features corresponding to each object box data, and the semantic features of at least one instance are input into the masked cross-attention mechanism for processing. Specifically, based on the mask features, the weights of the noise map target features are adjusted so that when generating global features, the influence of important areas is amplified, while the influence of unimportant areas is suppressed. After the masked cross-attention process, the global features obtained will comprehensively consider the potential structure in the noise map, the location information of the target box, and the semantic information of the instance. This global feature can more comprehensively reflect the characteristics of the target instance, ensuring that the subsequently generated images can accurately express these features. Through self-attention processing and masked cross-attention mechanism, the noise map, target box mask features, and instance semantic features can be effectively fused to generate high-quality global features. This process not only improves the accuracy of feature extraction, but also enhances the semantic consistency and detail expression of the generated images, laying a solid foundation for subsequent image generation tasks.
[0041] In some embodiments of the present disclosure, based on the noise map and the gated semantic fusion features of at least one instance, generating a target image includes: performing self-attention processing on the noise map to obtain the target features of the noise map; performing cross-attention fusion processing on the gated semantic fusion features of at least one instance and the target features of the noise map to obtain the cross-attention fusion features of at least one instance; adding the target features of the noise map and the cross-attention fusion features of at least one instance through a residual mechanism to obtain the residual fusion features of at least one instance; and processing the residual fusion features of at least one instance through a feedforward network to obtain the target image. For example, the self-attention mechanism is used to capture long-distance dependencies in the noise map and help the model identify important features in the image. When performing self-attention processing on the noise map, the noise map is first mapped to the query, key, and value space. By calculating the similarity between the query and the key, the importance of which areas in the noise map can be determined, thereby weighting the corresponding values and extracting the target features of the noise map. After the self-attention processing, the target features of the noise map obtained contain the potential structure and pattern in the image, providing a basis for subsequent feature fusion. The cross-attention mechanism aims to combine the target features of the noise map and the gated semantic fusion features to generate a richer feature representation. In this process, the gated semantic fusion features interact with the target features of the noise map, and the fusion mode of the features is determined by calculating the attention weight between the two. Through the cross-attention fusion process, the cross-attention fusion features of at least one instance can be generated. These features integrate the structural information of the noise map and the semantic information of the instance, and can better reflect the characteristics of the target instance. The residual mechanism is an effective technique in deep learning to alleviate the gradient vanishing problem in deep networks. It maintains the flow of information by adding the input features to the transformed features. The target features of the noise map are added to the cross-attention fusion features of at least one instance to obtain the residual fusion features of at least one instance. This process ensures the effective combination of the target features and the fusion features, and enhances the model's ability to capture details. The feedforward network is a network structure that performs nonlinear transformations on the input features, usually consisting of multiple fully connected layers. It can further extract and optimize the feature representation. The residual fusion features of at least one instance are input into the feedforward network for processing. The feedforward network generates the final target image through a series of activation functions and linear transformations. This process ensures that the generated image can fully reflect the fusion effect of the noise map and instance features, and the final output image has high-quality details and accurate semantic information. Through the combination of self-attention processing, cross-attention fusion, residual mechanism and feedforward network, the target image can be effectively generated from the noise map and gated semantic fusion features. This process not only improves the effect of feature extraction and fusion, but also ensures the quality and consistency of the generated image, providing strong support for image generation tasks.
[0042] In some embodiments of the present disclosure, the image generation method is applied to an image generation model, and the image generation model includes a semantic feature generation module and a diffusion model. The semantic feature generation module may include a CLIP text encoder, a T5 text encoder, and a multi-layer perceptron (MLP). The diffusion model may include multiple submodules. Figure 3 , is a schematic diagram of the processing flow of a single image prompt text and the corresponding object frame data. The details are as follows:
[0043] 1) Input the "text description (i.e. the above-mentioned image prompt text)" into the CLIP text encoder to obtain text semantic features, which are used to capture the semantic information of the text and obtain the first semantic features corresponding to the image prompt text.
[0044] 2) Input the “text description” into the T5 text encoder to obtain text features. This is used to extract the deep features of the text, including grammatical and semantic information, and obtain the second semantic features corresponding to the image prompt text.
[0045] 3) Taking the learnable query as input, it is fused with the second semantic feature corresponding to the image prompt text through a cross-attention method to obtain the first fused feature.
[0046] The first fused feature is fused with the first semantic feature corresponding to the image prompt text through cross attention, and the target semantic feature corresponding to the image prompt text is obtained to enhance the model's understanding of the text description.
[0047] 4) Use the Fourier algorithm to process the target frame (i.e., the object frame data corresponding to the image prompt text) to capture the spatial information of the target frame and obtain the spatial information of the object frame. Then input it into the MLP (multi-layer perceptron) to obtain the target frame features of the object frame data.
[0048] 5) Feature integration output:
[0049] The fusion features are added to the target box features to integrate text and spatial information; the results are concatenated with the text semantic features to generate the final instance semantic feature information as the output of this module.
[0050] refer to Figure 4 The diffusion model is composed of multiple identical sub-modules in series (the algorithm flowchart only shows the specific structure of a single module), as follows:
[0051] 1) Input the output of the previous module into the self-attention module, so that the model can capture the long-distance dependencies between features and enhance the expressiveness of features.
[0052] 2) Convert the target box of each instance into a mask, take the semantic features of multiple instances as Q, take the output of the self-attention module as K and V, combine the mask of each instance, calculate the mask cross attention, and obtain the comprehensive feature O of each instance to avoid the influence of irrelevant features on each instance
[0053] 3) The comprehensive features of multiple instances obtained in step 2 are integrated into the network through the proposed gated semantic fusion module. Gated semantic fusion dynamically adjusts the weight of feature fusion according to the importance and spatial relationship of the instance to improve the model's ability to understand the scene. The following is the process of gated semantic fusion: The comprehensive features of each instance are input into a trainable lightweight network f (such as convolution + MLP) to learn the importance S of each instance. Use the Softmax operation to normalize the importance of different instances and obtain their respective feature importance weights W. i (), i represents the i-th instance. For each instance, the gating weight G is calculated based on its area size and the IoU value with other instances. i The gate weight G i Determined by:
[0054]
[0055] Among them, A i is the area size of the i-th instance, IoU i,j is the intersection-over-union ratio of the i-th instance to the j-th instance. These weights reflect the importance and spatial relationship of the instances. e represents the base of the natural logarithm.
[0056] Using the gate weight G i and feature importance weight W i , for each instance feature F i Perform weighted fusion to generate a comprehensive feature representation.
[0057] The fusion formula can be expressed as:
[0058] Among them, G i is the gating weight of the ith instance, F i is the feature of the i-th instance, W i is the feature importance weight of the ith instance.
[0059] 4) The features after gated semantic fusion are taken as K and V, and the features output by self-attention are taken as Q, and they are fused through cross-attention.
[0060] 5) The output of the self-attention module is added and fused with the features of the previous step through the residual network to learn deeper features and prevent the gradient vanishing problem in deep networks.
[0061] 6) Input the features of the previous step into the feedforward network as the output of this module.
[0062] Figure 5 It is a block diagram of an image generating device provided by an embodiment of the present disclosure.
[0063] like Figure 5 As shown, the image generating device 500 includes an acquiring module 510 , a determining module 520 , a generating module 530 and a denoising module 540 .
[0064] Specifically, the acquisition module 510 is used to acquire at least one image prompt text and object frame data corresponding to each image prompt text.
[0065] The determination module 520 is used to determine at least one instance semantic feature according to at least one image prompt text and the object frame data corresponding to each image prompt text.
[0066] The generating module 530 is used to generate target frame mask features corresponding to each object frame data based on the object frame data corresponding to each image prompt text.
[0067] The denoising module 540 is used to perform denoising processing on the noise image based on at least one instance semantic feature and a target frame mask feature corresponding to each object frame data through a diffusion model to obtain a target image.
[0068] The image generation device 500 can convey the user's intention more accurately by acquiring at least one image prompt text and its corresponding object frame data, thereby effectively improving the positioning accuracy of each instance in the generated image. Clear object frame data enables the model to better understand the location and characteristics of each instance. By determining at least one instance semantic feature, the generated image can better reflect the details and features of each instance. The enhancement of this semantic feature helps to generate higher quality multi-instance images to meet the user's needs for details and complexity. The target frame mask feature is generated based on the object frame data, so that the model can clearly distinguish the areas of different instances. This mask feature provides important spatial information for subsequent image generation and ensures the clarity and independence between instances. The noise map is denoised by the diffusion model, and a high-quality target image can be generated by combining the instance semantic features and the target frame mask features. The denoising process effectively reduces the noise in the generated image and improves the visual effect and detail performance of the image.
[0069] In some embodiments of the present disclosure, the determination module 520 is configured to: process at least one image prompt text through a CLIP text encoder to obtain a first semantic feature corresponding to each image prompt text; process at least one image prompt text through a T5 text encoder to obtain a second semantic feature corresponding to each image prompt text; fuse the first semantic feature corresponding to each image prompt text, the second semantic feature corresponding to each image prompt text, and a preset learnable query feature to obtain a target semantic feature corresponding to each image prompt text; process the object frame data corresponding to each image prompt text to obtain a target frame feature of each object frame data; and splice the target semantic feature corresponding to each image prompt text and the target frame feature of each object frame data to obtain at least one instance semantic feature.
[0070] In some embodiments of the present disclosure, the first semantic feature corresponding to each image prompt text, the second semantic feature corresponding to each image prompt text, and a preset learnable query feature are fused to obtain a target semantic feature corresponding to each image prompt text, including: fusing the second semantic feature corresponding to each image prompt text and the preset learnable query feature through a cross-attention mechanism to obtain a first fused feature; fusing the first semantic feature corresponding to each image prompt text and the first fused feature through a cross-attention mechanism to obtain a target semantic feature corresponding to each image prompt text.
[0071] In some embodiments of the present disclosure, object frame data corresponding to each image prompt text are processed to obtain target frame features of each object frame data, including: processing the object frame data corresponding to each image prompt text by a Fourier algorithm to obtain spatial information of each object frame; processing the spatial information of each object frame by a multi-layer perceptron to obtain target frame features of each object frame data.
[0072] In some embodiments of the present disclosure, the denoising module 540 is configured to: fuse at least one instance semantic feature, target frame mask features corresponding to each object frame data, and a noise map to obtain a global feature of at least one instance; perform gated semantic fusion processing on at least one instance global feature to obtain a gated semantic fusion feature of at least one instance; and generate a target image based on the noise map and the gated semantic fusion feature of at least one instance.
[0073] In some embodiments of the present disclosure, at least one instance semantic feature, target frame mask features corresponding to each object frame data, and a noise map are fused to obtain a global feature of the at least one instance, including: performing self-attention processing on the noise map to obtain a target feature of the noise map; performing mask cross-attention processing on the target feature of the noise map, target frame mask features corresponding to each object frame data, and at least one instance semantic feature through a mask cross-attention mechanism to obtain a global feature of the at least one instance.
[0074] In some embodiments of the present disclosure, generating a target image based on a noise map and gated semantic fusion features of at least one instance includes: performing self-attention processing on the noise map to obtain target features of the noise map; performing cross-attention fusion processing on the gated semantic fusion features of at least one instance and the target features of the noise map to obtain cross-attention fusion features of at least one instance; adding the target features of the noise map and the cross-attention fusion features of at least one instance through a residual mechanism to obtain residual fusion features of at least one instance; processing the residual fusion features of at least one instance through a feedforward network to obtain a target image.
[0075] Figure 6 Schematic diagram of an electronic device 6 provided in an embodiment of the present disclosure. Figure 6 As shown, the electronic device 6 of this embodiment includes: a processor 601, a memory 602, and a computer program 603 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the computer program 603, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 601 executes the computer program 603, the functions of the modules in the above-mentioned device embodiments are implemented.
[0076] The electronic device 6 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 6 may include, but is not limited to, a processor 601 and a memory 602. Those skilled in the art will appreciate that Figure 6 The electronic device 6 is merely an example and does not limit the electronic device 6 . The electronic device 6 may include more or less components than those shown in the figure, or different components.
[0077] The processor 601 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0078] The memory 602 may be an internal storage unit of the electronic device 6, for example, a hard disk or memory of the electronic device 6. The memory 602 may also be an external storage device of the electronic device 6, for example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 6. The memory 602 may also include both an internal storage unit of the electronic device 6 and an external storage device. The memory 602 is used to store computer programs and other programs and data required by the electronic device.
[0079] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units.
[0080] If the integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present disclosure implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. The computer program may include computer program code, and the computer program code may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electric carrier signals and telecommunication signals.
[0081] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.
Claims
1. An image generation method, characterized in that: include: Obtain at least one image prompt text and object frame data corresponding to each image prompt text; Determining at least one instance semantic feature according to at least one image prompt text and object frame data corresponding to each image prompt text; Based on the object frame data corresponding to each image prompt text, generate the target frame mask feature corresponding to each object frame data; The noise map is denoised by using a diffusion model based on at least one instance semantic feature and a target frame mask feature corresponding to each object frame data to obtain a target image.
2. The method according to claim 1, characterized in that Determining at least one instance semantic feature according to at least one image prompt text and object frame data corresponding to each image prompt text includes: Processing at least one image prompt text by using a CLIP text encoder to obtain a first semantic feature corresponding to each image prompt text; Processing at least one image prompt text by using a T5 text encoder to obtain a second semantic feature corresponding to each image prompt text; The first semantic feature corresponding to each image prompt text, the second semantic feature corresponding to each image prompt text, and the preset learnable query feature are fused to obtain the target semantic feature corresponding to each image prompt text; Processing the object frame data corresponding to each image prompt text to obtain the target frame features of each object frame data; The target semantic features corresponding to each image prompt text and the target frame features of each object frame data are spliced to obtain at least one instance semantic feature.
3. The method according to claim 2, characterized in that The first semantic features corresponding to each image prompt text, the second semantic features corresponding to each image prompt text, and the preset learnable query features are fused to obtain the target semantic features corresponding to each image prompt text, including: The second semantic features corresponding to the prompt texts of each image and the preset learnable query features are fused through a cross attention mechanism to obtain a first fused feature; The first semantic features corresponding to each image prompt text and the first fusion features are fused through a cross-attention mechanism to obtain target semantic features corresponding to each image prompt text.
4. The method according to claim 2, characterized in that: The object frame data corresponding to each image prompt text is processed to obtain the target frame features of each object frame data, including: The object frame data corresponding to each image prompt text is processed by Fourier algorithm to obtain the spatial information of each object frame; The spatial information of each object frame is processed by a multi-layer perceptron to obtain the target frame features of each object frame data.
5. The method according to claim 1, characterized in that The noise image is denoised by using a diffusion model based on at least one instance semantic feature and a target frame mask feature corresponding to each object frame data, and the target image obtained includes: fusing at least one instance semantic feature, target frame mask features corresponding to each object frame data, and the noise map to obtain a global feature of at least one instance; Performing gated semantic fusion processing on the global feature of at least one instance to obtain a gated semantic fusion feature of at least one instance; The target image is generated based on the noise map and the gated semantic fusion features of at least one instance.
6. The method according to claim 5, characterized in that The global feature of at least one instance is obtained by fusing the semantic feature of at least one instance, the target frame mask feature corresponding to each object frame data, and the noise map, and comprising: Performing self-attention processing on the noise map to obtain target features of the noise map; The target features of the noise image, the target frame mask features corresponding to each object frame data, and at least one instance semantic feature are subjected to mask cross-attention processing through a mask cross-attention mechanism to obtain a global feature of at least one instance.
7. The method according to claim 5, characterized in that Generating the target image based on the noise map and the gated semantic fusion feature of at least one instance includes: Performing self-attention processing on the noise map to obtain target features of the noise map; Performing cross-attention fusion processing on the gated semantic fusion feature of at least one instance and the target feature of the noise map to obtain a cross-attention fusion feature of at least one instance; Adding the target feature of the noise image and the cross-attention fusion feature of at least one instance through a residual mechanism to obtain a residual fusion feature of at least one instance; The residual fusion features of at least one instance are processed through a feed-forward network to obtain the target image.
8. An image generating device, characterized in that: include: An acquisition module, used for acquiring at least one image prompt text and object frame data corresponding to each image prompt text; A determination module, configured to determine at least one instance semantic feature according to at least one image prompt text and object frame data corresponding to each image prompt text; A generation module, used to generate target frame mask features corresponding to each object frame data based on the object frame data corresponding to each image prompt text; The denoising module is used to denoise the noise image based on at least one instance semantic feature and a target frame mask feature corresponding to each object frame data through a diffusion model to obtain a target image.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Search enhancement-based text graph method, apparatus and device, and medium
CN120747274A
Multi-object scene-oriented training-condition-free diffusion image generation method
CN122265425A