Image generation method, device, electronic device, and readable storage medium
By integrating line, style and text features into an image generation method and using a diffusion model to iteratively generate images, the problem of inaccurate interpretation of user intent in existing technologies is solved, and high-quality image generation is achieved.
Patent Information
- Application Number
- CN202411884446.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Existing technologies are unable to accurately interpret the user's true intentions when generating images, resulting in the inability to generate high-quality images that meet user needs.
By obtaining line reference images, style reference images and text descriptions, fusing line image features, text features and style image features, and iterating using the diffusion model, the target image is generated.
The generated image more accurately meets the user's expectations in terms of overall perception and specific details, ensuring image quality and avoiding the problem of insufficient image quality in existing technologies caused by inaccurate interpretation of user intentions.
Smart Images

Figure CN119338951B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an image generation method, device, electronic device and readable storage medium. Background Art
[0002] With the application and advancement of artificial intelligence in the field of imaging, image generation is becoming increasingly efficient and convenient. In related technologies, the continuous advancement of multimodal technology is rapidly advancing the ability to create user-defined images based on text and reference images. However, existing technologies still have limitations in interpreting users' true intent, making it difficult to generate high-quality images that meet user needs. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide an image generation method, device, electronic device and readable storage medium to solve the problem that the existing technology cannot interpret the user's true intention when generating images, resulting in the inability to generate high-quality images that meet user needs.
[0004] According to a first aspect of an embodiment of the present invention, there is provided an image generation method, the method comprising:
[0005] A line reference image, a style reference image and a text description are obtained; the line image features obtained based on the line reference image and the text features obtained based on the text description are fused to obtain image-text control features; the image-text control features, the text features and the style image features obtained based on the style reference image are fused to obtain image-text fusion features; a noise image used to generate a target image is obtained, and the noise image, the structure control features, the image-text control features, the style image features and the image-text fusion features obtained based on the line reference image are input into a diffusion model for iteration to obtain a target image output by the diffusion model.
[0006] According to a second aspect of the embodiments of the present invention, there is provided an image generating apparatus, the apparatus comprising:
[0007] The acquisition module is configured to acquire a line reference image, a style reference image and a text description; the first fusion module is configured to fuse the line image features obtained based on the line reference image and the text features obtained based on the text description to obtain image-text control features; the second fusion module is configured to fuse the image-text control features, the text features and the style image features obtained based on the style reference image to obtain image-text fusion features; the image generation module is configured to acquire a noise image for generating a target image, and input the noise image, the structural control features obtained based on the line reference image, the image-text control features, the style image features and the image-text fusion features into the diffusion model for iteration to obtain the target image output by the diffusion model.
[0008] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0009] According to a fourth aspect of an embodiment of the present invention, a readable storage medium is provided, wherein the readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0010] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0011] The present invention obtains a line reference image, a style reference image, and a text description to determine the multiple, real-world needs for a target image. It then fuses line image features derived from the line reference image with text features derived from the text description to generate image-text control features. The image-text control features, text features, and style image features derived from the style reference image are then fused to generate image-text fusion features. This ensures that the resulting image-text control features and image-text fusion features meet both structural and text content requirements, more accurately aligning with user expectations for the overall look and feel of the target image and its specific details. A noise image is obtained to generate the target image. The noise image, structural control features, image-text control features, style image features, and image-text fusion features derived from the line reference image are then input into a diffusion model for iteration, resulting in the target image output by the diffusion model. This approach combines multiple features to continuously adjust and optimize image generation, ensuring image quality and ensuring that the generated image more closely matches the user's true intent. This overcomes the limitations of existing technologies, which often fail to fully and accurately interpret user intent and thus fail to generate high-quality images that meet user needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0013] Figure 1 is a flow chart of an image generation method provided by an embodiment of the present invention;
[0014] Figure 2 This is a structural example diagram of a fusion module provided by an embodiment of the present invention;
[0015] Figure 3 This is a structural diagram of a diffusion submodule provided by an embodiment of the present invention;
[0016] Figure 4 is a structural diagram of an image generating device provided by an embodiment of the present invention;
[0017] Figure 5 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0018] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration and not limitation to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present invention with unnecessary detail.
[0019] An image generation method, apparatus, electronic device, and readable storage medium according to embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0020] Figure 1 FIG. 1 is a flow chart of an image generation method provided by an embodiment of the present invention. Figure 1 As shown, the method includes:
[0021] S101, obtaining a line reference image, a style reference image, and a text description;
[0022] S102, fusing the line image features obtained based on the line reference image and the text features obtained based on the text description to obtain the image-text control features;
[0023] S103, fusing the image-text control feature, the text feature, and the style image feature obtained based on the style reference graph to obtain an image-text fusion feature;
[0024] S104, obtaining a noise image for generating a target image, and inputting the noise image, structural control features, image-text control features, style image features, and image-text fusion features obtained based on a line reference image into a diffusion model for iteration to obtain a target image output by the diffusion model.
[0025] Specifically, the image generation method of this embodiment can be executed by a client or a server, or can be executed jointly by a client and a server. The following description will be made using the client as an example. The line reference image, style reference image, and text description are reference information used to generate the target image, wherein the line reference image can be a line drawing and / or a simple drawing. The line reference image can represent the general structure, outline, and other information of the generated target image. The style reference image can represent the style characteristics of the target image that the user expects to generate, such as realistic, cartoon, abstract, and other styles. The text description can supplement the user's expectations of the target image from more specific semantic information and other aspects. By comprehensively obtaining these different forms of input, the user's intention can be more comprehensively understood.
[0026] It should be noted that after obtaining the line reference image, style reference image, and text description, corresponding encoders can be used to perform feature processing on the line reference image, style reference image, and text description to obtain line image features and structural control features corresponding to the line reference image, style image features corresponding to the style reference image, and text features corresponding to the text description. For example, an image encoder, such as a convolutional neural network encoder or a Transformer encoder, can be used to perform feature processing on the line reference image and style reference image, respectively, to obtain line image features corresponding to the line reference image and style image features corresponding to the style reference image. The line image features contain visual information of the line reference image, and the style image features contain style and content information of the style reference image, which will be used to influence the style of the generated image. Furthermore, a ControlNet network is used to perform feature processing on the line reference image to obtain structural control features corresponding to the line reference image, and a corresponding text encoder is used to process the text description to obtain text features corresponding to the text description. The structural control features will be used to guide the subsequent image generation process, ensuring that the image structure of the generated target image is consistent with the line reference image. The text features contain semantic information in the text description and will be used to guide the generated image to conform to the content of the text description.
[0027] It can be understood that the image-text control features are obtained by fusing the line image features obtained from the line reference image and the text features obtained from the text description. This organically combines the image's structural information (line image features) with the semantic information contained in the text description (text features). The line image features reflect the image's basic framework, while the text features contain the user's more detailed requirements for the image. The fused image-text control features can, at a comprehensive feature level, consider both the image's structural layout and the user's specific expectations through the text description. This allows the subsequent generation process to proceed in a direction that meets both structural requirements and text content needs based on this fused feature.
[0028] Furthermore, the image-text control features, text features, and style image features derived from a style reference image are fused to produce an image-text fusion feature. Based on the fusion of line image features and text features, style image features are further incorporated. Style is a key factor influencing image quality and whether it meets user preferences. By fusing image-text control features (which already incorporate line image features and text features), text features (to further reinforce text requirements), and style image features (to clarify style requirements) into an image-text fusion feature, the generated image closely aligns with the user's diverse needs in terms of content, structure, and style, more accurately meeting the user's expectations for the target image's overall appearance and specific details.
[0029] Furthermore, a noisy image is introduced and combined with the multiple features previously fused (structural control features, image-text control features, style image features, and image-text fusion features) to feed into the diffusion model for iteration. As the iterations progress, the generated image is continuously adjusted and optimized based on the various input features. This comprehensive input of multiple features enables the diffusion model to fully consider the user's needs from different perspectives in each iteration, gradually generating a high-quality image that meets the user's true intent from multiple aspects, such as image structure, content filling, and style shaping. This iterative generation approach, leveraging the synergy of multiple features, results in a target image that avoids the limitations of existing technologies, which struggle to fully and accurately interpret user intent and, therefore, fail to generate high-quality images that meet user needs.
[0030] According to the technical solution provided by the embodiments of the present invention, a line reference image, a style reference image, and a text description are obtained to determine the multiple real-world requirements for a target image. Line image features derived from the line reference image and text features derived from the text description are then fused to obtain image-text control features. The image-text control features, text features, and style image features derived from the style reference image are then fused to obtain image-text fusion features. The resulting image-text control features and image-text fusion features meet both structural and text content requirements, more accurately aligning with user expectations for the overall look and feel of the target image and specific details. A noise image is obtained to generate the target image, and the noise image, structural control features, image-text control features, style image features, and image-text fusion features derived from the line reference image are input into a diffusion model for iteration, resulting in the target image output by the diffusion model. This approach combines multiple features to continuously adjust and optimize image generation, ensuring the quality of the generated image while ensuring that the generated image more closely matches the user's true intent. This avoids the limitations of existing technologies, which often fail to fully and accurately interpret user intent and thus fail to generate high-quality images that meet user needs.
[0031] In some embodiments, line image features obtained based on a line reference image and text features obtained based on a text description are fused to obtain graphic and text control features, including: semantically enhancing text features to obtain enhanced text features; fusing line image features and enhanced text features to obtain cross-fusion features; and determining graphic and text control features based on the cross-fusion features, text features, and line image features.
[0032] Specifically, before fusing the line image features and the text features, the text features are first semantically enhanced to obtain enhanced text features. For example, the text features can be processed using a self-attention mechanism to capture the long-distance dependencies within the text features, enhance the semantic representation of the text, and obtain enhanced text features.
[0033] Furthermore, the line image features and enhanced text features are fused to obtain cross-fusion features. Line image features primarily reflect structural information such as line shape, direction, and connection relationships in the line reference image, while enhanced text features carry the user's demand for richer and more accurate semantics of the target image. Fusion of these two features organically combines the image's structural information with the semantic requirements described by text at a new feature level, so that the generated image conforms to both the basic structural framework set by the lines and the content conveyed by the text. During the fusion process, either a simple splicing approach or a cross-attention approach can be used. Exemplarily, this embodiment preferably uses a cross-attention approach to fuse line image features and enhanced text features. Specifically, the enhanced text features can be used as the query (Q), and the line image features can be used as the key (K) and value (V) to fuse the enhanced text features and line image features. For each dimension of the line image features and each dimension of the enhanced text features, their contribution to the fusion is adjusted based on the attention weight. In this way, the relationship between the two features can be more reasonably reflected in the cross-fusion features obtained by fusion, so that those feature parts that are more relevant to the user's intention can be given more attention, thereby achieving more accurate and richer multimodal information expression and generation.
[0034] Furthermore, after obtaining the cross-fusion features, the image-text control features are determined based on the cross-fusion features, text features, and line image features. This balances the image structure (derived from the line image features), the user's semantic needs (derived from the text features), and the structural and semantic information embodied by the initially fused cross-fusion features. This ensures that when generating images, no one aspect is overly emphasized at the expense of others. For example, focusing solely on line structure will not produce an image lacking semantic content, nor will emphasizing only semantic needs lead to a chaotic image structure. This allows for more precise guidance of image generation towards meeting user needs.
[0035] In some embodiments, the image-text control feature is determined based on the cross-fusion feature, text feature and line image feature, including: using a learnable cross-attention module to process the line image feature to obtain a learnable image feature; fusing the learnable image feature with the cross-fusion feature to obtain an image cross-fusion feature; and performing feature connection on the image cross-fusion feature and the text feature to obtain the image-text control feature.
[0036] Specifically, the learnable cross-attention module allows its parameters to be continuously adjusted and optimized during model training. By learning from large amounts of labeled or unsupervised data, the module can automatically find the optimal associations and fusion weights between different modal data, thereby better adapting to specific tasks and data characteristics.
[0037] It is understandable that the line image features are input into a learnable cross-attention module. This module will re-weight the line image features by calculating the attention weights between each part of the line image features and other possible related information based on its own learnable parameters and the input line image features, thereby obtaining learnable image features. For example, a line image feature may represent the features of a portrait of a person outlined by simple lines, including information such as the person's contour lines and limb lines. After processing by the cross-attention module, the line features of the person's head may be dynamically assigned different weights based on the subsequent fusion requirements with text features, etc., so that when fused with other features later, certain line parts can be more prominent or weakened to better fit the overall generation goal.
[0038] Furthermore, the learnable image features are fused with the cross-attention features to produce cross-fusion image features. This fusion further comprehensively and deeply combines the structural information represented by the line image features, the learnable image features adjusted by cross-attention, and the semantic information conveyed by the text features at a new feature level. For example, the learnable image features may provide a weighted representation that better aligns with the overall objective for certain key structural components in line images (such as the eye lines in a portrait). When fused with the cross-fusion features, this optimized representation of key structures can be further integrated with other elements in the cross-fusion features (such as features corresponding to the "vivid image" description related to text semantics), resulting in a cross-fusion image feature that captures both the key image structures and the text semantics.
[0039] Furthermore, the image cross-fusion features and text features are concatenated to obtain the image-text control feature. This directly connects the image structure and semantic information contained in the re-fused image cross-fusion features with the original, purest text semantic information retained in the text features. This allows the image-text control feature to encompass all key image information after multiple processing and fusion, as well as the complete semantic information of the text, in a unified feature representation. For example, the image cross-fusion feature may contain all the information related to the character's image, from the line structure to the fusion with the text semantics, such as the character's outline and expression, which is related to the text description of "lively and cute." The text feature may be the feature vector corresponding to the original description of "lively and cute." Through feature concatenation, the image-text control feature can fully integrate this information. In the subsequent image generation process, it can serve as a comprehensive control factor, guiding the generated image to meet the structural and fusion semantic requirements of the image while accurately reflecting the semantic connotation conveyed by the text, thereby more accurately meeting the user's needs for generated images.
[0040] It should be noted that feature connection can use residual connection to connect image cross-fusion features and text features to prevent the problem of gradient disappearance and gradient explosion during network training. In some examples, feature connection is performed on image cross-fusion features and text features, including: using a feedforward network to extract features of image cross-fusion features, and performing residual connection on the extracted image cross-fusion features and text features to obtain image-text control features.
[0041] In some embodiments, the image-text control feature, the text feature, and the style image feature obtained based on the style reference graph are fused to obtain the image-text fusion feature, including:
[0042] Cross-attention calculation is performed on the text features and style image features to obtain the initial fusion features; cross-attention calculation is performed on the initial fusion features and image-text control features to obtain image-text fusion features.
[0043] Specifically, individual text features and style image features can only describe part of the image from their respective perspectives. The initial fusion features, obtained through cross-attention calculation, fuse information from two different dimensions, enriching the feature representation while avoiding the problem of semantic and style separation that may result from simple splicing. The image-text control features are obtained by integrating multiple aspects of information, including line image features, text features, and fused cross-fusion features, and to a certain extent control the content and structure of the generated image. The initial fusion features are mainly a fusion of text and style. The image-text fusion features obtained through cross-attention calculation can integrate the fusion results of these two important stages, so that the generated image has both the correct structure and the appropriate style and semantic atmosphere.
[0044] In some embodiments, the diffusion model includes multiple diffusion sub-modules arranged in succession; the noise image, the structural control features, the image-text control features, the style image features, and the image-text fusion features obtained based on the line reference image are input into the diffusion model for iteration to obtain a target image output by the diffusion model, including: using the noise image as the target input of the first diffusion sub-module, and inputting the structural control features, the image-text control features, the style image features, and the image-text fusion features into each diffusion sub-module; using the first diffusion sub-module to process the noise image, the structural control features, the image-text control features, the style image features, and the image-text fusion features to obtain the target output of the first diffusion sub-module; using the target output of the first diffusion sub-module as the target input of the second diffusion sub-module, and using the second diffusion sub-module to process the target output, the structural control features, the image-text control features, the style image features, and the image-text fusion features of the first diffusion sub-module to obtain the target output of the second diffusion sub-module; using the target output of the second diffusion sub-module as the target input of the third diffusion sub-module, and repeating the processing flow of the diffusion sub-module until the target image output by the last diffusion sub-module is obtained.
[0045] Specifically, the diffusion model connects multiple diffusion submodules in sequence. The first diffusion submodule takes a noisy image as the target input and denoises the noisy image using structural control features, image-text control features, style image features, and image-text fusion features to obtain the corresponding target output. Each of the input features carries important information about different aspects of the target image. Structural control features help determine the basic structure and layout of the image, such as the general outline structure of a portrait or the overall framework of an architectural image. Image-text control features combine information from text descriptions with image-related features to accurately control the content and semantic direction of image generation. Style image features clarify the style characteristics of the image, such as realism, cartoon, retro, etc. The image-text fusion feature further integrates the results of the previous fusion of multiple features, influencing image generation from a more comprehensive perspective. Multiple features can work together to guide the image generation process.
[0046] Furthermore, the target output of the first diffusion submodule serves as the target input of the second diffusion submodule. This target input is then denoised using structural control features, image-text control features, style image features, and image-text fusion features. This yields the target output of the second diffusion submodule. This output then serves as the target input of the third diffusion submodule, and the process of these submodules is repeated until the target image is output by the final diffusion submodule. By sequentially feeding the noise image and various features (structural control features, image-text control features, style image features, and image-text fusion features) into each diffusion submodule for processing, each diffusion submodule further optimizes the image based on the output of the previous submodule and incorporating these features. Through this iterative process, the feature representation of the image increasingly aligns with user needs and the desired target image.
[0047] In some embodiments, the processing flow of each diffusion submodule includes:
[0048] The target input, image-text fusion features and style image features are fused to obtain the initial fusion features; the image-text control features and the time steps of the corresponding diffusion sub-module are processed using the multi-layer perceptron module respectively, and the processed image-text control features and the corresponding time steps are fused to obtain the time fusion features; the target output is determined based on the target input, initial fusion features, structural control features and time fusion features.
[0049] Specifically, during the processing of each diffusion submodule, the target input, the image-text fusion features, and the style image features are first fused to obtain the initial fused features. Specifically, the target input, the image-text fusion features, and the style image features are fused to obtain the initial fused features, including: processing the target input using a self-attention mechanism, and fusing the processed target input with the image-text fusion features to obtain the initial image-text fusion features; fusing the initial image-text fusion features with the style image features to obtain the initial fused features. When fusing the processed target input with the image-text fusion features, the processed target input can be used as the query (Q), and the image-text fusion features can be used as the key (K) and value (V) for cross-attention calculation to obtain the initial image-text fusion features. When fusing the initial image-text fusion features with the style image features to obtain the initial fused features, the initial image-text fusion features can be used as the query (Q), and the style image features can be used as the key (K) and value (V) for cross-attention calculation to obtain the initial fused features.
[0050] Furthermore, a multi-layer perceptron (MLP) module is used to process the image-text control features and the time steps of the corresponding diffusion submodules. The processed image-text control features and the corresponding time steps are then fused to produce a temporal fusion feature. After each MLP process, the time step and the global image-text control feature each generate a corresponding output feature vector. These two output feature vectors are then fused together, combining the information about the generation stage contained in the time step with the image generation control information contained in the image-text control feature to form a new temporal fusion feature. For example, if the feature vector obtained after MLP processing of the time step represents information such as the current generation stage's focus on the image contour, while the feature vector obtained after MLP processing of the global image-text control feature contains information about the specific content and style of the image, then the fused temporal fusion feature will contain both aspects of information, providing a more comprehensive feature foundation for subsequent integration into the network.
[0051] Furthermore, the output target output corresponding to each diffusion submodule is obtained based on the target input, initial fusion features, structural control features, and time fusion features. The specific method of obtaining the output target output will be described in detail in the subsequent embodiments and will not be elaborated here.
[0052] In some embodiments, a target output is determined based on a target input, initial style image features, structural control features, and temporal fusion features, including: normalizing the initial style image features and the temporal fusion features, and feature-connecting the normalized initial style image features and the temporal fusion features with the target input to obtain residual fusion features; extracting features from the residual fusion features, and normalizing the extracted residual fusion features and the temporal fusion features to obtain target fusion features; fusing the structural control features, the temporal fusion features, and the target fusion features to obtain a target output.
[0053] Specifically, the initial style image features and temporal fusion features are first normalized to ensure that the feature values operate on the same scale, which helps improve the stability and convergence speed of the model. Normalization can be achieved through various methods, such as adaptive layer normalization (AdaLN), z-score normalization, and min-max normalization. This embodiment preferably uses adaptive layer normalization (AdaLN) to normalize the style image features and temporal fusion features. The normalized features are then concatenated with the target input to form residual fusion features. This preserves the original information of the target input while introducing guidance from the style and temporal fusion features. For example, the normalized features are residually concatenated with the target input to form residual fusion features.
[0054] Furthermore, feature extraction is performed on the residual fusion features. Specifically, a feedforward network can be used to extract features from the residual fusion features. This is to extract more valuable feature information that can better reflect the target output requirements from the residual fusion features. At the same time, some redundant information is removed and key features are highlighted, so that subsequent fusion operations can be performed based on more refined and effective features, thereby improving the quality of the final target output. The residual fusion features and time fusion features after feature extraction are normalized to obtain target fusion features. The reason for performing normalization again is similar to the previous one, that is, to ensure that the residual fusion features and time fusion features after feature extraction maintain a suitable state in terms of feature value range, etc., so as to better perform subsequent fusion operations, further stabilize the model training process, and improve the model training effect and the quality of the final output.
[0055] Finally, the structural control features, temporal fusion features, and target fusion features are fused to produce the final target output. This fusion process can employ various strategies, such as weighted summation and concatenation, to ensure effective interaction and complementation between different features. This allows the image generation method to comprehensively consider multiple aspects of information, including structure, style, content, and time. Ultimately, through multiple iterations, it generates high-quality images that meet user needs.
[0056] Figure 2 This is a schematic diagram of the structure of a fusion module provided by an embodiment of the present invention. Figure 2 As shown:
[0057] It can be understood that the fusion module can be used to fuse line image features and text features to obtain image and text control features. The specific processing flow is as follows:
[0058] First, the self-attention layer is used to semantically enhance the text features to obtain enhanced text features;
[0059] Then, the line image features and enhanced text features are fused using the cross attention layer to obtain the cross fusion features;
[0060] Then, the line image features are processed using a learnable cross-attention module to obtain learnable image features;
[0061] Secondly, the learnable image features are fused with the cross-fusion features to obtain the image cross-fusion features;
[0062] Finally, the feedforward network is used to extract the image cross-fusion features, and the extracted image cross-fusion features are residually connected with the text features to obtain the image-text control features.
[0063] Figure 3 FIG. 1 is a structural diagram of a diffusion submodule provided by an embodiment of the present invention. Figure 3 As shown,
[0064] It can be understood that the diffusion submodule can iterate the noise image at each time step according to the input data. The specific processing flow of each diffusion submodule is as follows:
[0065] First, the target input is processed using the self-attention layer, and the processed target input is fused with the image-text fusion feature using the cross-attention layer to obtain the initial image-text fusion feature;
[0066] Then, the cross attention layer is used again to fuse the initial image-text fusion features and the style image features to obtain the initial fusion features;
[0067] Then, MLP is used to process the image-text control features and the time steps of the corresponding diffusion submodules respectively, and the processed image-text control features and the corresponding time steps are fused to obtain the time fusion features;
[0068] Secondly, the initial style image features and temporal fusion features are input into the AdaLN layer for normalization, and the normalized initial style image features and temporal fusion features are residually connected with the target input to obtain the residual fusion features;
[0069] Then, the residual fusion features are extracted using the feedforward network layer, and the residual fusion features and time fusion features after feature extraction are input into the AdaLN layer for normalization to obtain the target fusion features;
[0070] Finally, the structural control features, time fusion features and target fusion features are fused to obtain the target output.
[0071] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the process of the embodiment of the present invention.
[0072] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present invention, and are not described in detail here.
[0073] The following are embodiments of the apparatus of the present invention, which can be used to implement the method embodiments of the present invention. For details not disclosed in the apparatus embodiments of the present invention, please refer to the method embodiments of the present invention.
[0074] Figure 4 FIG. 1 is a structural diagram of an image generating device provided by an embodiment of the present invention. Figure 4 As shown, the device includes:
[0075] An acquisition module 401 is configured to acquire a line reference image, a style reference image, and a text description;
[0076] The first fusion module 402 is configured to fuse the line image features obtained based on the line reference image and the text features obtained based on the text description to obtain the image-text control features;
[0077] The second fusion module 403 is configured to fuse the image-text control feature, the text feature, and the style image feature obtained based on the style reference map to obtain an image-text fusion feature;
[0078] The image generation module 404 is configured to obtain a noise image for generating a target image, and input the noise image, the structural control features, the image-text control features, the style image features, and the image-text fusion features obtained based on the line reference image into the diffusion model for iteration to obtain the target image output by the diffusion model.
[0079] In some embodiments, the first fusion module 402 is further configured to perform semantic enhancement on text features to obtain enhanced text features; fuse line image features and enhanced text features to obtain cross-fusion features; and determine graphic control features based on the cross-fusion features, text features, and line image features.
[0080] In some embodiments, the first fusion module 402 is further configured to process line image features using a learnable cross-attention module to obtain learnable image features; fuse the learnable image features with the cross-fusion features to obtain image cross-fusion features; and perform feature connection on the image cross-fusion features and text features to obtain image-text control features.
[0081] In some embodiments, the second fusion module 403 is further configured to perform cross-attention calculation on text features and style image features to obtain initial fusion features; and perform cross-attention calculation on initial fusion features and image-text control features to obtain image-text fusion features.
[0082] In some embodiments, the image generation module 404 is further configured to use the noise image as the target input of the first diffusion sub-module, and input the structural control features, image-text control features, style image features and image-text fusion features into each diffusion sub-module; use the first diffusion sub-module to process the noise image, structural control features, image-text control features, style image features and image-text fusion features to obtain the target output of the first diffusion sub-module; use the target output of the first diffusion sub-module as the target input of the second diffusion sub-module, use the second diffusion sub-module to process the target output, structural control features, image-text control features, style image features and image-text fusion features of the first diffusion sub-module to obtain the target output of the second diffusion sub-module; use the target output of the second diffusion sub-module as the target input of the third diffusion sub-module, and repeat the processing flow of the diffusion sub-module until the target image output by the last diffusion sub-module is obtained.
[0083] In some embodiments, the image generation module 404 is further configured to fuse the target input, the image-text fusion features, and the style image features to obtain an initial fusion feature; use a multi-layer perceptron module to process the image-text control features and the time steps of the corresponding diffusion submodule respectively, and fuse the processed image-text control features and the corresponding time steps to obtain a time fusion feature; determine the target output based on the target input, the initial fusion feature, the structural control feature, and the time fusion feature.
[0084] In some embodiments, the image generation module 404 is further configured to normalize the initial style image features and the time fusion features, and feature-connect the normalized initial style image features and the time fusion features with the target input to obtain residual fusion features; perform feature extraction on the residual fusion features, and normalize the extracted residual fusion features and the time fusion features to obtain target fusion features; fuse the structural control features, the time fusion features and the target fusion features to obtain the target output.
[0085] Figure 5 Schematic diagram of an electronic device 5 provided by an embodiment of the present invention. Figure 5 As shown, the electronic device 5 of this embodiment includes: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable by the processor 501. When the processor 501 executes the computer program 503, the steps of the above-mentioned method embodiments are implemented. Alternatively, when the processor 501 executes the computer program 503, the functions of the modules / units in the above-mentioned device embodiments are implemented.
[0086] The electronic device 5 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 5 may include but is not limited to a processor 501 and a memory 502. Those skilled in the art will appreciate that Figure 5 This is merely an example of the electronic device 5 and does not limit the electronic device 5 . The electronic device 5 may include more or fewer components than shown in the figure, or different components.
[0087] The processor 501 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0088] The memory 502 may be an internal storage unit of the electronic device 5, such as a hard drive or memory of the electronic device 5. The memory 502 may also be an external storage device of the electronic device 5, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. The memory 502 may also include both an internal storage unit of the electronic device 5 and an external storage device. The memory 502 is used to store computer programs and other programs and data required by the electronic device.
[0089] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0090] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the present invention can also implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a readable storage medium, and when executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program can include computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The readable storage medium can include: any entity or device that can carry computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium.
[0091] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. An image generation method, characterized in that: include: Get line references, style references, and text descriptions; fusing the line image features obtained based on the line reference image and the text features obtained based on the text description to obtain image-text control features; fusing the image-text control feature, the text feature, and the style image feature obtained based on the style reference graph to obtain an image-text fusion feature; Obtaining a noise image for generating a target image, and inputting the noise image, the structural control features obtained based on the line reference image, the image-text control features, the style image features, and the image-text fusion features into a diffusion model for iteration to obtain the target image output by the diffusion model; the diffusion model includes a plurality of diffusion submodules connected in sequence; The fusing of the line image features obtained based on the line reference image and the text features obtained based on the text description to obtain the image-text control features includes: Performing semantic enhancement on the text features to obtain enhanced text features; Fusing the line image features and the enhanced text features to obtain cross-fusion features; Processing the line image features using a learnable cross-attention module to obtain learnable image features; Fusing the learnable image feature with the cross-fusion feature to obtain an image cross-fusion feature; The image cross-fusion feature and the text feature are feature-connected to obtain the image-text control feature.
2. The method according to claim 1, characterized in that The step of fusing the image-text control feature, the text feature, and the style image feature obtained based on the style reference graph to obtain an image-text fusion feature includes: Performing cross-attention calculation on the text features and the style image features to obtain initial fusion features; A cross-attention calculation is performed on the initial fusion feature and the image-text control feature to obtain the image-text fusion feature.
3. The method according to claim 1, characterized in that The step of inputting the noise image, the structural control feature obtained based on the line reference image, the image-text control feature, the style image feature, and the image-text fusion feature into a diffusion model for iteration to obtain the target image output by the diffusion model includes: The noise image is used as the target input of the first diffusion submodule, and the structure control feature, the image-text control feature, the style image feature, and the image-text fusion feature are input into each diffusion submodule; Using a first diffusion submodule, the noise image, the structural control feature, the image-text control feature, the style image feature, and the image-text fusion feature are processed to obtain a target output of the first diffusion submodule; using the target output of the first diffusion submodule as the target input of the second diffusion submodule, and using the second diffusion submodule to process the target output of the first diffusion submodule, the structure control feature, the image-text control feature, the style image feature, and the image-text fusion feature to obtain the target output of the second diffusion submodule; The target output of the second diffusion submodule is used as the target input of the third diffusion submodule, and the processing flow of the diffusion submodule is repeatedly executed until the target image output by the last diffusion submodule is obtained.
4. The method according to claim 3, characterized in that The processing flow of each diffusion submodule includes: Fusing the target input, the image-text fusion feature, and the style image feature to obtain an initial fusion feature; The image-text control feature and the time step of the corresponding diffusion submodule are processed respectively using a multi-layer perceptron module, and the processed image-text control feature and the corresponding time step are fused to obtain a time fusion feature; The target output is determined according to the target input, the initial fusion feature, the structural control feature, and the time fusion feature.
5. The method according to claim 4, characterized in that The determining the target output according to the target input, the initial style image feature, the structural control feature, and the temporal fusion feature includes: Normalizing the initial style image features and the time fusion features, and performing feature concatenation on the normalized initial style image features and the time fusion features with the target input to obtain residual fusion features; Extracting features from the residual fusion features, and normalizing the extracted residual fusion features and the time fusion features to obtain target fusion features; The structural control feature, the time fusion feature and the target fusion feature are fused to obtain the target output.
6. An image generating device, characterized in that: include: an acquisition module configured to acquire a line reference image, a style reference image, and a text description; The first fusion module is configured to fuse the line image features obtained based on the line reference image and the text features obtained based on the text description to obtain a picture-text control feature; the fusing the line image features obtained based on the line reference image and the text features obtained based on the text description to obtain the picture-text control feature includes: semantically enhancing the text features to obtain enhanced text features; fusing the line image features and the enhanced text features to obtain a cross-fusion feature; processing the line image features using a learnable cross-attention module to obtain a learnable image feature; fusing the learnable image feature with the cross-fusion feature to obtain an image cross-fusion feature; and feature-connecting the image cross-fusion feature and the text feature to obtain the picture-text control feature; A second fusion module is configured to fuse the image-text control feature, the text feature, and the style image feature obtained based on the style reference graph to obtain an image-text fusion feature; The image generation module is configured to obtain a noise image for generating a target image, and input the noise image, the structural control features obtained based on the line reference image, the image-text control features, the style image features, and the image-text fusion features into a diffusion model for iteration to obtain the target image output by the diffusion model, wherein the diffusion model includes multiple diffusion sub-modules connected in sequence.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Image generation method and device and computer storage medium
CN117408899A
Video generation method and device, electronic equipment and readable storage medium
CN119052529A