Image generation method, system, equipment and medium
Through the image generation model and adaptation module based on the diffusion mechanism, efficient and diversified generation of marketing images is achieved, the problems of low generation efficiency and inconsistent style in the existing technology are solved, and the demand for rapid output in multiple scenarios is met.
Patent Information
- Application Number
- CN202510448437.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, marketing images are generated inefficiently, unable to adapt to the rapid output of multiple scenes, and manual drawing costs are high, making it difficult to maintain the consistency of style.
By introducing scene keywords and image keywords, semantic encoding is used using an image generation model based on the diffusion mechanism, and feature fusion and image decoding are combined with scene low-rank adaptation module and image low-rank adaptation module to generate marketing images.
It improves the efficiency of marketing images generation, ensures semantic independence and style consistency between background and image, reduces the cost of manual drawing, and meets the needs of diversified and rapid generation.
Smart Images

Figure CN120298530A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to an image generation method, system, device and medium. Background Art
[0002] In the operation of Internet platforms, the number of users is a key indicator to measure their development level. Excellent marketing activities can not only attract new users, but also effectively improve the activity and retention rate of old users. Therefore, marketing activities have become an important means to promote the growth of platforms. As the first interface for users to contact the brand, the style and performance of marketing visual content directly reflect the brand tone of the platform and have an important impact on user experience and brand perception. The currently still commonly used manual drawing method is used to produce marketing images through flat drawing or three-dimensional rendering. Although this method has certain advantages in personalized expression, its production cycle is long, the labor cost is high, and it is difficult to maintain the consistency of style. Especially when facing actual requirements such as multiple scenarios, multiple versions, and rapid updates, every time a new marketing scenario is added, a complete design process needs to be experienced again, and its drawing efficiency is difficult to meet the fast-paced needs of modern Internet marketing. Therefore, there is a need to provide an image generation method, system, device and medium. Summary of the Invention
[0003] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide an image generation method, system, device and medium, which improves the problems of low image generation efficiency in the prior art and inability to adapt to rapid output in multiple scenarios.
[0004] To achieve the above object and other related objects, the present invention provides an image generation method, including: obtaining a scene keyword and an image keyword; wherein, the scene keyword is used to represent the background environment in the marketing image to be generated, and the image keyword is used to represent the target image in the marketing image to be generated; inputting the scene keyword and the image keyword into a text encoder of an image generation model for encoding to generate a text embedding feature; wherein, the image generation model is a text-to-image model based on a diffusion mechanism; inputting the text embedding feature and a preset latent space noise into a diffusion network of the image generation model for feature fusion to generate an image feature fused with an image feature and a background feature; inputting the image feature into an image decoder of the image generation model to generate a marketing image.
[0005] In an embodiment of the present invention, the inputting the scene keyword and the image keyword into a text encoder of an image generation model for encoding to generate a text embedding feature includes: splicing the scene keyword and the image keyword to form a prompt text; inputting the prompt text into the text encoder for encoding to generate a text embedding feature.
[0006] In one embodiment of the present invention, the text embedding feature and the preset latent space noise are input into the diffusion network of the image generation model for feature fusion to generate image features fused with image features and background features, including: for each preset diffusion time step: the latent space noise feature of the current diffusion time step and the text embedding feature are input into the basic model of the diffusion network, and the background features used to generate the background environment are adjusted according to the scene low-rank adaptation module of the diffusion network, and the image features used to generate the target image are adjusted based on the image low-rank adaptation module of the diffusion network; wherein the diffusion network includes a basic model, a scene low-rank adaptation module and an image low-rank adaptation module, and the latent space noise feature of the first diffusion time step is an initialized Gaussian noise; the background feature and the image feature are fused to generate the latent space prediction feature of the current diffusion time step, and the latent space noise feature of the next diffusion time step is updated accordingly; the latent space prediction feature of the last preset diffusion time step is used as the image feature fused with background features and image features.
[0007] In one embodiment of the present invention, the step of inputting the image features into the image decoder of the image generation model to generate a marketing image includes: inputting the image features into the image decoder, and using the image decoder to perform layer-by-layer deconvolution and upsampling on the image features to obtain a marketing image; wherein the image decoder is a decoder module of a variational autoencoder.
[0008] In one embodiment of the present invention, after inputting the image features into the image decoder of the image generation model to generate a marketing image, the method further includes: analyzing the marketing image based on an image quality assessment method to determine the area to be optimized in the marketing image; and performing local redrawing processing on the area to be optimized to generate a final marketing image.
[0009] In an embodiment of the present invention, the scene low-rank adaptation module is trained. The training process of the scene low-rank adaptation module includes: obtaining scene images labeled with scene keywords; wherein, the scene keywords are used to characterize the background environment in the scene images; inputting the scene keywords into the text encoder of the image generation model to encode the scene keywords and obtain scene text embedding features; inputting the scene images into the image encoder of the image generation model to encode the scene images and obtain scene image features; training the scene low-rank adaptation module at each diffusion time step according to a preset total number of diffusion time steps to obtain a trained scene low-rank adaptation module; wherein, the following process is performed at each diffusion time step: inputting the scene text embedding features and the scene image features added with random noise into the diffusion network of the image generation model to generate predicted noise; wherein, the diffusion network includes a base model with frozen parameters and a trainable scene low-rank adaptation module, and the scene low-rank adaptation module acts on a preset network layer in the base model; calculating the difference degree between the predicted noise and the random noise at the current diffusion time step, and updating the parameters of the scene low-rank adaptation module based on the difference degree.
[0010] In an embodiment of the present invention, the image low-rank adaptation module is trained. The training process of the image low-rank adaptation module includes: obtaining image images labeled with image keywords; wherein, the image keywords are used to characterize the background environment in the image images; inputting the image keywords into the text encoder of the image generation model to encode the image keywords and obtain image text embedding features; inputting the image images into the image encoder of the image generation model to encode the image images and obtain image image features; training the image low-rank adaptation module at each diffusion time step according to a preset total number of diffusion time steps to obtain a trained image low-rank adaptation module; wherein, the following process is performed at each diffusion time step: inputting the image text embedding features and the image image features added with random noise into the diffusion network of the image generation model to generate predicted noise; wherein, the diffusion network includes a base model with frozen parameters and a trainable image low-rank adaptation module, and the image low-rank adaptation module acts on a preset network layer in the base model; calculating the difference degree between the predicted noise and the random noise at the current diffusion time step, and updating the parameters of the image low-rank adaptation module based on the difference degree.
[0011] In an embodiment of the present invention, an image generation system is further provided. The system includes: a data acquisition module for acquiring a scene keyword and an image keyword, where the scene keyword is used to represent the background environment in the marketing image to be generated, and the image keyword is used to represent the target image in the marketing image to be generated; a text encoding module for inputting the scene keyword and the image keyword into the text encoder of the image generation model for encoding to generate text embedding features, where the image generation model is a text-to-image model based on a diffusion mechanism; a diffusion module for inputting the text embedding features and a preset latent space noise into the diffusion network of the image generation model for feature fusion to generate image features fused with image features and background features; and an image generation module for inputting the image features into the image decoder of the image generation model to generate a marketing image.
[0012] In an embodiment of the present invention, an electronic device is further provided, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device implements the image generation method of any one of the above.
[0013] In an embodiment of the present invention, a computer-readable storage medium is further provided, on which a computer program is stored, and when the computer program is executed by the processor of the computer, the computer executes the image generation method of any one of the above.
[0014] As described above, an image generation method, system, device, and medium of the present invention have the following beneficial effects: By introducing a scene keyword and an image keyword to respectively represent the background environment and the target image in the marketing image, and performing semantic encoding on the keywords based on the text encoder to perform refined semantic guidance on the image generation content. It effectively improves the semantic independence and clarity of the background and the image during the generation process, thereby greatly improving the content composition, style coordination, and expression consistency ability of the finally generated marketing image. The present invention greatly reduces the manual drawing cost, so as to meet the rapid generation demand for high-quality and diverse images in the marketing scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a flowchart of an image generation method provided by an embodiment of the present invention;
[0016] Figure 2 It is a structural block diagram of an image generation system provided by an embodiment of the present invention;
[0017] Figure 3 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0018] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0019] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0020] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0021] The present invention provides a method for generating an image. By introducing a scene keyword and an image keyword, which respectively represent the background environment and the target image in the marketing image, and performing semantic encoding on the keywords based on a text encoder, semantic guidance for the image generation content can be refined. The semantic independence and clarity of the background and the image in the generation process are effectively improved, thereby greatly enhancing the content composition, style coordination, and expression consistency of the finally generated marketing image. The manual drawing cost is significantly reduced, so that the rapid generation demand for high-quality and diverse images in the marketing scenario can be met.
[0022] As Figure 1 shown, the method for generating an image includes the following steps:
[0023] S1. Obtain a scene keyword and an image keyword; wherein, the scene keyword is used to represent the background environment in the marketing image to be generated, and the image keyword is used to represent the target image in the marketing image to be generated.
[0024] When a specific marketing image needs to be generated, the user can directly input preset scene keywords and image keywords according to specific marketing needs to specify the background environment and target image content of the image. Exemplarily, the scene keywords can be festival street, night city, etc., which are used to characterize the background setting style and spatial atmosphere in the image; the image keywords can be truck, a man in formal dress, etc., which are used to describe the main characters or image features in the image. In addition, the user can also input natural language words they want to express, and through methods such as keyword matching or semantic similarity calculation, automatically determine the scene keywords and image keywords that are semantically closest to them from the existing keyword library, so as to improve the flexibility and intelligence of keyword input. It should be noted that both the scene keywords and the image keywords can be one or multiple, and those skilled in the art can adaptively select based on the actual marketing activities and are not limited here.
[0025] S2. Input the scene keywords and image keywords into the text encoder of the image generation model for encoding to generate text embedding features; wherein, the image generation model is a text-to-image model based on the diffusion mechanism.
[0026] Input the scene keywords and image keywords into the text encoder in the image generation model for semantic encoding to generate text embedding features. As a conditional input, during the diffusion process, it is input into the diffusion network together with the latent space noise at each diffusion time step. By gradually removing the added random noise and adjusting the features, the generated image gradually tends to the visual content consistent with the keyword semantics, thereby realizing the directional generation of the background environment and the target image. It should be noted that the image generation model can be any text-to-image model, as long as it adopts a diffusion-based generation architecture, uses text as the main control condition, and is used to gradually generate an image that matches the semantics of the input text from the noise, and as long as it has a text encoder, an image decoder, and a diffusion network. Among them, the diffusion network includes but is not limited to models such as StableDiffusion and Imagen.
[0027] Specifically, the process of generating text embedding features includes steps S21 and S22 (not shown in the figure):
[0028] S21. Concatenate the scene keywords and image keywords to form a prompt text.
[0029] According to the preset concatenation rule, concatenate the scene keywords and image keywords to form a semantically coherent prompt text. Exemplarily, the scene keywords can be placed in the front, and the image keywords can be placed behind the scene keywords to obtain the prompt text, such as "a man in formal dress walking on the street with a mobile phone in his hand".
[0030] S22. Input the prompt text into the text encoder for encoding to generate text embedding features.
[0031] Use the text encoder to perform semantic extraction operations such as word segmentation, embedding mapping, and context modeling on the input prompt text to generate structured high-dimensional text embedding features. This text embedding feature retains scene information and image information and is input as semantic guidance in the form of a tensor for use by the diffusion model during image generation. Among them, the text encoder can be any Transformer-based encoder, including but not limited to the CLIP model, BERT model, etc. Preferably, in order to better capture cross-modal semantic relationships, the text encoder is the CLIP model.
[0032] S3. Input the text embedding features and the preset latent space noise into the diffusion network of the image generation model for feature fusion to generate image features that fuse image features and background features.
[0033] Input the text embedding features obtained after encoding and the preset latent space noise into the diffusion network together. The text embedding features are encoded by scene keywords and image keywords, carrying the semantic information of the background environment and the target image in the image, and participate in the regulation of the latent space features through the cross-modal attention mechanism in the diffusion network, thereby guiding the generation direction of the image content. The latent space noise vector is a high-dimensional noise representation randomly sampled from a Gaussian distribution and is used to initialize the diffusion process, which is the starting point for the gradual restoration of the image. The diffusion network can only include a basic diffusion model. Without introducing additional structures, through the feature iteration and semantic guidance of multiple preset diffusion time steps, it is possible to gradually control the latent space state of the image under the guidance of semantic information, thereby generating image features containing background and image semantics for subsequent decoding of the image to generate marketing images. Among them, the image features obtained by the diffusion network are latent space image features.
[0034] Furthermore, the diffusion network includes a basic diffusion model as the backbone structure, and a Scene Low-Rank Adaptation (LoRA) module and an Image Low-Rank Adaptation module integrated into the basic model. At each diffusion time step of the diffusion process, the current latent space noise feature and the text embedding feature are input into the diffusion network together, where the latent space noise feature at the current diffusion time step is initialized random noise / the latent space prediction feature obtained at the previous diffusion time step. The basic diffusion model and the two low-rank adaptation modules work together to semantically guide and locally adjust the latent space noise feature. Among them, the Scene Low-Rank Adaptation module is used to adjust the background-related features in the latent space according to the scene keywords, so that the background of the generated image is consistent with the expected environmental atmosphere. The Image Low-Rank Adaptation module is used to adjust the target image features in the latent space according to the image keywords, so as to ensure that the generated character or IP image conforms to the input description in terms of structure, style, etc. After all diffusion time steps are processed according to the above process, the background features and image features respectively adjusted by the Scene Low-Rank Adaptation module and the Image Low-Rank Adaptation module are fused in the latent space to obtain the complete image features including the background environment and the target image.
[0035] Specifically, the process of generating the above image features includes steps S31 to S3 (not shown in the figure):
[0036] S31. For each preset diffusion time step, execute the following process:
[0037] First, input the latent space noise feature and the text embedding feature at the current diffusion time step into the basic model of the diffusion network, and adjust the background feature for generating the background environment according to the Scene Low-Rank Adaptation module, and adjust the image feature for generating the target image based on the Image Low-Rank Adaptation module of the diffusion network; where the diffusion network includes a basic model, a Scene Low-Rank Adaptation module and an Image Low-Rank Adaptation module, and the latent space noise feature at the first diffusion time step is initialized Gaussian noise.
[0038] For the first diffusion time step, execute the following process:
[0039] Input the initialized Gaussian noise and text embedding features into the base model, model the Gaussian noise, and use the scene low-rank adaptation module and the image low-rank adaptation module integrated inside the base model for semantic adjustment. Among them, the scene low-rank adaptation module adjusts the background-related features according to the scene keywords, making the layout and color style of the scene in the image continuously consistent with the text semantic description of the scene keywords, and obtaining the scene features used to represent the image background. The image low-rank adaptation module adjusts the local details of the character or target image, making its appearance structure, posture, and style features gradually approach the target image expressed in the image keywords, and obtaining the image features used to represent the IP image. After obtaining the background features and image features, fuse the background features and image features to generate the latent space prediction features at the first diffusion time step, and obtain the latent space noise features required for the next diffusion time step according to the diffusion mechanism.
[0040] For each of the remaining diffusion time steps: After obtaining the latent space noise features at the current diffusion time step based on the latent space prediction features at the previous diffusion time step, input the latent space noise features at the current diffusion time step and the text embedding features into the base model of the diffusion network together. The base model further models the current latent space state and continues to call the scene low-rank adaptation module and the image low-rank adaptation module integrated inside it for semantic adjustment. Among them, the scene low-rank adaptation module continuously optimizes the background information in the latent space according to the scene keywords included in the text embedding features, making the spatial composition, light and shadow effects, etc. in the image gradually approach the text description of the target scene, and further enhancing the stability and expression consistency of the scene features. The image low-rank adaptation module refines the performance effect of the character or target image on the basis of the previous diffusion time step, further adjusts its appearance contour, posture structure, and style details, making the target image in the generated image more accurately and naturally match the semantic features required by the image keywords.
[0041] S32: Use the latent space prediction features at the last preset diffusion time step as the image features that fuse the background environment and the target image.
[0042] After all diffusion time steps are executed, through the gradual adjustment of the scene low-rank adaptation module and the image low-rank adaptation module during the entire diffusion process, the finally obtained latent space prediction features not only retain the diversity information in the original noise but also completely encode the visual semantic information guided by the text embedding features. At this time, the latent space prediction features generated at the last diffusion time step can be used as image features for subsequent image generation.
[0043] The present invention realizes the separate control and joint construction of the image background environment and the target image by introducing scene keywords and image keywords, and performing semantic adjustment through the scene low-rank adaptation module and the image low-rank adaptation module in the image generation model respectively, which improves the controllability and semantic accuracy of the image generation process, greatly improves the fitting degree of the generated marketing image with the scene keywords and image keywords, and can ensure the consistency of the scene and the target image through the scene low-rank adaptation module and the image low-rank adaptation module.
[0044] S4. Input the image features into the image decoder of the image generation model to generate a marketing image.
[0045] Specifically, the image generation process includes: inputting the image features into the image decoder, and using the image decoder to perform layer-by-layer deconvolution and upsampling on the image features to obtain a marketing image; wherein, the image decoder is the decoder module of the variational autoencoder. Specifically, input the image features fused with the semantic information of the background environment and the target image into the image decoder in the image generation model, and perform layer-by-layer deconvolution (transpose convolution) operations and upsampling processing on the input image features. By performing step-by-step restoration processing on the image features, a marketing image with a complete spatial structure and visual details can be decoded. The marketing image has a preset resolution, can accurately reflect the background environment and the target image described in the text prompt, and is applicable to various marketing scenarios such as brand communication and advertising display.
[0046] Further, after generating the marketing image, the following secondary optimization processing process is also included:
[0047] First, analyze the marketing image based on the image quality assessment method to determine the area to be optimized in the marketing image.
[0048] After the image generation is completed, in order to improve the visual quality and expression effect of the marketing image, the image quality assessment method is also used to analyze the generated marketing image to identify the parts in the marketing image that are blurred, distorted or lack details, so as to determine the area that needs to be optimized. Mark these areas as the areas to be optimized for subsequent local redrawing. Among them, the image quality assessment method includes, but is not limited to, gradient analysis based on edge sharpness, structural similarity index, etc. Those skilled in the art can adaptively select according to actual needs and are not limited here.
[0049] Then perform local redrawing processing on the area to be optimized to generate the final marketing image.
[0050] Specifically, based on the contextual information and text embedding features of the original marketing image, the image generation model is guided to regenerate the area to be optimized, ensuring that the redrawn content is semantically consistent with the overall image, and is naturally connected in style, structure, color, etc., thereby obtaining the final marketing image.
[0051] The scene low-rank adaptation module in the present invention is obtained by training, and the training process of the scene low-rank adaptation module includes the following process:
[0052] First, a scene image annotated with scene keywords is obtained; wherein the scene keywords are used to represent the background environment in the scene image.
[0053] In the present invention, the scene low-rank adaptation module is obtained through training and is used to semantically guide and optimize background environment features during image generation. During training, it is necessary to obtain scene images annotated with scene keywords, and the scene keywords are used to characterize the background environment information in the corresponding image.
[0054] After obtaining the scene keywords and scene images, the scene keywords are input into the text encoder of the image generation model to encode the scene keywords and obtain the scene text embedding features.
[0055] The scene keywords are input into the text encoder in the image generation model, and the scene keywords are segmented, embedded and mapped, and context modeled to generate high-dimensional scene text embedding features for representing scene semantic information. This embedding feature lays the foundation for semantic alignment in the training phase, enabling the image generation model to learn the semantic mapping relationship between keywords and image background.
[0056] The scene image is input into the image encoder of the image generation model, the scene image is encoded, and the scene image features are obtained.
[0057] The acquired scene image is input into the image encoder in the image generation model. The image encoder performs layer-by-layer convolution, downsampling and feature extraction operations on the input scene image, compresses the high-dimensional image data and converts it into a low-dimensional latent space representation, and obtains the scene image features used for modeling. It can be understood that the scene image features correspond one-to-one to the scene text embedding features obtained by encoding the scene keywords, which are used for semantic alignment in the training stage, thereby enhancing the model's perception and generation capabilities of background semantics.
[0058] According to the preset total number of diffusion time steps, the scene low-rank adaptation module is trained at each diffusion time step to obtain a trained scene low-rank adaptation module; wherein the following process is performed at each diffusion time step:
[0059] First, input the scene text embedding feature and the scene image feature with added random noise into the diffusion network of the image generation model to generate predicted noise. Among them, the diffusion network includes a base model with frozen parameters and a trainable scene low-rank adaptation module, and the scene low-rank adaptation module acts on a preset network layer in the base model. Specifically, input the scene text embedding feature and the scene image feature with added random Gaussian noise corresponding to the current diffusion time step into the diffusion network of the image generation model to generate the predicted noise at the current diffusion time step. This image feature is obtained by adding random Gaussian noise corresponding to the diffusion time step to the original image encoding result and is used to simulate the intermediate state in the image generation process. The text embedding feature provides semantic guidance information corresponding to this image. In this embodiment, the diffusion network includes a base model and a trainable scene low-rank adaptation module. Among them, the parameters of the base model are frozen and do not participate in parameter update; the scene low-rank adaptation module is inserted into a preset network layer of the base model to fine-tune and optimize the response of the base model in the background feature dimension so that the generated background feature is more suitable for the scene keyword. Then, calculate the difference degree between the predicted noise and the random noise at the current diffusion time step, and update the parameters of the scene low-rank adaptation module based on the difference degree. Repeat the above process until the final difference degree converges or reaches the preset number of iterations to obtain a trained scene low-rank adaptation module.
[0060] In the present invention, the image low-rank adaptation module is trained, and the training process of the image low-rank adaptation module includes:
[0061] Obtain an image with an image keyword marked; where the image keyword is used to characterize the background environment in the image.
[0062] Input the image keyword into the text encoder of the image generation model to encode the image keyword and obtain an image text embedding feature.
[0063] Input the image into the image encoder of the image generation model to encode the image and obtain an image feature.
[0064] According to the preset total number of diffusion time steps, train the image low-rank adaptation module at each diffusion time step to obtain a trained image low-rank adaptation module. Among them, at each diffusion time step, perform the following process:
[0065] First, input the image text embedding feature and the image feature with added random noise into the diffusion network of the image generation model to generate predicted noise. Among them, the diffusion network includes a base model with frozen parameters and a trainable image low-rank adaptation module, and the image low-rank adaptation module acts on a preset network layer in the base model.
[0066] Then, calculate the difference between the predicted noise and the random noise at the current diffusion time step, and update the parameters of the image low-rank adaptation module based on the difference.
[0067] The training process of the image low-rank adaptation module is similar to that of the above-mentioned scene low-rank adaptation module, and will not be elaborated here.
[0068] As Figure 2 shown, the image generation system 100 includes: a data acquisition module 110, a text encoding module 120, a diffusion module 130, and an image generation module 140. The above-mentioned data acquisition module 110 is used to acquire a scene keyword and an image keyword; wherein, the scene keyword is used to characterize the background environment in the marketing image to be generated, and the image keyword is used to characterize the target image in the marketing image to be generated. The text encoding module 120 is used to input the scene keyword and the image keyword into the text encoder of the image generation model for encoding to generate text embedding features; wherein, the image generation model is a text-to-image model based on the diffusion mechanism. The diffusion module 130 is used to input the text embedding features and a preset latent space noise into the diffusion network of the image generation model for feature fusion to generate image features fused with image features and background features. The image generation module 140 is used to input the image features into the image decoder of the image generation model to generate a marketing image.
[0069] For the specific limitations of the image generation system, reference can be made to the limitations on the image generation method in the above text, which will not be elaborated here. Each module in the above image generation system can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor of the computer device in a hardware format or independent of it, or stored in the memory of the computer device in a software format, so that the processor can call the corresponding operations of the above modules.
[0070] It should be noted that, in order to highlight the innovative part of the present invention, modules not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other modules in this embodiment.
[0071] As Figure 3 shown, the electronic device 1 may include a memory 11, a processor 12, and a bus, and may also include a computer program stored in the memory 11 and executable on the processor 12, such as an image generation program.
[0072] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 11 can be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 11 can also be an external storage device of the electronic device 1 in other embodiments, such as a plug-in mobile hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 can also include both the internal storage unit and the external storage device of the electronic device 1. The memory 11 can be used not only to store application software installed on the electronic device 1 and various types of data, such as the code for image generation, etc., but also to temporarily store data that has been output or will be output.
[0073] In some embodiments, the processor 12 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged together, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 12 is the control core (Control Unit) of the electronic device 1, connecting various components of the entire electronic device 1 through various interfaces and circuits. By running or executing programs or modules stored in the memory 11 (such as the image generation program, etc.), and calling data stored in the memory 11, it performs various functions of the electronic device 1 and processes data.
[0074] The processor 12 executes the operating system of the electronic device 1 and various installed application programs. The processor 12 executes the application program to implement the steps in the above-mentioned image generation method.
[0075] Exemplarily, the computer program can be divided into one or more modules. One or more modules are stored in the memory 11 and executed by the processor 12 to complete this application. One or more modules can be a series of computer program instruction segments capable of completing specific functions, and these instruction segments are used to describe the execution process of the computer program in the electronic device 1. For example, the computer program can be divided into a data acquisition module 110, a text encoding module 120, a diffusion module 130, and an image generation module 140.
[0076] The integrated unit implemented in the form of software function modules described above can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The above software function modules are stored in a storage medium and include several instructions for causing a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute part of the functions of the method for generating images according to various embodiments of the present application.
[0077] In summary, a method, system, device, and medium for generating images disclosed by the present invention combines multi-model invocation with generative artificial intelligence technology to construct a method for generating images applicable to various marketing scenarios. Compared with the traditional manual drawing method, the present invention uses the trained scene low-rank adaptation and image low-rank adaptation modules to separately control and adjust the background environment and the target image in the image generation process, which not only significantly improves the design efficiency and shortens the production cycle of marketing images and three-dimensional scenes, but also ensures the consistency and stability of the visual style in marketing activities. At the same time, this method of the present invention also supports various forms such as two-dimensional images, photos, and three-dimensional visual maps, has good scalability and multi-scenario adaptability, and meets the content generation requirements of the brand under diverse marketing needs. Through differentiated visual performance, this method helps to uniformly shape the brand image and enhance the market influence, and supports the brand to implement a differentiation strategy at the visual content level in the fierce competition. Compared with the traditional solution, the present invention has obvious advantages in aspects such as multi-model collaborative regulation, semantic-guided generation, and scene adaptation ability, and provides a more intelligent, efficient, and innovative display marketing image generation solution. Therefore, the present invention effectively overcomes various shortcomings in the prior art and has high industrial utilization value.
[0078] The above embodiments are only illustrative of the principles and effects of the present invention and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.
Claims
1. A method for generating an image, characterized in that, The method includes: Obtaining a scene keyword and an image keyword; wherein, the scene keyword is used to characterize the background environment in the marketing image to be generated, and the image keyword is used to characterize the target image in the marketing image to be generated; Inputting the scene keyword and the image keyword into a text encoder of an image generation model for encoding to generate a text embedding feature; wherein, the image generation model is a text-to-image model based on a diffusion mechanism; Inputting the text embedding feature and a preset latent space noise into a diffusion network of the image generation model for feature fusion to generate an image feature integrating an image feature and a background feature; Inputting the image feature into an image decoder of the image generation model to generate a marketing image.
2. The method for generating an image according to claim 1, wherein The step of inputting the scene keyword and the image keyword into a text encoder of an image generation model for encoding to generate a text embedding feature includes: Concatenating the scene keyword and the image keyword to form a prompt text; Inputting the prompt text into the text encoder for encoding to generate a text embedding feature.
3. The method for generating an image according to claim 1, wherein The step of inputting the text embedding feature and a preset latent space noise into a diffusion network of the image generation model for feature fusion to generate an image feature integrating an image feature and a background feature includes: For each preset diffusion time step: Inputting the latent space noise feature at the current diffusion time step and the text embedding feature into a basic model of the diffusion network, and adjusting the background feature for generating the background environment according to a scene low-rank adaptation module of the diffusion network, and adjusting the image feature for generating the target image based on an image low-rank adaptation module of the diffusion network; wherein, the diffusion network includes a basic model, a scene low-rank adaptation module and an image low-rank adaptation module, and the latent space noise feature at the first diffusion time step is an initialized Gaussian noise; Fusing the background feature and the image feature to generate a latent space prediction feature at the current diffusion time step, and updating the latent space noise feature at the next diffusion time step accordingly; Taking the latent space prediction feature at the last preset diffusion time step as the image feature integrating the background feature and the image feature.
4. The method for generating an image according to claim 3, wherein The scene low-rank adaptation module is obtained through training, and the training process of the scene low-rank adaptation module includes: Obtaining a scene image labeled with a scene keyword; wherein, the scene keyword is used to characterize the background environment in the scene image; Inputting the scene keyword into a text encoder of the image generation model to encode the scene keyword to obtain a scene text embedding feature; Inputting the scene image into an image encoder of the image generation model to encode the scene image to obtain a scene image feature; Training the scene low-rank adaptation module at each diffusion time step according to a total amount of preset diffusion time steps to obtain a trained scene low-rank adaptation module; wherein, the following process is executed at each diffusion time step: Input the scene text embedding features and the scene image features after adding random noise into the diffusion network of the image generation model to generate predicted noise; wherein, the diffusion network includes a base model with frozen parameters and a trainable scene low-rank adaptation module, and the scene low-rank adaptation module acts on a preset network layer in the base model; Calculate the difference degree between the predicted noise and the random noise at the current diffusion time step, and update the parameters of the scene low-rank adaptation module based on the difference degree.
5. The method for generating an image according to claim 3, wherein The image low-rank adaptation module is obtained through training, and the training process of the image low-rank adaptation module includes: Obtain an image with image keywords; wherein, the image keywords are used to characterize the background environment in the image. Input the image keywords into the text encoder of the image generation model to encode the image keywords and obtain image text embedding features. Input the image into the image encoder of the image generation model to encode the image and obtain image features. Train the image low-rank adaptation module at each diffusion time step according to the total number of preset diffusion time steps to obtain a trained image low-rank adaptation module; wherein, the following process is executed at each diffusion time step: Input the image text embedding features and the image features after adding random noise into the diffusion network of the image generation model to generate predicted noise; wherein, the diffusion network includes a base model with frozen parameters and a trainable image low-rank adaptation module, and the image low-rank adaptation module acts on a preset network layer in the base model; Calculate the difference degree between the predicted noise and the random noise at the current diffusion time step, and update the parameters of the image low-rank adaptation module based on the difference degree.
6. The method for generating an image according to claim 1, wherein The step of inputting the image features into the image decoder of the image generation model to generate a marketing image includes: inputting the image features into the image decoder, and using the image decoder to perform layer-by-layer transposed convolution and upsampling on the image features to obtain a marketing image; wherein, the image decoder is the decoder module of a variational autoencoder.
7. The method for generating an image according to claim 1, wherein, After inputting the image features into the image decoder of the image generation model to generate a marketing image, it further includes: Analyze the marketing image based on an image quality assessment method to determine the area to be optimized in the marketing image; Perform local redrawing processing on the area to be optimized to generate a final marketing image.
8. An image generation system, characterized in that, The system includes: A data acquisition module, configured to acquire scene keywords and image keywords; wherein, the scene keywords are used to characterize the background environment in the marketing image to be generated, and the image keywords are used to characterize the target image in the marketing image to be generated; A text encoding module, configured to input the scene keywords and the image keywords into the text encoder of the image generation model for encoding to generate text embedding features; wherein, the image generation model is a text-to-image model based on a diffusion mechanism; A diffusion module for inputting the text embedding features and a preset latent space noise into a diffusion network of the image generation model for feature fusion to generate image features fused with figurative features and background features; An image generation module for inputting the image features into an image decoder of the image generation model to generate a marketing image.
9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the method for generating an image according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored thereon, which, when executed by a processor of a computer, causes the computer to execute the method for generating an image according to any one of claims 1 to 7.